Itay Itzhak @ COLM 🍁
itay-itzhak.bsky.social
Itay Itzhak @ COLM 🍁
@itay-itzhak.bsky.social
NLProc, deep learning, and machine learning. Ph.D. student @ Technion and The Hebrew University.
https://itay1itzhak.github.io/
We vibe-tested our own paper, but the reviewers gave us the actual metrics:
"From Feelings to Metrics" has been accepted to #COLM2026! 🎉

We turn "this model just feels better" into structured, user-aware evaluation!

#COLM people - let’s grab a coffee and vibe-test some research in person! ☕️🌴
July 9, 2026 at 12:57 PM
Results: standard evaluation mask model quality. 🎭

Top models like GPT-5.1 were initially less preferred by beginners, but their win rate skyrocketed on personalized tasks (e.g., 9% to 94%). This shows how superior models are better for all users when evaluation is *user-aware*. 📈
April 23, 2026 at 11:35 AM
Can we scale it? We built a pipeline mirroring vibe-testing structure:
👤 Profile: Turns user descriptions into structured profiles.
✍️ Rewrite: Personalizes prompts for specific contexts.
⚖️ Judge: Compares models using user-defined criteria.
Automation meets personalization. 🤖✨
April 23, 2026 at 11:35 AM
The anatomy of a vibe-test:
👉 Input: Users adapt task complexity and context to mimic their daily life.
👉 Output: The judge is the user. What counts as "clear" or "good tone" is defined by the individual’s perspective, not a static definition.
April 23, 2026 at 11:35 AM
Why do we "vibe-test" and ignore leaderboards? We ran a survey to find out.

Our findings:
❌ 86% said they’ve used a model that "felt" significantly better (or worse) than its reported scores.
✅ 82% of you are "vibe-testing" models through direct interaction.
April 23, 2026 at 11:35 AM
Ever used a top-ranked LLM that just... felt wrong for you?

You’re not alone. Instead of leaderboards, many of us turn to "vibe-testing" - manually comparing models to our own needs. But can we turn these feelings into a structured evaluation?

New paper: "From Feelings to Metrics" 🧵
April 23, 2026 at 11:35 AM
In Rio for #ICLR2026 🇧🇷 and already had my first açaí! 🍧
Come chat LLM safety and evaluation, and stop by our ManagerBench poster (w/ @adisimhi)!
- Tomorrow (Friday) @ 10:30
- Poster Session 3, Pavilion 4
April 23, 2026 at 11:32 AM
Had a blast at CoLM! It really was as good as everyone says, congrats to the organizers 🎉
This week I’ll be in New York giving talks at NYU, Yale, and Cornell Tech.
If you’re around and want to chat about LLM behavior, safety, interpretability, or just say hi - DM me!
October 13, 2025 at 4:19 PM
In Vienna for #ACL2025, and already had my first (vegan) Austrian sausage!

Now hungry for discussing:
– LLMs behavior
– Interpretability
– Biases & Hallucinations
– Why eval is so hard (but so fun)
Come say hi if that’s your vibe too!
July 27, 2025 at 6:11 AM
🔄 Step 2: Cross-tuning.
We swap instruction datasets between models with different pretraining.
Result: Biases follow the pretrained model!

PCA clearly shows models group by pretraining base, not by instruction.
The bias “signature” stays intact, no matter the finetuning!
July 15, 2025 at 1:38 PM
🎲 Step 1: Training randomness.
We finetune the same model 3× with different seeds.
Result: Some variation in bias scores, but behavior patterns stay stable compared to MMLU variance.
✅ Aggregating across seeds reveals consistent trends.
July 15, 2025 at 1:38 PM
🧪 We introduce a two-step causal framework to disentangle the effects of:
- Pretraining
- Instruction tuning
- Training randomness

- 🍁 Bottom line - pretraining is the origin of bias. Finetuning? Just the messenger
#CausalInference #TrustworthyAI #NLP
July 15, 2025 at 1:38 PM
🚨New paper alert🚨

🧠
Instruction-tuned LLMs show amplified cognitive biases — but are these new behaviors, or pretraining ghosts resurfacing?

Excited to share our new paper, accepted to CoLM 2025🎉!
See thread below 👇
#BiasInAI #LLMs #MachineLearning #NLProc
July 15, 2025 at 1:38 PM