https://itay1itzhak.github.io/
"From Feelings to Metrics" has been accepted to #COLM2026! 🎉
We turn "this model just feels better" into structured, user-aware evaluation!
#COLM people - let’s grab a coffee and vibe-test some research in person! ☕️🌴
Top models like GPT-5.1 were initially less preferred by beginners, but their win rate skyrocketed on personalized tasks (e.g., 9% to 94%). This shows how superior models are better for all users when evaluation is *user-aware*. 📈
Top models like GPT-5.1 were initially less preferred by beginners, but their win rate skyrocketed on personalized tasks (e.g., 9% to 94%). This shows how superior models are better for all users when evaluation is *user-aware*. 📈
👤 Profile: Turns user descriptions into structured profiles.
✍️ Rewrite: Personalizes prompts for specific contexts.
⚖️ Judge: Compares models using user-defined criteria.
Automation meets personalization. 🤖✨
👤 Profile: Turns user descriptions into structured profiles.
✍️ Rewrite: Personalizes prompts for specific contexts.
⚖️ Judge: Compares models using user-defined criteria.
Automation meets personalization. 🤖✨
👉 Input: Users adapt task complexity and context to mimic their daily life.
👉 Output: The judge is the user. What counts as "clear" or "good tone" is defined by the individual’s perspective, not a static definition.
👉 Input: Users adapt task complexity and context to mimic their daily life.
👉 Output: The judge is the user. What counts as "clear" or "good tone" is defined by the individual’s perspective, not a static definition.
Our findings:
❌ 86% said they’ve used a model that "felt" significantly better (or worse) than its reported scores.
✅ 82% of you are "vibe-testing" models through direct interaction.
Our findings:
❌ 86% said they’ve used a model that "felt" significantly better (or worse) than its reported scores.
✅ 82% of you are "vibe-testing" models through direct interaction.
You’re not alone. Instead of leaderboards, many of us turn to "vibe-testing" - manually comparing models to our own needs. But can we turn these feelings into a structured evaluation?
New paper: "From Feelings to Metrics" 🧵
You’re not alone. Instead of leaderboards, many of us turn to "vibe-testing" - manually comparing models to our own needs. But can we turn these feelings into a structured evaluation?
New paper: "From Feelings to Metrics" 🧵
Come chat LLM safety and evaluation, and stop by our ManagerBench poster (w/ @adisimhi)!
- Tomorrow (Friday) @ 10:30
- Poster Session 3, Pavilion 4
Come chat LLM safety and evaluation, and stop by our ManagerBench poster (w/ @adisimhi)!
- Tomorrow (Friday) @ 10:30
- Poster Session 3, Pavilion 4
This week I’ll be in New York giving talks at NYU, Yale, and Cornell Tech.
If you’re around and want to chat about LLM behavior, safety, interpretability, or just say hi - DM me!
This week I’ll be in New York giving talks at NYU, Yale, and Cornell Tech.
If you’re around and want to chat about LLM behavior, safety, interpretability, or just say hi - DM me!
Now hungry for discussing:
– LLMs behavior
– Interpretability
– Biases & Hallucinations
– Why eval is so hard (but so fun)
Come say hi if that’s your vibe too!
Now hungry for discussing:
– LLMs behavior
– Interpretability
– Biases & Hallucinations
– Why eval is so hard (but so fun)
Come say hi if that’s your vibe too!
We swap instruction datasets between models with different pretraining.
Result: Biases follow the pretrained model!
PCA clearly shows models group by pretraining base, not by instruction.
The bias “signature” stays intact, no matter the finetuning!
We swap instruction datasets between models with different pretraining.
Result: Biases follow the pretrained model!
PCA clearly shows models group by pretraining base, not by instruction.
The bias “signature” stays intact, no matter the finetuning!
We finetune the same model 3× with different seeds.
Result: Some variation in bias scores, but behavior patterns stay stable compared to MMLU variance.
✅ Aggregating across seeds reveals consistent trends.
We finetune the same model 3× with different seeds.
Result: Some variation in bias scores, but behavior patterns stay stable compared to MMLU variance.
✅ Aggregating across seeds reveals consistent trends.
- Pretraining
- Instruction tuning
- Training randomness
- 🍁 Bottom line - pretraining is the origin of bias. Finetuning? Just the messenger
#CausalInference #TrustworthyAI #NLP
- Pretraining
- Instruction tuning
- Training randomness
- 🍁 Bottom line - pretraining is the origin of bias. Finetuning? Just the messenger
#CausalInference #TrustworthyAI #NLP
🧠
Instruction-tuned LLMs show amplified cognitive biases — but are these new behaviors, or pretraining ghosts resurfacing?
Excited to share our new paper, accepted to CoLM 2025🎉!
See thread below 👇
#BiasInAI #LLMs #MachineLearning #NLProc
🧠
Instruction-tuned LLMs show amplified cognitive biases — but are these new behaviors, or pretraining ghosts resurfacing?
Excited to share our new paper, accepted to CoLM 2025🎉!
See thread below 👇
#BiasInAI #LLMs #MachineLearning #NLProc