https://itay1itzhak.github.io/
By bridging the gap between benchmarks and real-world vibe-testing, we can evaluate AI the way humans actually use it.
arxiv.org/abs/2604.14137
technion-cs-nlp.github.io/vibe-testin...
By bridging the gap between benchmarks and real-world vibe-testing, we can evaluate AI the way humans actually use it.
arxiv.org/abs/2604.14137
technion-cs-nlp.github.io/vibe-testin...
Top models like GPT-5.1 were initially less preferred by beginners, but their win rate skyrocketed on personalized tasks (e.g., 9% to 94%). This shows how superior models are better for all users when evaluation is *user-aware*. 📈
Top models like GPT-5.1 were initially less preferred by beginners, but their win rate skyrocketed on personalized tasks (e.g., 9% to 94%). This shows how superior models are better for all users when evaluation is *user-aware*. 📈
👤 Profile: Turns user descriptions into structured profiles.
✍️ Rewrite: Personalizes prompts for specific contexts.
⚖️ Judge: Compares models using user-defined criteria.
Automation meets personalization. 🤖✨
👤 Profile: Turns user descriptions into structured profiles.
✍️ Rewrite: Personalizes prompts for specific contexts.
⚖️ Judge: Compares models using user-defined criteria.
Automation meets personalization. 🤖✨
👉 Input: Users adapt task complexity and context to mimic their daily life.
👉 Output: The judge is the user. What counts as "clear" or "good tone" is defined by the individual’s perspective, not a static definition.
👉 Input: Users adapt task complexity and context to mimic their daily life.
👉 Output: The judge is the user. What counts as "clear" or "good tone" is defined by the individual’s perspective, not a static definition.
We analyzed public reports of vibe-tests—from YouTube to Reddit—to see how users evaluate models. Our analysis reveals vibe-testing recurring patterns, allowing us to formalize "vibe-testing" as a structured evaluation practice.
We analyzed public reports of vibe-tests—from YouTube to Reddit—to see how users evaluate models. Our analysis reveals vibe-testing recurring patterns, allowing us to formalize "vibe-testing" as a structured evaluation practice.
Our findings:
❌ 86% said they’ve used a model that "felt" significantly better (or worse) than its reported scores.
✅ 82% of you are "vibe-testing" models through direct interaction.
Our findings:
❌ 86% said they’ve used a model that "felt" significantly better (or worse) than its reported scores.
✅ 82% of you are "vibe-testing" models through direct interaction.
@boknilev @GabiStanovsky!
Preprint: arxiv.org/abs/2507.07186
Webpage: itay1itzhak.github.io/planted-in-...
We’d love your thoughts, critiques, and ideas 📬
Let’s talk about building more interpretable and trustworthy LLMs!
#NLProc #Bias #CognitiveAI
@boknilev @GabiStanovsky!
Preprint: arxiv.org/abs/2507.07186
Webpage: itay1itzhak.github.io/planted-in-...
We’d love your thoughts, critiques, and ideas 📬
Let’s talk about building more interpretable and trustworthy LLMs!
#NLProc #Bias #CognitiveAI
Cognitive biases are not introduced during instruction tuning.
They’re planted in pretraining and only surfaced by finetuning.
If we want fairer models, we need to look deeper into the pretraining pipeline.
Cognitive biases are not introduced during instruction tuning.
They’re planted in pretraining and only surfaced by finetuning.
If we want fairer models, we need to look deeper into the pretraining pipeline.
We swap instruction datasets between models with different pretraining.
Result: Biases follow the pretrained model!
PCA clearly shows models group by pretraining base, not by instruction.
The bias “signature” stays intact, no matter the finetuning!
We swap instruction datasets between models with different pretraining.
Result: Biases follow the pretrained model!
PCA clearly shows models group by pretraining base, not by instruction.
The bias “signature” stays intact, no matter the finetuning!
We finetune the same model 3× with different seeds.
Result: Some variation in bias scores, but behavior patterns stay stable compared to MMLU variance.
✅ Aggregating across seeds reveals consistent trends.
We finetune the same model 3× with different seeds.
Result: Some variation in bias scores, but behavior patterns stay stable compared to MMLU variance.
✅ Aggregating across seeds reveals consistent trends.
- Pretraining
- Instruction tuning
- Training randomness
- 🍁 Bottom line - pretraining is the origin of bias. Finetuning? Just the messenger
#CausalInference #TrustworthyAI #NLP
- Pretraining
- Instruction tuning
- Training randomness
- 🍁 Bottom line - pretraining is the origin of bias. Finetuning? Just the messenger
#CausalInference #TrustworthyAI #NLP