#evaleval
⏳ 9 more days! We extended the submission deadline for the EvalEval Workshop @ ACL 2026.

If your work touches AI evaluation, submit!

We welcome:
✅ Regular papers
✅ ARR submissions
✅ Non-archival work
✅ Position papers
✅ Extended abstracts

📅 Deadline: March 19
🌐 evalevalai.com/events/2026-...
ACL 2026 Workshop EvalEval
Welcome to the OpenReview homepage for ACL 2026 Workshop EvalEval
openreview.net
March 11, 2026 at 4:55 AM
Evaleval was such an amazing workshop! I learned a lot and enjoyed it immensely
EvalEval recap time! This was our first full day #NeurIPS2024 workshop that specifically looked at the emergent field of model evaluations with a deep and critical lens - and it went over incredibly well! 👇
December 16, 2024 at 6:17 PM
🚀 Excited to share that our provocation paper "Evaluations Using Wikipedia without Data Contamination: From Trusting Articles to Trusting Edit Processes" has been accepted at the NeurIPS workshop Evaluating Evaluations! 🌐📚

#NeurIPS2024 #evaleval #AIEvaluation
November 28, 2024 at 9:41 AM
@jackbandy.com and I are presenting new work on gender bias in story generation at the EvalEval workshop @ #ACL2026! We know tropes in film and TV are often implicitly gendered, i.e., nothing in the trope's description marks it as gendered, but one gender is overrepresented in creative depictions.
July 4, 2026 at 7:11 PM
It's a wrap on EvalEval in San Diego! A jam packed day of learning, making new friends, critically examining the field of evals, and walking away with renewed energy and new collaborations!

We have a lot of announcements coming, but first: EvalEval will be back for #ACL2026!
December 10, 2025 at 10:56 PM
AIベンチマーク、結果の再現性確保はマジで重要。UK AISIとEvalEvalの取り組み、そのストイックさには頭が下がるな。正直、個人レベルで同等の rigor を追求するのは、なかなか骨が折れる。
September 22, 2026 at 11:01 PM
EvalEval recap time! This was our first full day #NeurIPS2024 workshop that specifically looked at the emergent field of model evaluations with a deep and critical lens - and it went over incredibly well! 👇
December 16, 2024 at 5:59 PM
Super excited for the Evaluating Evaluations workshop at @neuripsconf.bsky.social today!!! evaleval.github.io #NeurIPS2024

@msftresearch.bsky.social's FATE group, Sociotechnical Alignment Center, and friends will be presenting several papers there. See below for details...
Home - EvalEval 2024
A NeurIPS 2024 workshop on best practices for measuring the broader impacts of generative AI systems
evaleval.github.io
December 15, 2024 at 4:19 PM
📰 Hugging Face
How UK AISI and EvalEval Are Making Benchmark Results Reproducible

https://huggingface.co/blog/evaleval-a
is#IA##AI##ML#ML
September 22, 2026 at 6:01 PM
🚀 Launching Every Eval Ever: Toward a Common Language for AI Eval Reporting 🚀

A shared schema + crowdsourced repository so we can finally compare evals across frameworks and stop rerunning everything from scratch 🔧

A tale of broken AI evals 🧵👇

evalevalai.com/projects/eve...
Every Eval Ever | EvalEval Coalition
evalevalai.com
February 17, 2026 at 3:00 PM
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
September 22, 2026 at 9:45 PM
Benchmarks can be superficial, but model explanations and evaluations are fundamentally intertwined. What if we used interpretability as principled, scientific evaluation? If it met scientific standards?

arxiv.org/abs/2605.05508
coming to EvalEval at ACL as oral 🧵
1/6
June 14, 2026 at 10:00 PM
I'll be there with @narijohnson.bsky.social to present our work at the EvalEval workshop! Would love to find a time for coffee & chat while we're there :)
December 3, 2024 at 1:26 AM
With less than a week to go for NeurIPS 2024, I wanted to make a small thread to celebrate our little workshop, EvalEval, and the incredible amount of interest and love we have received in our first time organizing it.

But first, What does Evaluate Evaluations mean?
December 5, 2024 at 6:25 PM
🚨 The next edition of EvalEval Workshop is coming to
@aclmeeting.bsky.social 2026!

🧠 Workshop on "AI Evaluation in Practice: Bridging Research, Development, and Real-World Impact" 🎇

📢 CFP is now open!!! More details ⏬

📍 San Diego
📝 Submission deadline: Mar 12, 2026
February 17, 2026 at 12:21 AM


I'm at Neurips, presenting this work with @glenberman.bsky.social , Ned Cooper and @wesleydeng.bsky.social . Come check out our poster at the EvalEval workshop on Sunday or DM me to chat!
In a paper we’re workshopping this week at #NeurIPS2024, Ned Cooper, @wesleydeng.bsky.social, and Ben Hutchinson and I ask: what is the model of societal impacts reflected in efforts to evaluate GenAI systems?

Paper: arxiv.org/abs/2410.22985
Workshop: evaleval.github.io
1/5
December 12, 2024 at 7:58 PM
Join us for the Eval Eval Coalition Social at @facct.bsky.social tomorrow Tuesday June 24th from 4-4:30 pm during the coffee break! We would love to have you join us and we look forward to seeing you there!! #FAccT2025 #EvalEval
June 23, 2025 at 2:41 PM
The EvalEval partnership: a shared schema (Every Eval Ever) that requires prompt format, sampling strategy, pass/fail criteria to be reported alongside the score. An open platform (Evaluation Cards) that makes it searchable and reproducible.
September 23, 2026 at 1:53 PM
Excited to have our mini paper on the ETHICS benchmark at the @neuripsconf.bsky.social #EvalEval workshop next week! We draw on moral theory, empirical research, and prompt evaluation to argue that the benchmark lacks validity. Stay tuned for future work on the practical consequences for evals.
December 2, 2024 at 4:59 PM
Come see @shira-a.bsky.social present our co-authored paper "Evaluating Refusal" at the #NeurIPS2024 EvalEval workshop today in E Meeting Room 16! Drawing from Indigenous and feminist studies, Shira will be presenting on why refusal needs to be incorporated into evaluation frameworks of #GenAI.
December 15, 2024 at 5:45 PM
At EvalEval, we'll share a short provocation piece. We invite folks working to make AI more inclusive to get specific about *who exactly benefits* from improved models.

arxiv.org/abs/2411.09102
with my co-authors 💜 @samanthadalal.bsky.social @smhall97.bsky.social
December 13, 2024 at 8:13 AM
🚀 Technical practitioners & grads — join to build an LLM evaluation hub!
Infra Goals:
🔧 Share evaluation outputs & params
📊 Query results across experiments

Perfect for 🧰 hands-on folks ready to build tools the whole community can use

Join the EvalEval Coalition here 👇
forms.gle/6fEmrqJkxidy...
[EvalEval Infra] Better Infrastructure for LM Evals
Welcome to EvalEval Working Group Infrastructure! Please help us get set up by filling out this form - we are excited to get to know you! This is an interest form to contribute/collaborate on a research project, building standardized infrastructure for AI evaluation. Status Quo: The AI evaluation ecosystem currently lacks standardized methods for storing, sharing, and comparing evaluation results across different models and benchmarks. This fragmentation leads to unnecessary duplication of compute-intensive evaluations, challenges in reproducing results, and barriers to comprehensive cross-model analysis. What's the project? We plan to address these challenges by developing a comprehensive standardized format for capturing the complete evaluation lifecycle. This format will provide a clear and extensible structure for documenting evaluation inputs (hyperparameters, prompts, datasets), outputs, metrics, and metadata. This standardization enables efficient storage, retrieval, sharing, and comparison of evaluation results across the AI research community. Building on this foundation, we will create a centralized repository with both raw data access and API interfaces that allow researchers to contribute evaluation runs and access cached results. The project will integrate with popular evaluation frameworks (LM-eval, HELM, Unitxt) and provide SDKs to simplify adoption. Additionally, we will populate the repository with evaluation results from leading AI models across diverse benchmarks, creating a valuable resource that reduces computational redundancy and facilitates deeper comparative analysis. Tasks? As a collaborator, you would be expected to: Work towards merging/integrating popular evaluation frameworks (LM-eval, HELM, Unitxt) Group 1 - Extend to Any Task: Design universal metadata schemas that work for ANY NLP task, extending beyond current frameworks like lm-eval/DOVE to support specialized domains (e.g., machine translation) Group 2 - Save the Relevant: Develop efficient query/download systems for accessing only relevant data subsets from massive repositories (DOVE: 2TB, HELM: extensive metadata) The result will be open infrastructure for the AI research community, plus an academic publication. When? We're looking for researchers who can join ASAP and work with us for at least 5 to 7 months. We are hoping to find researchers who would take this on as an active project (8 hours+/week) in this period.
forms.gle
June 12, 2025 at 3:01 PM
Benchmark-Zahlen sind oft nicht reproduzierbar – UK AISI und EvalEval wollen das mit standardisierten Evals aendern. Erst wenn ein Dritter dasselbe Ergebnis erhaelt, ist die Zahl belastbar. #KIEvaluation

Mehr von mir: linkedin.com/in/maurice-putinas
September 23, 2026 at 9:00 AM
I'll be at @neuripsconf.bsky.social next week!
- Reach out to chat about AI regulation, data and data transparency, public resources, and mechanisms for AI accountability,
- Come to our EvalEval workshop on Sunday where we talk about what's needed to make evaluation meaningful!
evaleval.github.io
Home - EvalEval 2024
A NeurIPS 2024 workshop on best practices for measuring the broader impacts of generative AI systems
evaleval.github.io
December 5, 2024 at 10:35 PM
✨New Work✨ forthcoming at the #NeurIPS2024 #EvalEval workshop: NLP researchers have developed myriad instruments (tools, benchmarks, etc.) for measuring representational harms caused by LLM-based systems. But to what extent are these instruments actually used (and useful) in practice?
November 26, 2024 at 5:36 PM