#evaluators
1000 evaluators given identical resumes & told candidate used AI. Only difference on resumes was candidate’s name: Emily vs James. Study found women using AI are evaluated as less competent than men who use AI...To evaluators, women’s AI use signaled inability, while men’s AI use signaled initiative
July 3, 2026 at 1:59 PM
#devtober day 5, sub evaluators

While I could poll the original map + camera data from an area, I could not do this with multiple sources.

Today's work was to add support for an evaluator to have "sub evaluators" which will help with interiors like dondoran and the monastery.
October 5, 2025 at 8:30 PM
Using LLMs as "evaluators" is just automatically making up data.

#acl2025
July 28, 2025 at 12:55 PM
kasane teto. woah.
she put her face on the pear sticker, and pear economy evaluators saw a near 4000% rise in pear stocks. she could, quote: "pear-ly believe it."
October 23, 2025 at 12:26 AM
Those Fifa evaluators.
November 30, 2024 at 12:17 PM
The AI industry is planning to have third-party evaluators monitor potentially risky AI models—but will these regulators have enough influence to make a real difference? Jasmine Sun reports:
AI Companies’ Shaky New Plan to Police Themselves
Third-party evaluators will be allowed inside companies to monitor the models. But how much access and influence will they really have?
bit.ly
September 25, 2026 at 9:24 PM
I think this is my neighbor a little ways away... We had someone asking about evaluators for trees on nextdoor and it matches the event almost 100% why he needed recs. LOL he absolutely is owed a ton the other person fucked up big time.
June 25, 2026 at 1:12 AM
The AI industry is planning to have third-party evaluators monitor potentially risky AI models—but will these regulators have enough influence to make a real difference? Jasmine Sun reports:
AI Companies’ Shaky New Plan to Police Themselves
Third-party evaluators will be allowed inside companies to monitor the models. But how much access and influence will they really have?
bit.ly
September 26, 2026 at 9:15 PM
There’s a bunch of research demonstrating that shorter hours and work weeks don’t harm productivity even vs. the 40-hour week, so it’s funny that the people who consider themselves the most rational evaluators of fact alive still just want to dominate their workers for purely ideological reasons
The grind culture that birthed many Big Tech companies from Google to Amazon is back.

As the AI race heats up, startups are promoting hardcore cultures like “996,” or working 9 a.m. to 9 p.m., six days a week.
Why these companies insist on a 72-hour work week
Start-ups are promoting hardcore cultures such as “996,” meaning working from 9 a.m. to 9 p.m. six days a week, as they race to compete in AI.
www.washingtonpost.com
October 21, 2025 at 12:56 AM
lol
September 22, 2026 at 5:15 PM
Nice way to test when AI can replace human evaluators & judges, it compares if LLMs align better with group consensus than individual human evaluators do

GPT-4 & Gemini pass the test in 8/10 tasks, but struggle with deep contextual understanding. Few-shot learning helps. arxiv.org/abs/2501.10970
January 26, 2025 at 9:48 PM
The most worrying part of this study:

"Older evaluators showed less gender bias than male Gen Z evaluators, who are more likely to use AI themselves. Among Gen Z males, 97% rated James as a “strong” candidate while only 76% rated Emily as “strong,” representing a 21 percentage point gender gap."
July 5, 2026 at 4:55 PM
Why are NFL teams and draft evaluators so obsessed with age, even when it comes to the QB position?

On the pod @natetice.bsky.social and I discussed how it affects the evaluation process.

🎧: apple.co/4hZC9eT

📺: youtu.be/KzLBJnvVEfU?...
April 3, 2025 at 2:08 PM
Women’s use of AI is perceived as a sign of incompetence, while men’s use of AI is seen as an indication of initiative. 🤦🏻‍♂️
Women Who Use AI Seen As Incompetent; Men Who Use AI Seen As Pragmatic
A 2026 study found that when women and men use AI tools to create identical resumes, evaluators view the women as less competent, while crediting the men with initiative.
www.forbes.com
July 2, 2026 at 8:40 PM
Mad Enough to Blog It™ www.indignity.net/the-worst-th...
March 6, 2025 at 2:27 PM
It is a job of mine to figure out how to engage students in the content. It is NOT my job to make content engaging.

There’s a difference and a lot of administrators and evaluators don’t know that.
September 10, 2025 at 12:45 AM
“The AI industry is planning to have third-party evaluators monitor potentially risky AI models—but will these regulators have enough influence to make a real difference?” Jasmine Sun reports:

www.theatlantic.com/technology/2...
AI Companies’ New Plan to Keep Themselves From Destroying Everything
Third-party evaluators will be allowed inside companies to monitor the models. But how much access and influence will they really have?
www.theatlantic.com
September 27, 2026 at 8:13 AM
Proud to share that I was able to make my NFL debut today! Unfortunately, after breaking 93 bones on my only snap, I have been ruled Dead by the NFL’s independent medical evaluators and am doubtful to return this season
September 14, 2026 at 1:45 AM
It's nice to know google likes my writing 🥰

eugeneyan.com/writing/llm-...
December 19, 2024 at 3:34 AM
The other aspect of this is that LLMs turn us from creators to evaluators and editors. But to be successful editors and evaluators, we need to know enough to (a) judge the output accurately and (b) fix the output as needed.
May 8, 2025 at 2:15 AM
NSF is perpetually aiming to broaden its pool of evaluators across geographical regions and institutional types. If you must decline an invitation to join the review process, consider nominating like-minded colleagues or former students who can contribute similar expertise. (5)
February 19, 2025 at 7:10 PM
the big ai labs should have to deal w/ third party evaluators that aren't a bunch of their lesswrongist ideological fellow travelers
September 13, 2026 at 4:42 AM
Beyond that, neither Redwood nor METR should be considered independent evaluators.
August 27, 2026 at 3:56 PM
This is a great take.

Anthropic is now, in seeking “embedded evaluators,” exactly what banks in the 2000s were in asking for oversight through credit rating agencies. . . and we all knowhow that turned out.
The first element of this strategy (external private monitors against systemic social risk) was tried in the aughts with credit rating agencies in the mortgage-backed securities market. It failed. darioamodei.com/post/we-must...
Dario Amodei — We Must Pace the Frontier
darioamodei.com
September 12, 2026 at 5:34 PM
I know what you're thinking: man, I wish @spencernusbaum.bsky.social and Andrew had a story detailing what they've heard from talent evaluators across baseball about the prospects acquired in the MacKenzie Gore trade.

Lucky for you, we came through:
They were traded for MacKenzie Gore. Here’s what MLB evaluators see.
The Post talked to several talent evaluators employed by teams across MLB, asking them about the return for MacKenzie Gore. Here’s what we learned.
www.washingtonpost.com
January 23, 2026 at 7:23 PM