#LLMsAsJudges
What is the optimal temperature for LLMs-as-judges?

doi.org/10.48550/arX...

Why read? The authors find that higher temperatures may help in complex evaluation tasks.
What's next? Exploration of optimal temperature for different tasks.

#ShortReview #AI #LLM #Temperature #LLMsAsJudges
The Necessity of Setting Temperature in LLM-as-a-Judge
Using large language models (LLMs) as judges for evaluating model outputs has emerged as an important paradigm for automated evaluation. However, the choice of decoding temperature in LLM-as-a-judge s...
doi.org
September 30, 2026 at 8:05 AM
📚 Just published an open-access paper with Magdalena Zdunek in @catclassquarterly.bsky.social, exploring the use of LLMs as judges in human subject indexing. We show the potential of this approach, but also its weaknesses.

doi.org/10.1080/0163...

#AI #LLM #LLMsAsJudges #SubjectIndexing #KO
Large Language Models as Judges for Human Subject Indexing
The study evaluates LLMs as judges for human subject indexing by comparing their performance against a manually curated gold standard across descriptions of forty Polish library and information sci...
doi.org
August 17, 2026 at 7:09 AM