#ragtruth
Track these benchmarks to evaluate Agentic RAG retrieval quality: NeedleInAHaystack, BeIR, FRAMES, RAGTruth, RULER, MMNeedle, FEVER.
September 14, 2026 at 8:29 AM
A new study shows the first token of a hallucinated span offers a stronger detection signal than later tokens, using RAGTruth token‑level data. This pattern holds for model sizes. https://getnews.me/first-hallucinated-token-shows-stronger-signal-for-llm-error-detection/ #llmhallucination #ragtruth
October 6, 2025 at 3:03 PM
Study shows the first token of a hallucinated span has a stronger logit deviation, yielding higher detection rates than later tokens, based on the RAGTruth extended dataset. Read more: https://getnews.me/first-hallucination-tokens-show-stronger-signals-than-later-tokens/ #hallucination #token
October 3, 2025 at 9:09 AM
HalluGuard, a 4 billion‑parameter model, achieves 84% balanced accuracy on the RAGTruth benchmark, matching a 7 billion‑parameter system. It will be released under Apache 2.0. Read more: https://getnews.me/haluguard-model-reduces-hallucinations-in-retrieval-augmented-generation/ #halluguard #ai
October 3, 2025 at 1:52 AM
Turk-LettuceDetect's Turkish ModernBERT scored 0.7266 F1 on token-level hallucination detection, handles up to 8,192 tokens, and comes with a 17,790‑instance RAGTruth dataset. https://getnews.me/turk-lettucedetect-ai-model-enhances-turkish-hallucination-detection/ #turkishnlp #ai
September 25, 2025 at 12:50 AM
Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, Kashun Shum, Randy Zhong, Juntong Song, Tong Zhang
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
https://arxiv.org/abs/2401.00396
May 20, 2024 at 4:08 AM
Yuanhao Wu, Juno Zhu, Siliang Xu, Kashun Shum, Cheng Niu, Randy Zhong, Juntong Song, Tong Zhang
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models. (arXiv:2401.00396v1 [cs.CL])
http://arxiv.org/abs/2401.00396
January 2, 2024 at 3:03 AM
introduce a perturbed multi-hop QA dataset with induced hallucinations. Via supervised fine-tuning on our dataset, we achieve better recall with a 7B model than GPT-4o on the RAGTruth hallucination detection benchmark and offer competitive performance [4/5 of https://arxiv.org/abs/2505.04844v1]
May 9, 2025 at 5:58 AM
HDMBench. Experimental results demonstrate that HDM-2 out-performs existing approaches across RagTruth, TruthfulQA, and HDMBench datasets. This work addresses the specific challenges of enterprise deployment, including computational efficiency, domain [4/5 of https://arxiv.org/abs/2504.07069v1]
April 10, 2025 at 6:00 AM
Evaluations on the RAGTruth corpus demonstrate an F1 score of 79.22% for example-level detection, which is a 14.8% improvement over Luna, the previous state-of-the-art encoder-based architecture. Additionally, the system can [5/6 of https://arxiv.org/abs/2502.17125v1]
February 25, 2025 at 6:30 AM
Building on ModernBERT's extended context capabilities (up to 8k tokens) and trained on the RAGTruth benchmark dataset, our approach outperforms all previous encoder-based models and most prompt-based models, while being [3/6 of https://arxiv.org/abs/2502.17125v1]
February 25, 2025 at 6:30 AM
今日のZennトレンド

RAGのウソを検知する新手法(LLM-as-a-Judgeを超えて)
この記事は、RAGのハルシネーションを高速に検出する手法「LettuceDetect」を紹介しています。
これは、軽量な言語モデルModernBERTをRAGTruthで学習させたもので、従来のLLMを用いた手法と同等の性能ながら高速かつ低コストで幻覚検知が可能です。
RAGシステムにおけるハルシネーション対策として有効な選択肢となりえます。
RAGのウソを検知する新手法(LLM-as-a-Judgeを超えて)
本記事では、RAGの幻覚(ハルシネーション)を検出するための「LettuceDetect」という手法について、ざっくり理解します。株式会社ナレッジセンスは、エンタープライズ企業向けにRAGを提供しているスタートアップです。 この記事は何この記事は、RAGのハルシネーションを高速に検出するための「LettuceDetect」の論文[1]について、日本語で簡単にまとめたものです。https://arx
zenn.dev
March 11, 2025 at 9:15 PM