← Tool Radar
dataset

Research on LLM Disagreement on Factual Claims

A novel dataset with thin documentation, good for LLM consistency research but limited practical impact.

5.3Overall
Utility4
Onboarding7
Craft4
Niche fit6
Longevity6

Five dimensions scored independently; overall is a weighted average. Scores are only comparable within this same rubric.

Good for

Academic researchers studying LLM factual disagreement

Not for

Engineering teams looking for ready‑to‑use production tools

Project description

zenodo.org

Alternatives

TruthfulQAOpenAI EvalsMMLU
Visit site1 · Stars at evalEvaluated 2026-08-12

Similar tools