dataset
Research on LLM Disagreement on Factual Claims
A novel dataset with thin documentation, good for LLM consistency research but limited practical impact.
5.3Overall
Utility4
Onboarding7
Craft4
Niche fit6
Longevity6
Five dimensions scored independently; overall is a weighted average. Scores are only comparable within this same rubric.
Good for
Academic researchers studying LLM factual disagreement
Not for
Engineering teams looking for ready‑to‑use production tools
Project description
zenodo.org
Alternatives
TruthfulQAOpenAI EvalsMMLU