SimpleQA: evaluating factuality and hallucinations
The SimpleQA benchmark contains 4,000 human-written factual questions with verified, unique answers. Learn how to use this dataset to evaluate models and interpret results carefully.
SimpleQA contains 4,000 factual questions written by people, each with a single indisputable answer. Two annotators verified the reference answers.
An autograder classifies model responses as correct, incorrect, or not attempted. This structure supports comparisons on questions with verified answers, but it does not guarantee that the benchmark covers every topic or real-world use.
To evaluate a model, submit the questions and compare its responses with the references, keeping conditions consistent across models. Also record the share of questions not attempted, rather than looking only at the share answered correctly.
If you use AI to study or apply the method, do not include personal data or internal information in prompts. Prefer public questions and review the results: a benchmark score measures performance, not accuracy in every context.