Skip to content
Rota Nacional

Guides ·

SimpleQA: evaluating factuality and hallucinations

The SimpleQA benchmark contains 4,000 human-written factual questions with verified, unique answers. Learn how to use this dataset to evaluate models and interpret results carefully.

SimpleQA contains 4,000 factual questions written by people, each with a single indisputable answer. Two annotators verified the reference answers.

An autograder classifies model responses as correct, incorrect, or not attempted. This structure supports comparisons on questions with verified answers, but it does not guarantee that the benchmark covers every topic or real-world use.

To evaluate a model, submit the questions and compare its responses with the references, keeping conditions consistent across models. Also record the share of questions not attempted, rather than looking only at the share answered correctly.

If you use AI to study or apply the method, do not include personal data or internal information in prompts. Prefer public questions and review the results: a benchmark score measures performance, not accuracy in every context.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free