The report gives an audio reasoning model a 92% score on Big Bench Audio and compares time to first token with and without its reasoning step.
Artificial Analysis reports that Gemini 2.5 Native Audio Thinking scored 92% on Big Bench Audio, a benchmark of 1,000 audio questions adapted from Big Bench Hard. According to the report, the model with thinking took an average of 3.87 seconds to produce its first token; without thinking, it took 0.63 seconds. These figures offer a comparison of reasoning and latency for speech-to-speech models, but do not by themselves establish which model suits a particular task.
If you use AI to study or apply these results, avoid submitting recordings or documents containing personal data unless necessary. Rota Nacional detects personal data before execution and applies the organization’s policy, such as replacing it with markers, removing it, or blocking the request. Audio transcription already passes through this barrier; voice features are in preview. The platform does not claim that these controls solve the performance problem measured by the benchmark.