A benchmark suite with eight datasets evaluates vector similarity search beyond recall and latency, examining its impact on downstream tasks such as RAG and image classification.
On December 21, 2025, researchers from Alibaba and partners presented Iceberg, a task-centric benchmark suite with eight datasets. The work evaluates vector similarity search beyond recall and latency, considering its impact on downstream applications, including RAG and image classification.
The authors note that traditional benchmarks focused on recall and latency can miss important effects in applications that use search results. To verify the scope and conclusions, consult the original publication and check its reported datasets, tasks, and metrics. If you use AI to study or apply the work, avoid submitting sensitive organizational data without authorization.