Skip to content
Rota Nacional

Radar ·

SGLang and vLLM: throughput on A6000 GPUs

Published on August 4, 2024, the report compares two serving engines with Mistral Large 2 on four A6000 GPUs: about 1,100 tokens/s with SGLang and under 750 with vLLM; performance was similar at smaller batch sizes.

On August 4, 2024, an engineer reported replacing vLLM with SGLang in a local data-generation pipeline. With Mistral Large 2 in 4-bit precision, running on four A6000 GPUs, they observed about 1,100 tokens per second with SGLang, versus under 750 with vLLM.

The report says performance was similar at smaller batch sizes. The comparison illustrates that throughput can vary with the serving engine and batch size, even when model and hardware are held constant; it does not establish that the results generalize to other configurations. Consult the original publication to verify the context and figures. If using AI to study or apply the report, avoid submitting sensitive organizational data unless your organization’s safeguards have been applied.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free