Published on August 4, 2024, the report compares two serving engines with Mistral Large 2 on four A6000 GPUs: about 1,100 tokens/s with SGLang and under 750 with vLLM; performance was similar at smaller batch sizes.
On August 4, 2024, an engineer reported replacing vLLM with SGLang in a local data-generation pipeline. With Mistral Large 2 in 4-bit precision, running on four A6000 GPUs, they observed about 1,100 tokens per second with SGLang, versus under 750 with vLLM.
The report says performance was similar at smaller batch sizes. The comparison illustrates that throughput can vary with the serving engine and batch size, even when model and hardware are held constant; it does not establish that the results generalize to other configurations. Consult the original publication to verify the context and figures. If using AI to study or apply the report, avoid submitting sensitive organizational data unless your organization’s safeguards have been applied.