Published on August 16, 2024, the summary reports a benchmark of Llama 3.1 8B in fp16 on an RTX 3090: token-per-second performance was described as reasonable with more than 100 simultaneous requests.
Published on August 16, 2024, a post on the Backprop blog reports an inference benchmark of Llama 3.1 8B in fp16 on an RTX 3090. According to the summary, the test found a token-per-second rate described as reasonable with more than 100 simultaneous requests.
The result offers a reference for assessing inference on a single GPU under concurrent load; this summary provides no numeric values or further configuration details. To check the finding, consult the original post and verify its parameters, workload, and reported metrics.