Skip to content
Rota Nacional

Radar ·

MoE kernels explore AWS EFA

On November 6, 2025, Perplexity described expert-parallelism kernels for serving MoE models across multiple AWS GPUs with EFA, without GPUDirect Async.

On November 6, 2025, Perplexity presented an expert-parallelism kernel approach for serving large MoE models across multiple AWS GPUs using EFA. According to the description, GPU dispatch and combine operations group tokens into RDMA writes; a host proxy thread coordinates transfers alongside grouped GEMM.

The proposal addresses multi-node MoE serving when EFA does not provide GPUDirect Async. To check the details and technical context, consult the original publication and verify its stated architecture and conditions. If you use AI to study or apply the material, avoid submitting sensitive organizational data and follow internal data-handling policies.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free