Published on May 13, 2026, these notes cover techniques and system components intended to improve language-model inference and serving.
Published on May 13, 2026, the class notes cover SmoothQuant, AWQ, INT4 inference kernels, pruning and sparsity, MoE, PagedAttention, FlashAttention, speculative decoding, and batching. Their stated scope is the use of these techniques and components to improve LLM inference and serving; the source provides no quantitative results or method comparisons.
Consult the original source to check the explanations and references for each topic before applying a technique. If you use AI to study the material or analyze organizational data, avoid submitting personal or confidential information without authorization, and follow your organization’s data-protection policy.