BF16 gradients in FlashAttention-3 may grow late in training
A reported 450-million-parameter transformer trained with BF16 FlashAttention-3 saw its gradient norm rise 1,000× after 25 billion tokens, ending with higher loss than FP32 attention and no NaNs.
According to the summary archived in Radar on 4 October 2026, an article reports that a 450-million-parameter transformer trained with FlashAttention-3 in BF16 had a 1,000× increase in gradient norm after 25 billion tokens. It finished with a higher loss than the run using FP32 attention, with no NaNs. The excerpt also mentions a mitigation: recomputing the backward pass of two attention layers in FP32. The archived text is cut off at that point, so the outcome of this mitigation should be checked in the original source.
For teams training models with fused BF16 attention, the finding is a warning about stability late in training. Rota Nacional does not train models or fix attention kernels; it offers an API compatible with OpenAI and Anthropic for inference. If anyone uses AI to study this material, they should keep logs, names, e-mail addresses and internal identifiers out of the prompt, because the platform detects CPF, CNPJ, e-mail, phone and person names before any model runs. To check relevance, consult the original article and compare its gradient-norm curves with your own training runs.