Skip to content
Rota Nacional

Guides ·

A unified view of Attention-MoE and FFN-MoE

A publication describes treating pre-mixed tokens and (V·Wo) as an FFN expert to unify Attention-MoE and FFN-MoE with shared experts.

The publication presents an approach that treats pre-mixed tokens and (V·Wo) as an FFN expert. With shared experts, it unifies Attention-MoE and FFN-MoE.

According to the publication, this formulation reduces perplexity with similar compute. The reported result may help engineers compare sparse attention and FFN designs; it does not mean every model will see the same gain.

To study the proposal, identify which components are shared and which calculations are treated as experts. Then compare designs under equivalent compute conditions and record perplexity and other relevant criteria.

If you use AI to analyze or apply the material, avoid sending personal or internal data unless necessary. Remove it or use test data; check your organization’s policy and verify conclusions against the original paper.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free