IT lexicon AI & ML Mixture of Experts

Mixture of Experts MoE

AI & ML På svenska → Updated: 2026-05-23

Neural network where only a subset of "experts" activate per token — cheaper inference at the same capacity.

The idea is old (Jacobs 1991) but was scaled by Google's Switch Transformer (2021) and GShard. A router module picks e.g. 2 out of 8 expert FFNs per token. Result: a model with 7B active parameters but 56B total holds the knowledge of a 56B model at the inference cost of a 7B. Mistral's Mixtral 8x7B (2023), DeepSeek-V3, and GPT-4 (reportedly) all use MoE. Trade-off: memory grows (all experts must be held) but compute drops.

← Back to the lexicon