Mixture of Experts

Mixture of Experts is a neural network architecture that activates selected expert networks for each input to reduce computation.

3 articles
Last mentioned

Mixture of Experts (MoE) is an AI model architecture in which a router selects only some of several expert subnetworks to process each token. Unlike a dense model, which uses all of its parameters for every token, it requires less computation per token relative to its total parameter count, although the weights of inactive experts must still be held in memory.

The idea was proposed in 1991, and large language models such as Mixtral and DeepSeek-V3 use it.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

…-weight model, meaning its trained weights are released for others to run, built as a mixture of experts that activates only part of the network for each token.

A mixture-of-experts model routes each request through only a subset of specialized sub-networks, while a dense model uses all of its parameters every time.

Mixture of Experts, or MoE, is another model design to recognize.


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.