Introduction
Scaling AI models with mixture of experts (MoE)
()
1. Introduction to Mixture of Experts (MoE)
What is MoE?
()
Why MoE, though?
()
The challenges of MoE
()
Final thoughts
()
2. MoE Architecture Breakdown
Intro to MoE architecture
()
MoE components
()
Gating mechanisms, part 1
()
Gating mechanisms, part 2
()
Types of MoE architectures
()
Routing trade-offs, part 1
()
Routing trade-offs, part 2
()
Fine-tuning MoEs
()
Instruction-tuning MoE
()
Final thoughts on MoE architecture
()
3. Hands-On MoE Implementation
Basic implementation of MoE, part 1
()
Basic implementation of MoE, part 2
()
Introducing load balancing
()
Hierarchical MoE: Hands-on
()
4. Applications and Future of MoE
Introduction to MoE applications
()
Multimodal MoEs
()
UniMoE
()
Google GateM2Former
()
Federated MoE
()
FedMoE
()
FedMoE-DA
()
Adaptive MoEs
()
Adap-MoE
()
SwiftMoE
()
Conclusion
Next steps in your MoE journey
()