Fine-Tuning and Serving Mixture-of-Experts Models in PyTorch
September 9, 2026 · IntelliSensei Team
Why MoE checkpoints break naive fine-tuning: how routing works, the load-balancing loss, adapter-based recipes that keep the router intact, expert-parallel sharding with FSDP2, and honest memory and throughput math for serving.
Read more