Summary

AMD published the Instella-MoE technical report (arXiv 2609.00791, Wed 2 Sep digest): a fully open 16B-total / 2.8B-active MoE trained from scratch on AMD Instinct MI300X and MI325X GPUs. It combines Gated Multi-head Latent Attention and FarSkip-Collective connectivity, and was built through a multi-stage pipeline: pre-training, mid-training, long-context extension, SFT with feedback-driven data curation, DPO, and RL with Multi-Teacher On-Policy Distillation.

Why it matters
The open-weight frontier is currently a Chinese-lab story with a US response forming via Nvidia/Poolside; Instella-MoE adds a third vector — open weights proven out on non-NVIDIA silicon at scale. At 16B it is not frontier-tier, but it is a credible reference for anyone evaluating MI300X/MI325X training stacks, and MLA-class architecture choices show AMD targeting the same serving-efficiency conventions as DeepSeek-family models.
Technical details
Arxiv 2609.00791, announced in the Wed 2 Sep 2026 digest
Scale 16B total / 2.8B active MoE
Hardware trained entirely on AMD Instinct MI300X and MI325X
Architecture Gated Multi-head Latent Attention (Gated MLA); FarSkip-Collective connectivity
Pipeline pre-training -> mid-training -> long-context extension -> SFT (feedback-driven data curation) -> DPO -> RL with Multi-Teacher On-Policy Distillation
Openness fully open (technical report + model family on Hugging Face under AMD org)
Tags
open-weightsamdmoemlatraining-infrastructureinstinct