TAAFT
Free mode
100% free
Freemium
Free Trial
Prompts Deals

Instella MoE 16B A3B Midtrain

By AMD
Instella-MoE-16B-A3B-Midtrain is the mid-training stage checkpoint in AMD's Instella-MoE family, a fully open Mixture-of-Experts language model with 16 billion total parameters and 2.8 billion active parameters activated per token across 64 experts (6 activated, 2 shared). Starting from the pretrain checkpoint, the model is further trained on high-quality data mixtures to refine key language capabilities before long-context extension and post-training. Training runs end-to-end on AMD Instinct MI300X and MI325X GPUs using AMD's ROCm software stack and Primus framework, and uses Gated Multi-head Latent Attention (Gated MLA) and FarSkip-Collective communication to reduce cross-expert communication overhead. This checkpoint precedes the long-context base, SFT, DPO, and reinforcement-learning stages that produce the final instruction-tuned and reasoning checkpoints in the release.
New Text Visit model
Released: July 24, 2026

Overview

Instella-MoE-16B-A3B-Midtrain is the mid-training checkpoint of AMD's Instella-MoE model, a fully open Mixture-of-Experts language model with 16 billion total parameters and 2.8 billion active parameters per token. Built on the pre-trained checkpoint, it is further trained on high-quality data mixtures to refine core language capabilities on AMD Instinct MI300X and MI325X GPUs, ahead of long-context extension and instruction tuning stages.

About AMD

Semiconductor company designing CPUs, GPUs, adaptive processors, and AI accelerators for PCs, gaming, data centers, embedded systems, and high-performance computing.

Industry: Computer Hardware Manufacturing
Company Size: 31000
Location: US
Website: amd.com
View Company Profile
Last updated: July 28, 2026
0 AIs selected
Clear selection
#
Name
Task