TAAFT
Free mode
100% free
Freemium
Free Trial
Prompts Deals

Instella MoE 16B A3B Base

By AMD
Instella-MoE-16B-A3B-Base is the long-context extension checkpoint that AMD uses as the final base model in its Instella-MoE family, a fully open Mixture-of-Experts language model with 16 billion total parameters and 2.8 billion active parameters activated per token across 64 experts (6 activated, 2 shared). Building on the mid-trained checkpoint, it receives additional long-context training to extend its ability to process and reason over longer sequences, and is evaluated on standard and long-context benchmarks such as HELMET and RULER. Training runs end-to-end on AMD Instinct MI300X and MI325X GPUs using AMD's ROCm software stack and Primus framework, with Gated Multi-head Latent Attention (Gated MLA) and FarSkip-Collective communication for efficient large-scale training. This base checkpoint is the foundation for the SFT, DPO, and reinforcement-learning stages that produce the final instruction-tuned and reasoning checkpoints in the release.
New Text Visit model
Released: July 24, 2026

Overview

Instella-MoE-16B-A3B-Base is the long-context checkpoint AMD designates as the final base model in its Instella-MoE family, a fully open Mixture-of-Experts language model with 16 billion total parameters and 2.8 billion active parameters per token. It extends the mid-trained checkpoint with long-context training so it can process and reason over longer sequences, and serves as the foundation for the SFT, DPO, and RL post-trained checkpoints.

About AMD

Semiconductor company designing CPUs, GPUs, adaptive processors, and AI accelerators for PCs, gaming, data centers, embedded systems, and high-performance computing.

Industry: Computer Hardware Manufacturing
Company Size: 31000
Location: US
Website: amd.com
View Company Profile
Last updated: July 28, 2026
0 AIs selected
Clear selection
#
Name
Task