Arcee AI Launches Trinity Large, a 400B Sparse MoE AI Model
- Arcee AI launched Trinity Large, a 400B-parameter sparse Mixture of Experts model with 13B active parameters per token, trained on 17 trillion tokens in 30-33 days using 2048 Nvidia B300 GPUs[1].
- Three variants released: Preview (lightly post-trained for chat), Base (full pretraining checkpoint), and TrueBase (10T token checkpoint without instruct data)[1][3].
- Model offers 2-3x faster inference than peers due to high sparsity with 256 experts and 4 active per token[1].
- Trinity Large Preview excels in creative writing, storytelling, role-play, and voice assistance[1][4].
- Compares to Meta's Llama 4 Maverick 400B and Z.ai's GLM-4.5 on benchmarks[3].
Arcee AI, a U.S.-based open-intelligence lab, has released Trinity Large, a 400 billion-parameter sparse Mixture of Experts (MoE) model. The model, with 13 billion active parameters per token, was trained on 17 trillion tokens over 30-33 days on 2048 Nvidia B300 GPUs and is available in three open-weight variants[1].
Model Architecture and Training
Trinity Large uses a sparse MoE architecture featuring 256 experts with 4 active per token, resulting in a high sparsity ratio. The company increased dense layers from 3 to 6 to maintain routing stability. Training incorporated momentum-based expert load balancing with tanh-clipped updates and per-sequence balance loss for stability. The full pretraining run completed in 33 days, excluding context extension and post-training[1]. Data curation was advanced through DatologyAI, with strict filtering and synthetic augmentation for diverse domains[1].
Variants and Capabilities
Arcee AI shipped three variants: Trinity Large Preview, lightly post-trained and chat-ready; Base, the full 17T pretraining checkpoint; and TrueBase, an early 10T checkpoint without instruct data. The Preview model supports a 512K token context window and performs well in creative writing, storytelling, role-play, chat, and real-time voice assistance. It is optimized for agent reliability, coherent multi-turn conversations, structured JSON outputs, and tool use[1][2][4]. All Trinity models, including smaller Nano (6B) and Mini (26B) variants, share similar skill profiles across sizes[2][3].
Performance and Efficiency
The model achieves 2-3x faster inference than peers in the same weight class, enabled by efficient attention and high sparsity. Benchmark tests show it compares to Meta's Llama 4 Maverick 400B and Z.ai's GLM-4.5 using base models. It is deployed via Preview API with 8-bit quantization at 128K context and is available for free download or use on platforms like OpenRouter and Kilo[1][3][4].
Further sources
The stories that matter, in one email. Free — unsubscribe anytime.