AI2 Launches Olmo-core 3: Open MoE Training Framework Scaling to 2.38 Trillion Parameters
AI2 released Olmo-core 3, an open-source training framework for mixture-of-experts models scaling to 2.38 trillion parameters. It achieves 2.7× higher throughput than the previous FSDP implementation on 8 NVIDIA B300 GPUs, with MXFP8 boosting throughput 21% and reducing memory from 103GB to 95GB.
Executive Overview
The Allen Institute for AI (AI2) has officially announced Olmo-core 3 via the Hugging Face blog, a major upgrade to its large language model development framework. This open training system is redesigned for mixture-of-experts (MoE) models and is available immediately on GitHub with a technical report and interactive demo. It is engineered to scale to trillion-parameter ranges while maintaining computational efficiency. Key benchmarks show a 2.7× throughput improvement over the previous FSDP-based implementation on 8 NVIDIA B300 GPUs, with MXFP8 delivering a 21% throughput gain and reducing active memory from 103GB to 95GB.
📊 Official Technical Specifications & Data Sheet
| Technical Aspect | Confirmed Official Data |
|---|---|
| 💰 Pricing & Usage Cost | Fully open source and free on GitHub. No subscription tiers or per-token pricing; users bear the cost of renting GPU units. Official benchmarks were run on NVIDIA B300. |
| 🌐 Platforms & Immediate Availability | Available immediately via GitHub (open code) and Hugging Face (technical report + interactive demo). Runs on NVIDIA B300 GPU clusters. Adaptable to different hardware per official statement. |
| ⚡ Performance & Speed Metrics | 52,000 tokens/second/GPU on 8 B300 units vs. 19,400 previously (2.7× throughput). MXFP8 boosts throughput by 21% over BF16. Peak measured throughput: 858 TFLOP/s/GPU on 512 B300 units. Scaling expert pool from 8 to 128 reduced throughput by less than 5%. |
| 🛡️ Security & Breach Resistance | The official announcement did not include security standards or prompt injection protection; focus is on training architecture and computational efficiency. |
| 🧠 Context Window | No context window size specified in this announcement. The statement mentions that the next generation of Olmo will be trained with the "longest context window" in Olmo history, without a specific number. |
| 🌍 Arabic Language & Regional Support | The framework is a training architecture, not a ready model; language support depends on the training data chosen by the developer. Available globally via GitHub without geographic restrictions, enabling Arab researchers to build custom Arabic MoE models. |
Deep-Dive Features & Architecture
Olmo-core 3 represents an architectural shift in how MoE models are trained within the Olmo family. While the previous OlmoE used a MoE architecture with 64 routed experts, and Olmo 3 adopted a dense architecture, Olmo-core 3 introduces a training system specifically designed for much larger MoE models. The most notable shift is the move from fully sharded data parallelism (FSDP) to distributed data parallelism (DDP), where weights remain resident on GPUs and data is routed to them, eliminating repeated weight gathering per micro-batch.
The framework integrates three techniques for model distribution:
- Expert parallelism to distribute experts across GPUs
- Pipeline parallelism to split layers across GPU groups
- Distributed optimizer to distribute optimizer state instead of replicating it
It also adds operational optimizations: rowwise expert parallelism to place routed data directly into expert input buffers, GPU-resident routing to keep routing data on GPU, and grouped GEMM to combine small expert computations. The framework supports MXFP8, a low-precision numeric format that reduces computation and data movement between GPUs.
In an extended benchmark, the team increased the expert pool from 8 to 128 while selecting only 4 experts per token, keeping active parameters fixed at approximately 3.2 billion. Total parameter capacity grew from 4.6 billion to 47 billion, while training throughput dropped by less than 5%—strong evidence of scaling efficiency.
Benchmark & Competitive Performance
Olmo-core 3 is directly compared to NVIDIA Megatron-Core, the established choice for training large MoE models. Olmo-core 3's advantage is delivering a complete MoE training package within the open framework behind Olmo, with throughput optimizations over the previous FSDP-based implementation. In an initial test on 8 NVIDIA B300 units, a 47-billion-parameter MoE model achieved 52,000 tokens/second/GPU versus 19,400 in the old implementation—a 2.7× throughput gain. On 4 B300 units, MXFP8 increased throughput by 21% compared to BF16 while reducing peak active memory from 103GB to 95GB, with most of the improvement coming from feed-forward computations and inter-expert data transfer. At a scale of 512 B300 units, peak observed throughput reached 858 TFLOP/s/GPU.
The framework was also tested on a 1.2-trillion-parameter model with 58.36 billion active parameters per token on 512 NVIDIA B300 GPUs. A DeepEP v2 test reached a configuration with 2.38 trillion total parameters, which was a short capacity test rather than a full training run. These results position Olmo-core 3 as a highly scalable and efficient open alternative for large-scale MoE training.
Industry Impact & Enterprise Adoption
Olmo-core 3 lowers the barrier for organizations and researchers seeking to train large MoE models without relying on proprietary frameworks. By open-sourcing the training stack, AI2 enables enterprises to adapt the framework to their own hardware and data, potentially accelerating innovation in multilingual and domain-specific models. The framework's compatibility with NVIDIA B300 GPUs and support for MXFP8 align with industry trends toward lower-precision, higher-efficiency training. For the Arabic NLP community, Olmo-core 3 provides a foundation to build custom MoE models tailored to Arabic, leveraging the framework's scalability and open availability.
Conclusion
Olmo-core 3 marks a significant step forward in open-source AI infrastructure, offering a scalable, efficient, and fully accessible training framework for MoE models up to 2.38 trillion parameters. With proven throughput gains and memory optimizations, it empowers developers and researchers to push the boundaries of large-scale language model training. As AI2 continues to advance the Olmo family, Olmo-core 3 is poised to become a cornerstone for next-generation open MoE development.
Media Source: Hugging Face | البيان الرسمي للشركة: المصدر الأصلي | Fact Verification & Analysis: AI Tools Oasis
Frequently Asked Questions
Olmo-core 3 is an open-source training framework for mixture-of-experts (MoE) models developed by the Allen Institute for AI (AI2), officially announced via the Hugging Face blog. It targets training MoE models up to trillion-parameter scale while maintaining computational efficiency, and serves as the foundation for the next generation of Olmo models.
On 8 NVIDIA B300 GPUs, a 47-billion-parameter MoE model processed 52,000 tokens/second/GPU versus 19,400 in the previous implementation—a 2.7× throughput gain. MXFP8 increased throughput by 21% compared to BF16 while reducing active memory from 103GB to 95GB. Peak measured throughput reached 858 TFLOP/s/GPU on 512 B300 GPUs.
The framework was tested on a 1.2-trillion-parameter model with 58.36 billion active parameters per token on 512 NVIDIA B300 GPUs. A DeepEP v2 test also reached a configuration with 2.38 trillion total parameters, which was a short capacity test rather than a full training run.
Yes, Olmo-core 3 is fully open source and available on GitHub with a technical report and interactive demo. Researchers and developers can use it to train their own MoE models, adapt it to different hardware, and experiment with routing and parallelism strategies.
The framework relies on three main techniques: expert parallelism to distribute experts across GPUs, pipeline parallelism to split model layers, and distributed optimizer to distribute optimizer state. It also uses rowwise expert parallelism, GPU-resident routing, grouped GEMM, and MXFP8 support to reduce data movement.

AI Tools Oasis Team
Bringing you the latest news and analysis in the world of Artificial Intelligence with accuracy and credibility. Follow us for all updates.