Release of Olmo-core 3
On October 1, 2026, the team introduced Olmo-core 3, a significant upgrade to the training infrastructure for mixture-of-experts (MoE) models. This version is designed to scale MoE training into the trillion-parameter range while maintaining computational efficiency, addressing the challenges faced by researchers in developing large AI models.
Evolution from Previous Versions
Olmo-core 3 builds upon the earlier Olmo versions, particularly the MoE architecture used in OlmoE, which featured 64 routed experts. The previous implementation relied on fully sharded data parallelism (FSDP), which required frequent gathering and resharing of model weights. In contrast, Olmo-core 3 adopts a distributed data parallelism (DDP) approach, allowing expert weights to remain resident on GPUs and routing inputs directly to these experts, thus enhancing efficiency.
Core Techniques and Improvements
The transition from FSDP to DDP is crucial as it reduces the overhead associated with weight management during training. Olmo-core 3 integrates several parallelism strategies to optimize performance:
- Expert Parallelism: Distributes experts across multiple GPUs, minimizing memory requirements.
- Pipeline Parallelism: Splits model layers across GPU groups, further reducing memory load.
- Distributed Optimizer: Spreads optimizer state across GPUs instead of duplicating it on each unit.
Additional optimizations include rowwise expert parallelism, which places routed data directly into expert buffers, and GPU-resident routing metadata that streamlines the process of queuing work.
Performance Benchmarks
In a benchmark test, the expert pool was expanded from 8 to 128 while maintaining the selection of four experts per token, resulting in an increase in total parameter capacity from 4.6 billion to 47 billion. Notably, training throughput decreased by less than 5%. In tests using eight NVIDIA B300 GPUs, a 47-billion-parameter MoE achieved a throughput of 52,000 tokens per second per GPU, a significant improvement over the previous implementation’s 19,400 tokens per second.
Scaling to Trillion-Parameter Models
Olmo-core 3 has been benchmarked with configurations reaching up to 1.2 trillion parameters, achieving a peak throughput of 858 TFLOP/s per GPU. An experiment utilizing DeepEP v2 demonstrated the capability to handle a configuration with 2.38 trillion total parameters, showcasing the potential scale of Olmo-core 3.
Technical Insights
The accompanying technical report outlines several key findings from the development process. For instance, a routing score can improve even when the workload becomes less balanced, a phenomenon referred to as token gerrymandering. Additionally, lowering the learning rates for experts did not yield better results, and overlapping communication and computation did not always enhance training speed, highlighting the complexity of optimizing training processes.
Open Source and Future Directions
Olmo-core 3 is released as open-source code, allowing researchers to inspect, adapt, and experiment with the framework. This commitment to open model development ensures that the infrastructure and training methodologies are accessible, fostering innovation in the field of AI.
Original source: Hugging Face Blog