While we hurtle toward the unknown, make a coffee and spend some time in the zooo.

AllenAI Wants Everyone Else to Train MoEs

AllenAI has released Olmo-core 3, an open-source training infrastructure tailored specifically for large Mixture of Experts (MoE) models. By open-sourcing the complex scaffolding required to wrangle these sprawling architectures, they are inviting the broader community to shoulder the computational misery of frontier model training.

  • Olmo-core 3 provides open scalable training infrastructure designed to handle the notoriously finicky routing mechanics of large MoE models.
  • By lowering the engineering barrier for distributed training, AllenAI hopes to democratize an architectural approach previously hoarded by well-funded labs.
  • Developers get a peek under the hood of modern cluster management, assuming they have the hardware to actually run it.

Why should I care? Could be big
Unless you have a spare cluster of H100s lying around, appreciating this open infrastructure from afar is your safest bet.

Read the original: Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Subscribe to Ueno Zooo

The State of the Zoo, our weekly round-up of what actually mattered, is on its way. Join the list to get it first.
[email protected]
Subscribe