Bagel Labs is a physical AI research lab building the model layer for autonomous robots. Our models learn how the world changes and how a robot should act within it. WorldDiT brings world modeling and action generation into one architecture. Paris shows how independently trained diffusion experts can work together as one model. Our systems team makes this research reliable, traceable, and fast enough to iterate.
We care more about the systems you have made dependable than conventional credentials. If you can turn fragile research code into infrastructure people trust, we want to hear from you.
Role Overview
You will build the systems layer behind Bagel Labs' physical AI research. Researchers should be able to launch a run, understand what happened, compare it fairly, and reproduce it later. You will own the infrastructure that makes this possible across training, evaluation, data, models, and heterogeneous compute.
What You'll Do
- Build training and orchestration pipelines for WorldDiT, Paris, and related physical AI research.
- Standardize launchers, configurations, checkpointing, logs, metrics, and research artifacts.
- Build evaluation harnesses for action-conditioned world models and robot policy experiments.
- Own dataset and model lineage so every result traces back to its exact inputs, code, and configuration.
- Instrument GPU use, data throughput, training stability, routing behavior, rollout performance, and failure modes.
- Diagnose distributed failures, performance bottlenecks, and unexpected variance across hardware.
- Turn research prototypes into dependable tools that the team can run, understand, and extend.
Who You Might Be
You have made research systems reliable under real pressure. You may come from machine learning infrastructure, distributed training, research engineering, GPU systems, or data platforms. You understand that good infrastructure is not only fast. It helps researchers trust a result and learn from a failed run. You build simple tools that people use.
Desired Skills
- Hands-on experience with distributed training, GPU workloads, experiment infrastructure, or large-scale machine learning systems.
- The ability to diagnose performance, reliability, and reproducibility problems across complex workflows.
- Strong Python and Linux skills, with experience operating training workloads on clusters or cloud infrastructure.
- Good judgment about when to build a system and when a simple tool is enough.
- Clear communication and strong ownership.
What We Offer
- Competitive compensation and meaningful equity.
- Direct ownership of the systems behind Bagel Labs' physical AI research.
- Close collaboration with researchers working on WorldDiT, Paris, and new model architectures.
- Paid travel to leading machine learning, robotics, and systems conferences.