We're partnering with a well-funded frontier AI lab to find a Machine Learning Infrastructure Engineer who will design and build the foundational ML/AI infrastructure powering the next generation of large-scale model training and deployment.
This is a rare opportunity to join a team at the absolute cutting edge — working on systems that underpin frontier research and production AI at massive scale.
About the Role
As an ML Infrastructure Engineer, you'll own and evolve the core infrastructure stack that researchers and engineers depend on daily. You'll work at the intersection of distributed systems, high-performance computing, and machine learning — building platforms that enable training runs across thousands of accelerators, efficient model serving, and seamless experimentation at scale.
What You'll Do
- Design and build foundational ML infrastructure — training orchestration, distributed compute frameworks, GPU/TPU cluster management, and experiment tracking systems
- Optimize large-scale training pipelines — improving throughput, fault tolerance, checkpointing, and resource utilization across multi-node, multi-accelerator environments
- Develop and maintain model serving infrastructure — low-latency inference systems, model packaging, deployment pipelines, and autoscaling
- Build internal platforms and tooling — enabling researchers to iterate faster with self-service compute, data pipelines, and reproducible experiment workflows
- Collaborate closely with research teams — translating bleeding-edge ML requirements into reliable, scalable systems
- Drive infrastructure reliability — monitoring, observability, capacity planning, and incident response for mission-critical ML workloads
What You Bring
- 5+ years of software engineering experience, with significant time spent on ML infrastructure, distributed systems, or platform engineering
- Expertise in Python
- Hands-on experience with large-scale distributed training frameworks (PyTorch, JAX, DeepSpeed, Megatron, or equivalent)
- Strong background in GPU/accelerator programming, CUDA, and high-performance networking (NCCL, InfiniBand, RoCE)
- Production experience with Kubernetes, container orchestration, and cloud infrastructure (AWS, GCP, or Azure)
- Familiarity with ML experiment tracking, data pipelines, and model lifecycle management tooling
- Experience building systems that operate at thousands-of-GPUs scale is a strong plus
- Excellent problem-solving instincts and a bias toward shipping reliable, well-instrumented systems
Nice to Have
- Experience at a frontier lab, hyperscaler, or research-heavy AI organization
- Background in HPC, scientific computing, or large-scale simulation infrastructure
- Contributions to open-source ML infrastructure projects
- Familiarity with custom accelerator hardware (TPUs, Trainium, custom ASICs)