Machine Learning / Reinforcement Learning Infrastructure Engineer

Eka Robotics

Boston (MA)

On-site

USD 140,000 - 190,000

Full time

17 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Eka Robotics is seeking a Reinforcement/Machine Learning Infrastructure Engineer to design, implement, and maintain large-scale model training systems powering our robot learning efforts.

You will focus on building a streamlined developer experience and scalable tooling to accelerate research cycles, from prototyping to production training runs, collaborating closely with researchers.

Qualifications

  • BS, MS or higher in Computer Science, Computer Engineering, Machine Learning or related technical field.
  • Proven track record building ML training infrastructure, internal dev platforms, or scalable systems.
  • Hands-on experience with large-scale training using JAX (preferred), PyTorch, or TensorFlow.
  • Distributed training on cloud platforms or orchestration systems such as Kubernetes, SLURM, GCP or AWS.
  • Strong cross-functional communication and ownership mindset.
  • Experience building automated testing pipelines and ML CI/CD workflows.

Responsibilities

  • Own Training Infrastructure: design, implement, and maintain large-scale model training systems (orchestration, scheduling, checkpointing, experiment tracking).
  • Develop tooling to launch, monitor, debug, and reproduce experiments to improve researcher productivity.
  • Scale distributed training pipelines across compute clusters with researchers.
  • Manage cloud compute resources efficiently for future scaling.
  • Collaborate with researchers to support cutting-edge methods and contribute to training code.

Skills

ML training infra
Distributed systems
Deep learning frameworks
DevOps
Communication

Education

Bachelor's/Master's in CS/CE/ML

Tools

Kubernetes
SLURM
GCP/AWS

Job description

Eka Robotics

Eka Robotics is on a mission to build intelligence for the physical world - robots that are fast, general, and reliable. Our approach, grounded in physics, unlocks superhuman capabilities. We are defining the frontier of robotics research and deployment.

Our team consists of pioneers in robotics and machine learning. We are now hiring to scale our R&D effort. We are looking for hands-on individuals who are excited to help shape the future of robotics.

The Role

We are looking for a Reinforcement/Machine Learning Infrastructure Engineer to shape our training infrastructure. In this role, you will be responsible for designing, implementing, and maintaining the large-scale model training systems that power our next generation of robot learning.

We believe that world-class infrastructure is the foundation for moving research into production. You will focus on building an exceptional developer experience, creating intuitive and efficient tooling that our engineers and scientists love to use. Your work will directly accelerate our research cycles, making it effortless to test new ideas and scale successful experiments into production training runs. You will work closely with researchers to ensure our infrastructure scales seamlessly from prototyping to large-scale distributed training.

This is a hands-on, high-impact role at the intersection of machine learning, software engineering, and scalable infrastructure.

Responsibilities
  • Own Training Infrastructure: Design, implement, and maintain robust systems for large-scale model training, including job orchestration, scheduling, checkpointing, and experiment tracking.
  • Developer Experience & Tooling: Build streamlined, intuitive abstractions for launching, monitoring, debugging, and reproducing experiments, minimizing friction and maximizing productivity for our research teams.
  • Scale Distributed Training: Work closely with researchers to reliably scale reinforcement learning and machine learning pipelines across compute clusters.
  • Resource Management: Ensure efficient allocation and utilization of cloud-based compute resources while building the foundational systems needed for future scaling.
  • Collaborate with Researchers: Partner with the research team to understand their needs, build infrastructure that supports cutting-edge methods, guide best practices for training at scale, and contribute to core JAX model and training code.
Minimum Qualifications
  • Education: BS, MS or higher in Computer Science, Computer Engineering, Machine Learning or a related technical field.
  • Software Engineering: Strong software engineering fundamentals with a proven track record of building ML training infrastructure, internal developer platforms, or scalable systems.
  • Deep Learning Frameworks: Hands-on experience with large-scale training using JAX (preferred), PyTorch, or TensorFlow.
  • Distributed Systems: Familiarity with distributed training, multi-host setups, data pipelines, and managing workloads on cloud platforms or orchestration systems (e.g., Kubernetes, SLURM, GCP, AWS).
  • Communication & Ownership: Strong cross-functional communication skills, a deep ownership mindset, and a passion for building tools that improve the developer experience.
  • Infrastructure & DevOps: Experience building automated testing pipelines, CI/CD for ML workflows, and custom logging/telemetry stacks.
Preferred Qualifications
  • Domain Experience: Background in robotics, reinforcement learning or other machine learning systems.
  • Systems Design: Experience designing abstractions that balance researcher flexibility with system reliability.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning / Reinforcement Learning Engineer
Machine Learning / Reinforcement Learning Engineer

Eka Robotics • Massachusetts

On-site
USD 140,000 - 210,000
Robotics ML Training Infrastructure Engineer
Robotics ML Training Infrastructure Engineer

Eka Robotics • Boston (MA)

On-site
USD 140,000 - 190,000
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Machine Learning / Computer Vision Engineer
Machine Learning / Computer Vision Engineer

Eka Robotics • Boston (MA)

On-site
USD 140,000 - 210,000
Machine Learning Engineer (Robotics and AI Institute LLC):
Machine Learning Engineer (Robotics and AI Institute LLC):

Rai • Cambridge (MA)

On-site
USD 186,000 - 238,000
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
Machine Learning / Computer Vision Engineer
Machine Learning / Computer Vision Engineer

Eka • Boston (MA)

On-site
USD 90,000 - 130,000
Research Engineer, Infrastructure, RL Systems
Research Engineer, Infrastructure, RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Machine Learning Engineer (Robotics and AI Institute LLC):
Machine Learning Engineer (Robotics and AI Institute LLC):

Robotics and AI Institute • Cambridge (MA)

On-site
USD 186,000 - 238,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Ultra • New York (NY)

Hybrid
USD 180,000 - 240,000