Machine Learning Engineer - Pre Training

Mindbeam

United States

On-site

USD 100,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mindbeam is looking for an expert to design and optimize large-scale pre-training systems for their generative AI models. You will build scalable pipelines and implement distributed training strategies while collaborating closely with researchers.

Ideal candidates possess strong Python skills, experience with large-scale training, and familiarity with ML frameworks like PyTorch and TensorFlow. This role promises impactful innovations in AI development within a collaborative environment.

Qualifications

  • 2+ years of experience with large-scale model training.
  • Strong coding skills in Python.
  • Experience with GPU scheduling and memory optimization.

Responsibilities

  • Build scalable pre-training pipelines for foundation models.
  • Implement distributed training strategies across GPUs/TPUs.
  • Collaborate with researchers to develop production-ready workflows.

Skills

Large-scale model training
Distributed systems
Python
ML frameworks (PyTorch, TensorFlow, JAX)
GPU scheduling
Memory optimization
Docker
Kubernetes

Education

Bachelor’s, Master’s, or PhD in Computer Science or related field

Job description

About Mindbeam

We are building the next-generation AI infrastructure for open source and enterprise. Our work is deeply research-oriented and passionate about developing ground-breaking innovations to take state-of-the-art AI applications to the next level.

What drives us is not only advancing technology, but empowering the people behind it. We are a community of researchers, engineers, and visionaries who believe that collaboration, curiosity, and openness fuel progress. If you’re motivated by impact and inspired to build tools that others can build upon, you’ll be in the right place.

Mission

Design and optimize large-scale pre-training systems that power Mindbeam’s generative AI models.

Role Expectations
  • Build scalable pre-training pipelines for foundation models, optimizing throughput and efficiency.
  • Implement distributed training strategies across GPUs/TPUs and high-performance clusters.
  • Collaborate with researchers to translate experimental setups into production-ready workflows.
  • Develop monitoring and fault-tolerance systems to ensure reliable large-scale training.
  • Continuously benchmark and tune performance across hardware and software stacks.
Background
  • Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or related field—or equivalent experience.
  • 2+ years of experience with large-scale model training and distributed systems.
  • Strong coding skills in Python and familiarity with ML frameworks (PyTorch, TensorFlow, JAX).
  • Experience with GPU scheduling, memory optimization, and parallelism strategies.
  • Comfort with containerized and orchestrated environments (Docker/Kubernetes).
  • Understanding of high-performance computing and networking bottlenecks.
About You

You thrive on scale and complexity. You enjoy solving system-level bottlenecks, pushing hardware and software to their limits, and working closely with researchers to accelerate cutting-edge AI development.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer - Post Training
Machine Learning Engineer - Post Training

Mindbeam • United States

On-site
USD 100,000 - 150,000
ML Engineer, Large-Scale Pre-Training
ML Engineer, Large-Scale Pre-Training

Mindbeam • United States

On-site
USD 100,000 - 140,000
Machine Learning Engineer - Kernels
Machine Learning Engineer - Kernels

Mindbeam • United States

On-site
USD 100,000 - 140,000
ML Engineer: Scale AI Deployment & Optimization
ML Engineer: Scale AI Deployment & Optimization

Mindbeam • United States

On-site
USD 100,000 - 150,000
Research, Pre-Training Science
Research, Pre-Training Science

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Unlimited PTO
Health, dental, and vision benefits
Paid parental leave
+1
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer
Research Engineer

Mind Robotics • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research, Pre-Training Science
Research, Pre-Training Science

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Paid parental leave
+1
Model Systems Engineer
Model Systems Engineer

Mind Robotics • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Machine Learning Operations Engineer
Machine Learning Operations Engineer

4Minds • Dallas (TX)

On-site
USD 130,000 - 200,000
Stock options
401(k) with company match
Unlimited PTO
+2