Lead Engineer, Machine Learning

Salt Digital Recruitment

United States

On-site

USD 180,000 - 260,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Salt Digital Recruitment is seeking a Lead Engineer - Machine Learning to own the production ML systems lifecycle, from data handling and training to inference and deployment. You will design scalable training pipelines and robust evaluation frameworks, balancing performance, reliability and cost while collaborating with research and app engineering teams.

The role requires deep expertise in modern large-model workflows, strong software fundamentals, and a track record of shipping production ML

Qualifications

  • Experience shipping machine learning systems to production.
  • Strong understanding of large-model training, fine-tuning, evaluation and inference.
  • Ability to design and optimize production ML infrastructure.

Responsibilities

  • Own end-to-end ML systems from data to production deployment.
  • Build and evolve training and fine-tuning pipelines for large models.
  • Design evaluation systems for model capability, safety and real-world performance.
  • Architect high-performance inference systems with optimized latency and resource use.
  • Establish reliable production infrastructure for deploying and monitoring models.
  • Collaborate with research and engineering to translate model capabilities into product improvements.
  • Diagnose and resolve production model and system issues.
  • Make pragmatic technical trade-offs and iterate based on real-world metrics.
  • Provide technical leadership to engineers across ML systems.

Skills

ML systems
GPU training
Python
PyTorch
JAX
Distributed ML
Model evaluation
Production ML

Tools

PyTorch
JAX

Job description

About the Opportunity

We are partnering with a fast-growing technology company developing a new generation of AI-native applications designed to make everyday tasks, communication, organization and workflows more intelligent and intuitive. The team is building proactive AI experiences with a strong focus on persistent context, reliable long-running workflows and successful real-world task completion. They are looking for a Lead Engineer - Machine Learning to own the execution layer that transforms advanced research and model capabilities into reliable, scalable production systems.

About the Role

As Lead Engineer - Machine Learning, you will work across the complete model lifecycle, including data, training, evaluation, inference and deployment. This is a hands‑on technical leadership position for someone who enjoys operating at the intersection of machine learning research, systems engineering and product development. You will take ownership of production ML systems while helping establish the technical standards and infrastructure required to deploy, monitor and continuously improve large-scale AI models.

What You'll Own
  • Own end-to-end ML systems, from data and training through to evaluation, inference and production deployment.
  • Build and evolve training and fine‑tuning pipelines for large models.
  • Design evaluation systems that measure model capability, robustness, safety and real‑world product performance.
  • Architect high-performance inference systems, optimizing latency, GPU utilization, memory, cost and reliability.
  • Build scalable data pipelines supporting high-quality real‑world and synthetic training data.
  • Establish reliable production infrastructure for deploying, monitoring and continuously improving models.
  • Partner closely with research and application engineering teams to translate model capabilities into meaningful product improvements.
  • Identify, diagnose and resolve model and system issues within production environments.
  • Make pragmatic technical trade‑offs and rapidly iterate based on measurable real‑world performance.
  • Provide technical leadership and support to engineers working across machine learning systems.
What We're Looking For
  • Proven experience building and shipping machine learning systems used in production, rather than solely developing research prototypes or demonstrations.
  • Strong understanding of modern large-model training, fine‑tuning, evaluation and inference.
  • Strong software engineering and systems engineering fundamentals.
  • Experience operating ML workloads at meaningful scale, particularly within GPU-based environments.
  • Experience building reliable training pipelines, inference systems and production ML infrastructure.
  • Strong understanding of large models and their potential failure modes.
  • Strong technical judgement with the ability to navigate complex and ambiguous engineering problems independently.
  • A bias toward experimentation, measurement, iteration and shipping.
  • High standards for correctness, reliability and production quality.
  • Ability to write strong, maintainable, production‑grade code.
Technology Environment
  • Python
  • PyTorch
  • JAX
  • GPU-based training and inference systems
  • Large Language Models (LLMs)
  • Model training and fine‑tuning
  • Model evaluation
  • Distributed ML systems
  • Production inference infrastructure
  • Real‑world and synthetic training data pipelines
Success Looks Like

Success in this role means research and model capabilities consistently translate into production-ready solutions with clearly defined performance and quality targets. ML pipelines, training loops and inference systems will be stable, efficient and maintainable, with production issues detected, diagnosed and resolved quickly. You will continuously improve model and system performance through experimentation, evaluation and monitoring, ensuring improvements are measurable and ultimately enhance the end‑user experience. You will also help create an environment where engineers are aligned, technically supported and able to deliver high‑impact ML work efficiently.

Salt is acting as an Employment Agency in relation to this vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Engineer
Applied AI Engineer

Salt Digital Recruitment • United States

On-site
USD 140,000 - 190,000
Machine Learning Engineer
Machine Learning Engineer

AI Squared • Washington

On-site
USD 110,000 - 140,000
AI / ML Engineer
AI / ML Engineer

Neuron Factory • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer
Machine Learning Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 150,000 - 190,000
Machine Learning Engineer
Machine Learning Engineer

5 Star Recruitment • Newark (NJ)

On-site
USD 120,000 - 160,000
Sr. Machine Learning Engineer
Sr. Machine Learning Engineer

Insilico Search Partners • Cambridge (MA)

On-site
USD 150,000 - 230,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Harnham • Tampa (FL)

On-site
USD 150,000 - 200,000
Machine Learning Engineer
Machine Learning Engineer

ExaCare AI • New York (NY)

Hybrid
USD 110,000 - 150,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Sierracorp • San Francisco (CA)

On-site
USD 150,000 - 200,000
Machine Learning Engineer
Machine Learning Engineer

Errgo • Town of Boston (NY)

Hybrid
USD 120,000 - 160,000
Medical, dental, and vision insurance
401(k)
Equity
+2