Senior ML Infrastructure Engineer

Harnham

New York (NY)

On-site

USD 150,000 - 200,000

Full time

18 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Harnham is seeking a senior, hands-on engineer to build and operate the infrastructure that enables ML teams to develop, deploy, and scale production systems.

You'll work at the intersection of software, cloud, platform engineering, and ML, focusing on reliability and operational excellence while collaborating with ML engineers and data scientists.

Qualifications

  • 7+ years of software engineering, platform engineering, DevOps, or similar.
  • Strong backend/systems engineering fundamentals with production operations experience.

Responsibilities

  • Design, build, and maintain infrastructure supporting ML training, deployment, and inference.
  • Develop and improve CI/CD pipelines, IaC, observability, and ops tooling.
  • Own production infrastructure with focus on availability, performance, scalability, and cost.
  • Build and operate cloud-native services and Kubernetes-based infrastructure for ML/data workloads.
  • Collaborate with ML/data teams to productionize models and improve deployments.
  • Create automation and platform abstractions to boost developer productivity.
  • Monitor health, troubleshoot production issues, and drive reliability improvements.
  • Participate in on-call rotations and take end-to-end ownership of critical components.
  • Establish best practices around infra, deployment, monitoring, and production ops.

Skills

Python
Scala
Distributed systems
DevOps / SRE
CI/CD

Education

Bachelor’s or Master’s degree in CS/Engineering

Tools

AWS
SageMaker
Docker
Kubernetes
Terraform
CI/CD tooling

Job description

We’re looking for a senior, hands-on engineer to build and operate the infrastructure that enables machine learning teams to develop, deploy, and scale production systems. This role sits at the intersection of software engineering, cloud infrastructure, platform engineering, and ML, with a strong emphasis on reliability and operational excellence.

You’ll work closely with ML engineers, data scientists, and software engineers to create scalable infrastructure, streamline deployment processes, and ensure production ML workloads are reliable, observable, and cost-efficient.

Key Responsibilities
  • Design, build, and maintain infrastructure supporting machine learning training, model deployment, and inference workloads.
  • Develop and improve CI/CD pipelines, infrastructure-as-code, observability, and operational tooling.
  • Own production infrastructure with a focus on availability, performance, scalability, and cost optimization.
  • Build and operate cloud-native services and Kubernetes-based infrastructure supporting ML and data workloads.
  • Collaborate with ML and data teams to productionize models and improve deployment and operational processes.
  • Create automation, internal tools, and platform abstractions that improve developer productivity and engineering workflows.
  • Monitor system health, troubleshoot production issues, and drive improvements to platform reliability.
  • Participate in on-call rotations and take end-to-end ownership of critical infrastructure and platform components.
  • Help establish best practices around infrastructure, deployment, monitoring, and production operations.
Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.
  • 7+ years of experience in software engineering, platform engineering, DevOps, SRE, or a similar discipline.
  • Strong backend/systems engineering fundamentals with significant experience operating production environments.
  • Strong Python experience and Scala microservices.
  • Proven experience building and maintaining backend, platform, or infrastructure services.
  • Hands‑on experience with cloud infrastructure and modern DevOps tooling, such as AWS, Sagemaker, Docker, Kubernetes, Terraform, and CI/CD platforms.
  • Experience designing, deploying, and supporting distributed systems in production.
  • Experience with machine learning infrastructure, model serving, data platforms, or ML deployment workflows is highly desirable.
  • Strong troubleshooting and problem-solving skills, with a proactive, operations-focused mindset.
  • Ability to navigate complex technical challenges and work effectively in a fast-moving engineering environment.
  • Strong communication and collaboration skills, with the ability to partner across engineering and data-focused teams.
Top 3 Resume Signals:

Experience building and operating real-time distributed systems at scale; ownership of production ML infrastructure and model-serving environments; and strong cloud-native engineering experience using AWS and Kubernetes.

What You’ll Bring

The ideal candidate is a strong backend/systems engineer who enjoys owning infrastructure from design through production. You’re comfortable diving into complex technical problems, automating repetitive processes, and improving the reliability of systems at scale. Experience supporting ML workloads is a major advantage, but strong platform and distributed-systems expertise is equally important.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer
ML Infrastructure Engineer

Clera • San Mateo (CA)

On-site
USD 180,000 - 240,000
AI/ML Platform Engineer
AI/ML Platform Engineer

Surge IT • Alexandria (VA)

On-site
USD 120,000 - 150,000
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
ML Engineer
ML Engineer

Blue Signal Search • United States

On-site
USD 120,000 - 180,000
Health insurance
Dental insurance
Life insurance
+1
Sr ML Engineer
Sr ML Engineer

dicedemo • Boston (AL)

On-site
USD 130,000 - 190,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Senior Machine Learning Engineer (DevOps/SRE)
Senior Machine Learning Engineer (DevOps/SRE)

Roku • Austin (TX)

On-site
USD 120,000 - 150,000
Senior Staff Engineer - Data & ML Ops Platform
Senior Staff Engineer - Data & ML Ops Platform

RAPSYS TECHNOLOGIES PTE LTD • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

ExaCare AI • New York (NY)

On-site
USD 100,000 - 140,000
Flexible PTO
Medical, dental, and vision coverage
Company off-sites
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000