Staff Platform Machine Learning Engineer – Engine

Jobtailor

New York (NY)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking an experienced platform/DevOps engineer to design and maintain infrastructure for ML training, deployment, and inference. You will own CI/CD, IaC, observability, and operational tooling to support reliable, scalable ML workloads.

Responsibilities include operating distributed systems in production, building Kubernetes- and cloud-based services, and collaborating with data scientists and engineers to optimize model deployment and data workflows.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
  • 7+ years of experience in software engineering, platform engineering, or DevOps/SRE roles.
  • Strong systems engineering background with production operations experience.
  • Hands-on experience with cloud infrastructure and tooling (AWS, Docker, Kubernetes, Terraform, CI/CD, monitoring).
  • Experience operating distributed systems in production.
  • Experience with ML infrastructure, model serving, or data platforms is a strong plus.

Responsibilities

  • Design, build, and maintain infrastructure supporting ML training, deployment, and inference workloads.
  • Own and improve CI/CD, infrastructure-as-code, observability, and operational tooling for ML systems.
  • Operate and evolve production systems with a strong focus on reliability, performance, and cost efficiency.
  • Build and maintain Kubernetes- and cloud-based services that support model execution and data workflows.
  • Partner with data scientists and engineers to support model deployment and operational needs.
  • Improve developer workflows through automation, tooling, and platform abstractions.
  • Participate in on-call and operational ownership for platform components.

Skills

Problem-Solving
Operations-First Mindset
Adaptability
Team Collaboration

Education

Bachelor’s Degree
Master’s Degree

Tools

AWS
Docker
Kubernetes
Terraform
CI/CD
Monitoring
Observability
Infrastructure-as-Code
Automation
Cloud Infrastructure

Job description

  • Design, build, and maintain infrastructure supporting ML training, deployment, and inference workloads.
  • Own and improve CI/CD, infrastructure-as-code, observability, and operational tooling for ML systems.
  • Operate and evolve production systems with a strong focus on reliability, performance, and cost efficiency.
  • Build and maintain Kubernetes- and cloud-based services that support model execution and data workflows.
  • Partner with data scientists and engineers to support model deployment and operational needs.
  • Improve developer workflows through automation, tooling, and platform abstractions.
  • Participate in on-call and operational ownership for platform components.
Requirements
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field
  • 7+ years of experience in software engineering, platform engineering, or DevOps/SRE roles
  • Strong systems engineering background with production operations experience
  • Strong experience with backend or platform services
  • Hands‑on experience with cloud infrastructure and tooling (e.g., AWS, Docker, Kubernetes, Terraform, CI/CD, monitoring)
  • Experience operating distributed systems in production
  • Experience with machine learning infrastructure, model serving, or data platforms is a strong plus
  • Strong problem‑solving skills and an operations‑first mindset
  • Ability to thrive in a fast‑paced, high‑tech environment and manage complex problems
Core Competencies

Demonstrates expertise in designing and maintaining infrastructure for machine learning workloads, with a strong focus on reliability, performance, and cost efficiency. Proficient in cloud infrastructure, CI/CD, and operational tooling to enhance developer workflows and support model deployment.

Highest-signal resume keywords
  • Cloud Infrastructure
  • Kubernetes
  • CI/CD
  • Machine Learning Infrastructure
  • Systems Engineering
ATS Optimization Keywords
Hard Skills
  • Software Engineering
  • Platform Engineering
  • DevOps
  • Backend Services
  • Production Operations
  • Distributed Systems
  • Model Serving
  • Data Platforms
  • Infrastructure-as-Code
  • Automation
Soft Skills
  • Problem‑Solving
  • Operations-First Mindset
  • Adaptability
Certifications & Qualifications
  • Bachelor’s Degree
  • Master’s Degree
Industry Keywords
  • Machine Learning
  • Observability
  • Operational Tooling
  • Data Workflows
Tools & Technologies
  • AWS
  • Docker
  • Terraform
  • Monitoring Tools
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior MLOps Architect: Scalable Cloud ML Platforms
Senior MLOps Architect: Scalable Cloud ML Platforms

Jobtailor • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI/ML Platform Engineer
AI/ML Platform Engineer

Planet Pharma • Indianapolis (IN)

On-site
USD 130,000 - 170,000
Senior Platform Engineer
Senior Platform Engineer

Jobtailor • Jersey City (NJ)

On-site
USD 140,000 - 180,000
Software Engineer, ML Platform
Software Engineer, ML Platform

Jobtailor • California (MO)

On-site
USD 140,000 - 200,000
Principal Platform Engineer, AI – Automation
Principal Platform Engineer, AI – Automation

Jobtailor • Phoenix (AZ)

On-site
USD 170,000 - 210,000
AI Engineering Technical Lead
AI Engineering Technical Lead

Jobtailor • Burbank (CA)

On-site
USD 180,000 - 240,000
AI Platform Engineer
AI Platform Engineer

Jobtailor • Massachusetts

On-site
USD 90,000 - 130,000
IT Manager – Platform Engineering, Data Science
IT Manager – Platform Engineering, Data Science

Jobtailor • Newport News (VA)

On-site
USD 180,000 - 240,000
AI Platform Delivery Director
AI Platform Delivery Director

Jobtailor • Massachusetts

On-site
USD 180,000 - 240,000
Sr. Platform Engineer, ML Infrastructure
Sr. Platform Engineer, ML Infrastructure

Insilico Search Partners • Cambridge (MA)

On-site
USD 140,000 - 210,000