Senior Machine Learning Ops Engineer

Jobtailor

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor seeks an experienced ML Infrastructure Engineer to architect and operate scalable cloud-based MLOps platforms for end-to-end model lifecycle management.

You will own technical strategy, drive cross‑functional initiatives, and mentor engineers while ensuring reliability, observability, and cost efficiency across deployments.

Qualifications

  • 6+ years of software engineering with Python and Linux.
  • Experience in infrastructure, cloud and/or MLOps.
  • Strong collaboration and communication skills.
  • Bachelor's or Master's in CS/EE or related field.

Responsibilities

  • Architect and deploy scalable cloud-based MLOps platforms and workflows for AI/ML models.
  • Own the technical strategy and evolution of ML infrastructure across teams.
  • Build automated systems to ship models rapidly with reliability and observability.
  • Define infra optimization strategies balancing performance and cost.
  • Evaluate new tools and lead their adoption to improve workflows.
  • Establish engineering best practices for ML infrastructure and CI/CD.
  • Provide mentorship, lead code reviews, and raise team capabilities.
  • Partner with ML engineers, researchers, and product teams to translate requirements into scalable solutions.
  • Drive complex infrastructure projects from strategy to production ownership.

Skills

Python
Linux
MLOps
Cloud infrastructure
Software engineering

Education

Bachelor's or Master's in CS/EE

Tools

AWS
GCP
CI/CD

Job description

Responsibilities
  • Architect, design, deploy, and operate scalable cloud-based MLOps platforms and workflows that enable efficient training, evaluation, deployment, monitoring, and lifecycle management of AI/ML models.
  • Own the technical strategy and evolution of ML infrastructure, identifying architectural bottlenecks and driving cross‑functional initiatives to improve developer productivity, experimentation velocity, scalability, and operational efficiency.
  • Build robust, reliable, and automated systems that enable teams to ship new models and features rapidly while maintaining high standards for quality, reproducibility, observability, security, and production reliability.
  • Define and implement infrastructure optimization strategies that balance performance, scalability, reliability, and cost across cloud and compute resources.
  • Evaluate emerging tools, technologies, and industry best practices in MLOps, cloud infrastructure, and ML systems, and lead their adoption where they can meaningfully improve ML development and production workflows.
  • Establish engineering best practices for ML infrastructure, including system design, code quality, testing, CI/CD, monitoring, documentation, and operational readiness.
  • Provide technical leadership and mentorship to engineers, lead design and code reviews, and help raise the engineering quality and technical capabilities of the broader team.
  • Partner closely with ML engineers, researchers, data engineers, and product teams to translate evolving AI/ML requirements into scalable and maintainable infrastructure solutions.
  • Drive complex, ambiguous infrastructure projects from technical strategy and architecture through implementation, production deployment, and long‑term operational ownership.
Requirements
  • A Bachelors Degree or a Masters Degree in Computer Science, Electrical Engineering, or a related field.
  • Core Skills: General Software Engineering skills with 6+ years of programming experience in Python and the surrounding tooling ecosystem along with familiarity in Linux and expertise in infrastructure, cloud and/or MLOps.
  • Personal Attributes: Team player, good communication skills, self‑starter.
  • Strong teamwork and communication skills to collaborate with cross‑functional teams, including ML and software engineers.
  • Nice to Have: Experience building MLOps pipelines for deep learning based perception solutions on AWS or GCP.
Core Competencies

Demonstrates expertise in architecting and deploying scalable cloud-based MLOps platforms, with a strong focus on optimizing ML infrastructure for performance, reliability, and cost. Proven ability to lead cross‑functional teams and mentor engineers while implementing best practices in software engineering and ML workflows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Ops Engineer
ML Ops Engineer

Hayden AI Technologies, Inc. • San Francisco (CA)

Hybrid
USD 100,000 - 135,000
ML Operations Engineer
ML Operations Engineer

NextGen Healthcare • Georgia

On-site
USD 80,000 - 120,000
Senior MLOps Architect: Scalable Cloud ML Platforms
Senior MLOps Architect: Scalable Cloud ML Platforms

Jobtailor • San Francisco (CA)

On-site
USD 140,000 - 210,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
Machine Learning Architect
Machine Learning Architect

Tiger Analytics Inc. • New Jersey

On-site
USD 120,000 - 160,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

LotusLynx • Pittsburgh

On-site
USD 120,000 - 150,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

ExaCare AI • New York (NY)

On-site
USD 100,000 - 140,000
Flexible PTO
Medical, dental, and vision coverage
Company off-sites
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Sr. Machine Learning Ops Engineer
Sr. Machine Learning Ops Engineer

Hayden AI • San Francisco (CA)

On-site
USD 170,000 - 260,000