Sr MLOps Engineer

Intuitive

Sunnyvale (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Intuitive is seeking an engineer to design, build, and maintain the infrastructure for the full ML lifecycle, from development to deployment. You will ensure seamless integration of models into production systems and deliver scalable value.

The role demands hands-on experience with Kubernetes, ML orchestration (Metaflow), and GPU hardware validation across diverse devices, plus strong problem-solving in a fast-paced environment.

Responsibilities

  • Bootstrap and maintain a production-grade Kubernetes cluster, including CNI networking and storage integration
  • Deploy and configure ML orchestration tooling (e.g., Metaflow) and artifact/dataset storage solutions to support reproducible ML workflows
  • Validate GPU node health and configuration across heterogeneous hardware (B200, L40S, A6000, V100), including driver/CUDA standardization and topology checks
  • Design and execute team migration playbooks, working directly with engineering teams to port workflows, migrate datasets/artifacts, and roll out tool updates
  • Write and maintain runbooks, architecture documentation, and disaster recovery procedures
  • Participate in on-call rotation and incident response for platform-level issues
  • Collaborate with IT/Security on identity integration, access control, and compliance requirements
  • Continuously evaluate and adopt infrastructure best practices for reliability, cost, and developer experience

Tools

Metaflow

Job description

Job Description

Primary Function of Position
In this role, you will be responsible for designing, building, and maintaining the infrastructure and tools necessary to support the entire machine learning lifecycle, from development to deployment. You will work closely with ML engineers and software developers across Intuitive to ensure that machine learning models are seamlessly integrated into our systems and deliver value at scale. The ideal candidate is an independent and fast-paced engineer with excellent problem-solving skills and practical working knowledge of modern ML development techniques.

Essential Job Duties

  • Bootstrap and maintain a production-grade Kubernetes cluster, including CNI networking and storage integration
  • Deploy and configure ML orchestration tooling (e.g., Metaflow) and artifact/dataset storage solutions to support reproducible ML workflows
  • Validate GPU node health and configuration across heterogeneous hardware (B200, L40S, A6000, V100), including driver/CUDA standardization and topology checks
  • Design and execute team migration playbooks, working directly with engineering teams to port workflows, migrate datasets/artifacts, and roll out tool updates
  • Write and maintain runbooks, architecture documentation, and disaster recovery procedures
  • Participate in on-call rotation and incident response for platform-level issues
  • Collaborate with IT/Security on identity integration, access control, and compliance requirements
  • Continuously evaluate and adopt infrastructure best practices for reliability, cost, and developer experience
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

Evlo AI • Seattle (WA)

On-site
USD 130,000 - 190,000
Senior MLOps Engineer: Kubernetes & ML Infra Architect
Senior MLOps Engineer: Kubernetes & ML Infra Architect

Intuitive • Sunnyvale (CA)

On-site
USD 120,000 - 180,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Confidential
Senior Machine Learning Engineer (DevOps/SRE)
Senior Machine Learning Engineer (DevOps/SRE)

Roku • Austin (TX)

On-site
USD 120,000 - 150,000
MLOps Engineer
MLOps Engineer

Codinix Consulting Services • California (MO)

On-site
USD 120,000 - 150,000
MLops Engineer
MLops Engineer

Arrayo • Massachusetts

On-site
USD 130,000 - 185,000
MLOps Engineer
MLOps Engineer

InfoVision Inc. • Irving (TX)

On-site
USD 100,000 - 130,000