ML Infrastructure Engineer

Gridware

San Francisco (CA)

On-site

USD 190,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, Dental & Vision plans
Paid parental leave
Alternating day off (every other Monday)
Two week paid break ('Off the Grid')
Commuter allowance
Company-paid training

Job summary

Gridware in San Francisco is looking for a Senior ML Infrastructure Engineer to design, build, and maintain their ML deployment infrastructure. This role will involve creating monitoring systems and collaborating with engineering teams to improve CI/CD pipelines.

The ideal candidate should have over 5 years of experience in production ML infrastructure and proficiency in Python, AWS, and Kubernetes. Competitive benefits include health coverage, a paid two-week break, and a commuter allowance.

Qualifications

  • 5+ years of experience building production ML infrastructure.
  • Strong software engineering skills and proficiency in Python.
  • Experience with cloud platforms (AWS) and container orchestration (Kubernetes).
  • Familiarity with feature stores, model registries, or centralized metadata systems.

Responsibilities

  • Design, build, and maintain the infrastructure for ML model deployment.
  • Develop monitoring systems to track model performance and data quality.
  • Create end-to-end testing frameworks for models and pipelines.
  • Collaborate with engineering teams to integrate ML systems with infrastructure.
  • Improve CI/CD pipelines for ML workloads.

Skills

5+ years of experience building production ML infrastructure
Strong software engineering skills and proficiency in Python
Experience with cloud platforms (AWS)
Container orchestration (Kubernetes)
Familiarity with feature stores, model registries, or centralized metadata systems (i.e. MLFlow)

Job description

Role Description

As a Senior ML Infrastructure Engineer, you will work directly in the Automation org with the core ML, Ops, and Analytics teams to help improve and build out the infrastructure around model deployment and monitoring. This role is essential to helping scale out the amount of time saving’s Gridware brings to customers.

Responsibilities
  • Design, build, and maintain the infrastructure, tooling, and workflows that enable reliable, scalable deployment of ML models to production.
  • Develop monitoring and observability systems to track model performance, data drift, data quality, and overall system health.
  • Create and maintain end-to-end testing frameworks and simulation environments to validate models and pipelines prior to deployment.
  • Work closely with Data Engineering and Platform Engineering teams to ensure ML systems integrate cleanly with broader Gridware infrastructure and operational standards.
  • Improve CI/CD pipelines for ML workloads, ensuring reproducibility, safe rollout, and automated rollback strategies.
Required Skills
  • 5+ years of experience building production ML infrastructure
  • Strong software engineering skills and proficiency in Python
  • Experience with cloud platforms (AWS) and container orchestration (Kubernetes)
  • Familiarity with feature stores, model registries, or centralized metadata systems (i.e. MLFlow)

At this time, Gridware is unable to provide visa sponsorship or immigration support for this role. We’re only able to consider candidates who are currently authorized to work in the country of employment without visa sponsorship now or in the future.

This describes the ideal candidate; many of us have picked up this expertise along the way. Even if you meet only part of this list, we encourage you to apply!

Benefits
  • Health, Dental & Vision (Gold and Platinum with some providers plans fully covered)
  • Paid parental leave
  • Alternating day off (every other Monday)
  • Off the Grid, a two week per year paid break for all employees.
  • Commuter allowance
  • Company-paid training

190000 - 260000 USD a year

  • Senior ML Engineer Base Salary- $190,000-$210,000
  • Staff ML Engineer Base Salary- $245,000-$260,000.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager, Machine Learning
Senior Manager, Machine Learning

Gridware Technologies Inc. • San Francisco (CA)

On-site
USD 245,000 - 295,000
Health, Dental & Vision (Gold/Platinum
Parental leave
Alternating day off
+3
Senior Applied Scientist, On-Device ML
Senior Applied Scientist, On-Device ML

Gridware • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, Dental & Vision (Gold/Platinum
Parental leave
Alternating day off
+3
Senior ML Infrastructure Engineer: Scale Production ML
Senior ML Infrastructure Engineer: Scale Production ML

Gridware • San Francisco (CA)

On-site
USD 190,000 - 260,000
Health, Dental & Vision plans
Paid parental leave
Alternating day off (every other Monday)
+3
Technical Program Manager, Automation
Technical Program Manager, Automation

Gridware • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, Dental & Vision
Paid parental leave
Alternating day off
+3
Senior Data Engineer
Senior Data Engineer

Gridware • San Francisco (CA)

On-site
USD 180,000 - 195,000
Health, Dental & Vision
Parental leave
Alternating day off
+3
ML Infrastructure Engineer
ML Infrastructure Engineer

Strativ Group • Menlo Park (CA)

On-site
USD 250,000 - 320,000
AI Infrastructure Engineer, Model Serving Platform
AI Infrastructure Engineer, Model Serving Platform

Segment (Twilio) • San Francisco (CA)

On-site
USD 175,000 - 220,000
Comprehensive health coverage
Retirement benefits
Learning and development stipend
+2
Senior ML Engineer
Senior ML Engineer

Grid Dynamics • United States

On-site
USD 140,000 - 190,000
Flexible schedule
Medical insurance
Vision and dental
+2
Senior Machine Learning Systems Engineer
Senior Machine Learning Systems Engineer

Tensec • Springfield (VA)

On-site
USD 216,000 - 304,000
Medical insurance
Dental insurance
Vision insurance
+3
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5