Staff Machine Learning Systems & Reliability Engineer (Moveworks)

ServiceNow

Mountain View (CA)

On-site

USD 250,000 - 320,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Generous family leave
Annual learning stipend
Flexible PTO
Competitive retirement plan
Paid volunteer time

Job summary

ServiceNow in Mountain View seeks a hands-on Staff Engineer to move ML models and agentic workflows from prototypes into secure, observable production systems. You will work at the intersection of ML systems, platform engineering, and SRE, partnering with ML, data, product, and infrastructure teams to create a paved path from experimentation to production.

The role covers data preparation, training, experiment tracking, evaluation, artifact management, serving, monitoring, and retraining with

Qualifications

  • Proven track record building and operating distributed production systems.
  • Experience with ML lifecycle: training, deployment, serving, monitoring, retraining.

Responsibilities

  • Move ML models and agentic workflows from prototype to production.
  • Design and operate complete ML lifecycle path: data prep, training, evaluation, serving, retraining.
  • Establish automated quality, safety, performance and compatibility checks.
  • Define and monitor SLIs/SLOs and incident response for ML pipelines.
  • Mentor engineers across ML, data, product, and platform teams.

Skills

Python
Go
Java
C++
Rust

Tools

Kubernetes
CI/CD
Infrastructure as Code

Job description

  • We are building AI-enabled product capabilities that improve through data, feedback, and real-world use
  • We need the production systems that make those capabilities dependable: repeatable delivery, measurable quality, controlled learning loops, and reliable operation at scale
  • We’re looking for a hands-on Staff Engineer who can move machine-learning models, agentic workflows, and self-learning approaches from promising prototypes into secure, observable, continuously deployable production systems
  • This role sits at the intersection of ML systems, platform engineering, and site reliability engineering
  • You will partner with ML, data, product, and infrastructure teams to create a paved path from experimentation to production—and take ownership of how those systems perform and evolve once deployed
  • Design and build the production path for the complete ML lifecycle: data and feature preparation, training, experiment tracking, evaluation, artifact and model management, serving, monitoring, feedback collection, and retraining
  • Build continuous-delivery workflows for models, prompts, agent workflows, data dependencies, and supporting services. Establish automated quality, safety, performance, and compatibility checks
  • Implement safe rollout patterns such as shadow traffic, canaries, progressive delivery, feature flags, versioned artifacts, automated rollback, and operational kill switches
  • Turn self-learning approaches into controlled production feedback loops. Build systems for collecting outcomes, validating feedback, maintaining lineage, triggering model refreshes, comparing candidates, and promoting changes under explicit guardrails
  • Define and operate SLIs, SLOs, alerts, and error budgets across infrastructure, data pipelines, inference services, model quality, and product behavior
  • Connect model analytics and product telemetry with traditional operational signals so teams can understand whether a problem originates in infrastructure, data, model behavior, or the surrounding product
  • Improve the scalability, availability, latency, and cost efficiency of distributed training, inference, and data-processing workloads. Own capacity planning and resource optimization, including GPU resources where applicable
  • Participate in production ownership across the service lifecycle: architecture reviews, deployment, on-call, incident response, blameless postmortems, and systemic remediation
  • Build self-service platforms and automation that reduce operational toil and shorten the time required for ML engineers and data scientists to reach production
  • Apply LLMs or agentic automation to evaluation, troubleshooting, and operational workflows where they produce reliable, measurable improvements
  • Establish practical standards for cloud infrastructure, Kubernetes, infrastructure as code, observability, security, and compliance
  • Provide technical leadership across ML, data, product, and platform teams, mentoring engineers and influencing architecture without relying on formal authority
Benefits
  • Generous family leave
  • Matched donations
  • Annual learning stipends
  • Flexible PTO
  • Competitive retirement plan
  • Paid volunteer time

Experience distinguishing service-health problems from data-quality or model-quality problemsPractical understanding of the ML lifecycle—including training, evaluation, model deployment, serving, monitoring, versioning, and retraining—and the ability to collaborate effectively with applied ML engineers or researchersStrong software-engineering skills in Python and at least one production systems language such as Go, Java, C++, or RustFamiliarity with SRE practices such as SLIs/SLOs, error budgets, sustainable on-call, incident management, and blameless postmortemsExcellent technical judgment and communication skills, especially when navigating ambiguity and coordinating across teams during production incidentsA track record of Staff-level technical ownership, typically gained through 7+ years of experience in software engineering, platform engineering, SRE, production engineering, or ML infrastructureExperience designing, operating, and troubleshooting distributed production systems, including failure analysis, capacity planning, and performance optimizationA strong automation and internal-customer mindset: you build platforms that are reliable, understandable, and pleasant for other engineers to useHands-on experience with cloud infrastructure, containers and Kubernetes, infrastructure as code, CI/CD, and modern observability

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Engineer, Machine Learning
Lead Engineer, Machine Learning

Salt Digital Recruitment • United States

On-site
USD 180,000 - 260,000
Senior ML Infrastructure Engineer
Senior ML Infrastructure Engineer

Harnham • New York (NY)

On-site
USD 150,000 - 200,000
MACHINE LEARNING ENGINEER (GENERAL)
MACHINE LEARNING ENGINEER (GENERAL)

MakerMaker • San Francisco (CA)

On-site
USD 180,000 - 260,000
MACHINE LEARNING ENGINEER (GENERAL)
MACHINE LEARNING ENGINEER (GENERAL)

MakerMaker.AI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 270,000
Staff Engineer, Machine Learning Systems & Reliability - Moveworks
Staff Engineer, Machine Learning Systems & Reliability - Moveworks

Servicenow • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

ExaCare AI • New York (NY)

On-site
USD 100,000 - 140,000
Flexible PTO
Medical, dental, and vision coverage
Company off-sites
Staff MLOps Engineer – ML Platform
Staff MLOps Engineer – ML Platform

BrightAI Corporation • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Senior Machine Learning Engineer Chicago, IL
Senior Machine Learning Engineer Chicago, IL

Attain • Chicago (IL), Northern (KY)

Hybrid
USD 170,000 - 240,000
Machine Learning Engineer
Machine Learning Engineer

Socket.dev • Shelton (CT)

On-site
USD 140,000 - 190,000
Senior Consultant, AI/ML Engineer
Senior Consultant, AI/ML Engineer

Hollstadt Consulting • Minnesota

On-site
USD 150,000 - 210,000