DevOps Engineer

LE0010 Stanford Health Care

Newark (CA)

On-site

USD 109,000 - 145,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Stanford Health Care in California seeks a Senior AI Platform & ML Ops Engineer to architect the layered infrastructure for autonomous systems. You will be the Master Engineer, blending Kubernetes, Terraform, LangGraph, and tool-calling logic into a robust AI factory that connects to healthcare APIs and patient data.

You will lead full-stack development of scalable microservices, implement ML model pipelines, monitor observability and safety guardrails, and mentor junior engineers while staying

Qualifications

  • Bachelor’s or higher degree in Computer Science, Engineering or related field.
  • Three or more years of directly related experience in MLOps/AI platforms.
  • Strong knowledge of cloud platforms (AWS, Azure, GCP) and IaC tools (Terraform, CloudFormation).

Responsibilities

  • Design, build, and maintain scalable AI/ML infrastructure and environments.
  • Develop and implement CI/CD pipelines for AI/ML models.
  • Collaborate with data scientists and software engineers on training, deployment, and inference pipelines.
  • Monitor AI/ML systems for high availability and performance.
  • Mentor junior team members and stay current with MLOps trends.

Skills

MLOps
Kubernetes
Terraform
Python
CI/CD
Cloud platforms
Containerization

Education

Bachelor’s degree in Computer Science or related field

Tools

Docker
GitLab CI/CD
CloudFormation

Job description

Senior AI Platform & ML Ops Engineer

We are seeking a high-caliber Senior AI Platform & ML Ops Engineer to architect the "layered" infrastructure required for autonomous, agentic systems within Stanford Healthcare. In this role, you will be the "Master Chef" of our AI ecosystem, seamlessly folding Expert-Level DevOps (Kubernetes, Terraform, DevOps orchestration) with Agentic Application Development (LangGraph, CrewAI, Tool-calling logic). You won't just manage servers; you will build the robust, full-stack "factory" where multi-agent frameworks interact with healthcare APIs, ensuring every autonomous action is governed by strict ML Ops observability (LangSmith, Arize) and safety guardrails. If you have the "crispy" coding skills to build RAG pipelines in Python and the "rich" architectural depth to deploy scalable microservices, extensive full stack software development expertise, we want you to lead the integration of reasoning‑based AI into the future of clinical and business workflow automations.

This is a Stanford Health Care job.

A Brief Overview

The MLOPs Engineer will play an integral role incorporating Artificial Intelligence (AI) within Stanford Health Care. The solutions will impact patient care, medical research, and operational services. This group is tasked to innovate, build, deploy and monitor production grade AI, machine learning (ML) and predictive algorithms into healthcare. The role will partner closely with lead researchers within the AI field and leaders across various clinical specialties and operations. This role will report to the Infrastructure group and have a dotted line relationship to the Data Science team. The role will be responsible for maintaining cloud‑based infrastructure as code repositories, maintaining infrastructure, deployment pipelines and designing the security landscape for the team and objects. The role will set the standards for the full SDLC of projects for the Data Science team.

Locations

Stanford Health Care

What you will do
  • Design, build and maintain scalable and robust infrastructure for AI/ML systems, including cloud-based environments, containerization and orchestration platforms.
  • Develop and implement CI/CD pipelines to automate the deployment, testing and monitoring of AI/ML models and applications.
  • Collaborate with data scientists, data engineers and software engineers to optimize model training, deployment and inference pipelines.
  • Monitor and troubleshoot AI/ML systems to ensure high availability, performance and reliability.
  • Maintain and monitor model training and inference pipelines across multi-cloud tenants especially around Large Language Models (LLMs).
  • Maintain Kubernetes pods, container registry and virtual machine image library and model registry.
  • Monitor infrastructure utilization and costs pertaining to model training, inference and GPU utilization.
  • Implement best practices for security, data privacy and compliance in AI/ML workflows and infrastructure.
  • Evaluate and integrate new tools, technologies and frameworks to improve the efficiency and effectiveness of our MLOps processes.
  • Mentor and provide technical guidance to junior members of the organization.
  • Stay up‑to‑date with the latest advancements and trends in MLOps, DevOps and cloud technologies and share them with the team.
Education Qualifications
  • Bachelor’s or higher degree in Computer Science, Engineering or a related field
Experience Qualifications
  • Three (3) or more years of directly related experience
Knowledge, Skills and Abilities
  • Proven experience as an MLOps Engineer.
  • Strong knowledge of cloud platforms such as AWS, Azure or Google Cloud and experience with infrastructure‑as‑code tools like Terraform or CloudFormation.
  • Proficiency in containerization technologies such as Docker and container orchestration platforms like Kubernetes.
  • Experience with CI/CD tools such as GitLab CI/CD, Github Actions or CiricleCI.
  • Solid programming skills in languages such as Python, Rust or Go and experience in scripting and automation.
  • Familiarity with machine learning frameworks and libraries such as PyTorch, Tensorflow and scikit‑learn.
  • Deep understanding of DevOps principles, agile methodologies and software development lifecycle.
  • Strong problem‑solving and trouble shooting skills, with the ability to analyze and resolve complex technical issues.
  • Excellent communication and collaboration skills with the ability to work effectively in cross‑functional teams.
Physical Demands and Work Conditions

Blood Borne Pathogens Category III - Tasks that involve NO exposure to blood, body fluids or tissues, and Category I tasks that are not a condition of employment. These principles apply to ALL employees: SHC Commitment to Providing an Exceptional Patient & Family Experience. Stanford Health Care sets a high standard for delivering value and an exceptional experience for our patients and families. Candidates for employment and existing employees must adopt and execute C-I-CARE standards for all of patients, families and towards each other. C‑I‑CARE is the foundation of Stanford’s patient‑experience and represents a framework for patient‑centered interactions. Simply put, we do what it takes to enable and empower patients and families to focus on health, healing and recovery.

C-I-CARE Standards
  • Know Me: Anticipate my needs and status to deliver effective care.
  • Show Me the Way: Guide and prompt my actions to arrive at better outcomes and better health.
  • Coordinate for Me: Own the complexity of my care through coordination.
Equal Opportunity Employer

Stanford Health Care (SHC) strongly values diversity and is committed to equal opportunity and non-discrimination in all of its policies and practices, including the area of employment. Accordingly, SHC does not discriminate against any person on the basis of race, color, sex, sexual orientation or gender identity and/or expression, religion, age, national or ethnic origin, political beliefs, marital status, medical condition, genetic information, veteran status, or disability, or the perception of any of the above. People of all genders, members of all racial and ethnic groups, people with disabilities, and veterans are encouraged to apply. Qualified applicants with criminal convictions will be considered after an individualized assessment of the conviction and the job requirements.

Base Pay Scale

Generally starting at $79.21 - $104.97 per hour. The salary of the finalist selected for this role will be set based on a variety of factors, including but not limited to, internal equity, experience, education, specialty and training. This pay scale is not a promise of a particular wage.

At Stanford Health Care, we seek to provide patients with the very best in diagnosis and treatment, with outstanding quality, compassion and coordination. With an unmatched track record of scientific discovery, technological innovation and translational medicine, Stanford Medicine physicians are pioneering leading edge therapies today that will change the way health care is delivered tomorrow. As part of our spirit of discovery, we also leverage our deep relationships with luminary Silicon Valley companies to develop new ways to deliver preeminent patient care. Learn about our awards and significant events.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director Engineering - AI/ML
Director Engineering - AI/ML

LE0010 Stanford Health Care • Palo Alto (CA)

On-site
USD 154,000 - 205,000
Sr. Enterprise AI Platform Engineer
Sr. Enterprise AI Platform Engineer

LE0010 Stanford Health Care • Palo Alto (CA)

On-site
USD 140,000 - 185,000
Data Scientist - AI/ML
Data Scientist - AI/ML

LE0010 Stanford Health Care • Palo Alto (CA)

On-site
USD 226,891,000 - 300,868,000
Senior Manager, Data Engineering
Senior Manager, Data Engineering

LE0010 Stanford Health Care • Palo Alto (CA)

On-site
USD 239,702,000 - 318,226,000
Lead Data Engineer - AI/ML
Lead Data Engineer - AI/ML

LE0010 Stanford Health Care • Palo Alto (CA)

On-site
USD 202,000 - 268,000
Senior Enterprise Info Mgmt Platform Admin
Senior Enterprise Info Mgmt Platform Admin

LE0010 Stanford Health Care • Palo Alto (CA)

On-site
USD 138,000 - 183,000
Sr. Enterprise AI Platform Engineer
Sr. Enterprise AI Platform Engineer

Stanford Health Care • California (MO)

Hybrid
USD 190,109,000 - 252,101,000
Manager, Enterprise Information Mgmt Platform Admin
Manager, Enterprise Information Mgmt Platform Admin

LE0010 Stanford Health Care • Palo Alto (CA)

On-site
USD 116,000 - 153,000
Assistant Patient Care Manager (RN), ICU Medicine - Days 0645
Assistant Patient Care Manager (RN), ICU Medicine - Days 0645

LE0010 Stanford Health Care • Palo Alto (CA)

On-site
USD 118,000 - 157,000
Clinical Systems Analyst - Data Integration & Software Development
Clinical Systems Analyst - Data Integration & Software Development

LE0010 Stanford Health Care • Palo Alto (CA)

On-site
USD 77,000 - 102,000