On-Site Senior MLOps Engineer: Real-Time LLM Infra

Nace.AI

Palo Alto (CA)

On-site

USD 210,000 - 280,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Nace AI in Palo Alto, CA is seeking a Senior MLOps Engineer to own the production-grade ML infrastructure that turns research into reliable systems. You will design and operate training orchestration, model registries, CI/CD for models, and automated evaluation pipelines across cloud and on-prem environments.

You'll lead LLM/SLM serving, scale low-latency inference with vLLM and related tooling, manage multi-GPU clusters, and implement observability and audit-ready controls to support enterprise

Qualifications

  • 5+ years of experience in MLOps, ML infrastructure, or platform engineering.
  • Proven experience deploying and scaling LLM/inference infrastructure in production, including model serving frameworks.
  • Strong proficiency with Kubernetes, Docker, and infrastructure-as-code (Terraform or similar).
  • Hands-on experience with GPU cluster management and distributed training/serving environments.
  • Proficient in Python with a track record of building substantial, maintainable systems.
  • Experience with ML pipelines and orchestration tools (Airflow, Kubeflow, Ray, MLflow, Weights & Biases).

Responsibilities

  • Design, build, and operate end-to-end ML infrastructure including training orchestration and evaluation pipelines.
  • Own LLM/SLM serving infrastructure to achieve low latency and high throughput.
  • Manage multi-GPU training and inference clusters across cloud and on-prem environments.
  • Implement observability: latency, throughput, drift, and alerting for models in production.
  • Harden stack for enterprise deployment: reproducibility, versioning, access controls, auditability.
  • Set MLOps best practices and tooling standards as a senior member of the team.

Skills

MLOps
ML infrastructure
Kubernetes
Python
Distributed training
GPU cluster management
CI/CD for models
Observability

Education

BS in CS or related field
MS in CS or related field

Tools

Docker
Terraform
Airflow
Kubeflow
Ray
MLflow
Weights & Biases
Spark

Job description

Nace AI in Palo Alto, CA is seeking a Senior MLOps Engineer to own the production-grade ML infrastructure that turns research into reliable systems. You will design and operate training orchestration, model registries, CI/CD for models, and automated evaluation pipelines across cloud and on-prem environments.

You'll lead LLM/SLM serving, scale low-latency inference with vLLM and related tooling, manage multi-GPU clusters, and implement observability and audit-ready controls to support enterprise

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior MLOps Engineer – Remote Production AI & LLMs
Senior MLOps Engineer – Remote Production AI & LLMs

Dynatron • United States

Remote
USD 150,000 - 230,000
Remote environment
Professional development
Ownership of production systems
+1
Senior AI Platform Engineer - LLMs & MLOps
Senior AI Platform Engineer - LLMs & MLOps

Namely • Mountain View (CA)

On-site
USD 120,000 - 160,000
Senior MLOps Architect for Production-Grade ML
Senior MLOps Architect for Production-Grade ML

S27a • Scottsdale (AZ)

On-site
USD 170,000 - 250,000
Unlimited PTO
Competitive parental leave
Annual bonus program
+1
MLOps Engineer — Scalable AI Infra & Deployment, Equity
MLOps Engineer — Scalable AI Infra & Deployment, Equity

Fundamental • United States

Remote
USD 180,000 - 260,000
Salary + equity
Health coverage for you and dependents
Parental leave for all
+2
Senior MLOps Engineer
Senior MLOps Engineer

Nace.AI • Palo Alto (CA)

On-site
USD 210,000 - 280,000
Senior LLM & ML Infrastructure Engineer (Remote)
Senior LLM & ML Infrastructure Engineer (Remote)

artificial intelligence and robotics laboratory (itu air lab) • Arlington (VA)

Hybrid
USD 140,000 - 210,000
Senior MLOps Engineer — Scalable ML Infra
Senior MLOps Engineer — Scalable ML Infra

Pharmadog • Boston (MA), Northern (KY)

Hybrid
USD 128,000 - 196,000
Senior MLOps Platform Engineer
Senior MLOps Platform Engineer

Roundel • Brooklyn Park (MN)

Hybrid
USD 98,000 - 176,000
Senior MLOps Engineer — Production ML Platforms
Senior MLOps Engineer — Production ML Platforms

Zeitview (formerly DroneBase) • San Francisco (CA)

On-site
USD 170,000 - 180,000
Stock options
Medical insurance (multiple plans)
Dental and vision insurance
+2
Executive Director, ML & MLOps Engineering
Executive Director, ML & MLOps Engineering

JPMorganChase • Palo Alto (CA)

On-site
USD 180,000 - 240,000