Staff ML Infrastructure Architect

Cacheflow

Palo Alto (CA)

On-site

USD 218,000 - 285,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Quince is seeking a Staff Engineer, Machine Learning Infrastructure to join our growing team in Palo Alto. You will design and operate production-grade ML systems at scale, from training pipelines to high-throughput inference serving, emphasizing extensibility, observability, and reliability.

You will mentor engineers, shape platform standards (CI/CD for ML, IaC, model versioning), and drive cost-aware performance optimizations.

Qualifications

  • 8+ years of industry experience, with at least 4+ years in ML Infrastructure, MLOps, or large-scale Data Platform engineering.
  • Proven track record designing and building MLOps platforms that support the full model lifecycle—from data ingestion and distributed training to real-time inference and model governance.
  • Deep expertise in cloud-native infrastructure (AWS), Kubernetes (EKS), Docker, and Infrastructure as Code tools (Terraform/Pulumi).
  • Hands-on mastery of ML frameworks such as PyTorch, TensorFlow, Kubeflow, or SageMaker, with strong opinions on building a cohesive, high-leverage developer experience.
  • Experience building Feature Stores and high-throughput data pipelines (Spark, Flink, Kafka).
  • CI/CD for ML, including model versioning, experiment tracking, and deployment strategies such as blue-green and canary rollouts.
  • Strong operational instincts with on-call discipline and reliability.
  • Startup mindset; capable of handling ambiguity and rapid pace.

Responsibilities

  • Architect the ML Infrastructure Foundation: end-to-end design of Quince’s ML platform, modular and scalable for long-term extensibility.
  • Build the 'Paved Road' for Production: develop the core developer experience for Data Scientists and AI Researchers to move from idea to production with minimal friction.
  • Drive Technical Excellence Across the Stack: set and uphold CI/CD for ML, IaC, model versioning, experiment tracking, and deployment strategies.
  • Own High-Impact System Design Decisions: lead evaluation and selection of core platform components with build-vs-buy balance.
  • Optimize Compute Performance & Cost: design GPU utilization optimizations, model batching, and cloud cost controls.
  • Ensure Production Scalability & Reliability: ML serving infrastructure that handles traffic surges with monitoring and automated recovery.
  • Mentor and Elevate the Engineering Team: code reviews and pairing to raise the technical bar.
  • Champion Operational Excellence: RCA for production failures and a rigorous on-call culture.

Skills

ML Infrastructure
MLOps platforms
CI/CD for ML
Cloud-native infra
DevOps collaboration

Tools

AWS
Kubernetes (EKS)
Docker
Terraform
Pulumi
Kubeflow
TensorFlow
PyTorch
SageMaker
Spark
Kafka

Job description

Quince is seeking a Staff Engineer, Machine Learning Infrastructure to join our growing team in Palo Alto. You will design and operate production-grade ML systems at scale, from training pipelines to high-throughput inference serving, emphasizing extensibility, observability, and reliability.

You will mentor engineers, shape platform standards (CI/CD for ML, IaC, model versioning), and drive cost-aware performance optimizations.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff ML Infra / MLOps Engineer: Scale & Automate AI
Staff ML Infra / MLOps Engineer: Scale & Automate AI

Quince • Palo Alto (CA)

On-site
USD 218,000 - 285,000
Staff ML Systems Engineer — Production ML Infra Lead
Staff ML Systems Engineer — Production ML Infra Lead

Quilter • Los Angeles (CA)

On-site
USD 130,000 - 160,000
Competitive salary and equity benefits
Health, dental, and vision insurance
Unlimited paid time off
+2
Sr. Engineering Manager, MLOps
Sr. Engineering Manager, MLOps

Quince • Palo Alto (CA)

On-site
USD 270,000 - 300,000
ML Infra Engineer - Scalable Training & Inference (Equity)
ML Infra Engineer - Scalable Training & Inference (Equity)

Snapchat • Palo Alto (CA)

On-site
USD 209,000 - 313,000
Senior ML Platform Engineer — Scale Research ML Infra
Senior ML Platform Engineer — Scale Research ML Infra

techire ai • San Francisco (CA)

On-site
USD 270,000 - 330,000
Stock options
Staff ML Platform Engineer — Lead & Scale ML Infra
Staff ML Platform Engineer — Lead & Scale ML Infra

Stripe • San Francisco (CA)

On-site
USD 224,000 - 336,000
Equity
401(k) plan
Medical, dental, and vision benefits
+1
Staff Engineer - ML Infra / MLOps
Staff Engineer - ML Infra / MLOps

Quince • Palo Alto (CA)

On-site
USD 218,000 - 285,000
Staff ML Infra Platform Architect
Staff ML Infra Platform Architect

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Remote Staff ML Systems Engineer — Production ML Infra
Remote Staff ML Systems Engineer — Production ML Infra

Devconnectplatform • Northern (KY)

Hybrid
USD 180,000 - 200,000
Senior ML Engineer, Data Infrastructure & Pipelines
Senior ML Engineer, Data Infrastructure & Pipelines

Unity Enterprise • Mountain View (CA)

Hybrid
USD 200,000 - 261,000
Comprehensive health insurance
Life and disability insurance
Employee stock ownership
+3