Staff Engineer, Machine Learning Infrastructure

Cacheflow

Palo Alto (CA)

On-site

USD 218,000 - 285,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Quince is seeking a Staff Engineer, Machine Learning Infrastructure to join our growing team in Palo Alto. You will design and operate production-grade ML systems at scale, from training pipelines to high-throughput inference serving, emphasizing extensibility, observability, and reliability.

You will mentor engineers, shape platform standards (CI/CD for ML, IaC, model versioning), and drive cost-aware performance optimizations.

Qualifications

  • 8+ years of industry experience, with at least 4+ years in ML Infrastructure, MLOps, or large-scale Data Platform engineering.
  • Proven track record designing and building MLOps platforms that support the full model lifecycle—from data ingestion and distributed training to real-time inference and model governance.
  • Deep expertise in cloud-native infrastructure (AWS), Kubernetes (EKS), Docker, and Infrastructure as Code tools (Terraform/Pulumi).
  • Hands-on mastery of ML frameworks such as PyTorch, TensorFlow, Kubeflow, or SageMaker, with strong opinions on building a cohesive, high-leverage developer experience.
  • Experience building Feature Stores and high-throughput data pipelines (Spark, Flink, Kafka).
  • CI/CD for ML, including model versioning, experiment tracking, and deployment strategies such as blue-green and canary rollouts.
  • Strong operational instincts with on-call discipline and reliability.
  • Startup mindset; capable of handling ambiguity and rapid pace.

Responsibilities

  • Architect the ML Infrastructure Foundation: end-to-end design of Quince’s ML platform, modular and scalable for long-term extensibility.
  • Build the 'Paved Road' for Production: develop the core developer experience for Data Scientists and AI Researchers to move from idea to production with minimal friction.
  • Drive Technical Excellence Across the Stack: set and uphold CI/CD for ML, IaC, model versioning, experiment tracking, and deployment strategies.
  • Own High-Impact System Design Decisions: lead evaluation and selection of core platform components with build-vs-buy balance.
  • Optimize Compute Performance & Cost: design GPU utilization optimizations, model batching, and cloud cost controls.
  • Ensure Production Scalability & Reliability: ML serving infrastructure that handles traffic surges with monitoring and automated recovery.
  • Mentor and Elevate the Engineering Team: code reviews and pairing to raise the technical bar.
  • Champion Operational Excellence: RCA for production failures and a rigorous on-call culture.

Skills

ML Infrastructure
MLOps platforms
CI/CD for ML
Cloud-native infra
DevOps collaboration

Tools

AWS
Kubernetes (EKS)
Docker
Terraform
Pulumi
Kubeflow
TensorFlow
PyTorch
SageMaker
Spark
Kafka

Job description

ABOUT QUINCE

Quince is a destination for builders, creators, innovators, and operators who want to come together and challenge the status quo. Our mission is simple: make really high quality essentials for really low prices, fairly and sustainably. We deliver on that mission through a unique manufacturer-to-consumer (M2C) model eliminating the layers of traditional retail that add cost and result in consumers paying more than they need to. We find, build relationships with, and work directly with the manufacturing partners behind some of the world’s finest products. From there, our teams design smart, efficient operational processes and build and deploy proprietary technology, AI, and analytics to help us scale fast.

What began with a small assortment of elevated basics has quickly grown into a cross-category brand spanning apparel, accessories, home goods, and more. Today, tens of millions of people across a growing number of countries come – and return – to Quince because they trust us to deliver.

OUR CULTURE

Quince is a culture built for builders by builders. Our way of working starts with a blank sheet of paper. We question conventional thinking, use technology and data to uncover new opportunities, and move quickly to turn ideas into reality. We aren’t interested in replicating how others do retail. We’re building a better way – at a speed and scale unlike anything that’s been done before.

  • We dream big and chase the hard problems others shy away from. Rejecting long-held assumptions is part of our company's DNA. Where conventional wisdom says you have to choose – soft or durable, speed or rigor, quality or price – we ask why that trade-off has to exist in the first place.
  • Our pace is fast and the bar is high because our customers expect a lot from us and we refuse to let them down. We believe the best results come from challenging ourselves, learning from one another, and building on each other's strengths.
  • Here, responsibility is not determined by role, tenure, or seniority. Every team member - no matter their level - has the opportunity to drive our business and shape our trajectory.

If you’re someone who likes to imagine new possibilities and build better systems rather than plug-in to outdated ones, Quince is the place for you.

THE ROLE
Staff Engineer, Machine Learning Infrastructure

We are seeking a Staff Engineer, Machine Learning Infrastructure to join our growing team.

The ideal candidate is a deeply technical ML infrastructure engineer who combines hands‑on mastery with system‑level thinking. You have built and operated production‑grade ML systems at scale — from distributed training pipelines and feature stores to high‑throughput inference serving — and you take pride in engineering platforms that other engineers love to use. You don’t just build for today’s requirements; you design for extensibility, observability, and resilience.

You are the kind of engineer who gravitates toward the hardest problems — whether that’s optimizing GPU utilization at the tail of the cost curve, designing a zero‑downtime model deployment system, or defining the architectural patterns that will define how Quince industrializes AI at scale. You operate with high autonomy, hold yourself to exceptional standards, and elevate the engineers around you through code reviews, technical mentorship, and by setting a bar for what great looks like.

Responsibilities
  • Architect the ML Infrastructure Foundation: Own the end‑to‑end technical design of Quince’s ML platform — including model training, serving, feature pipelines, and monitoring — ensuring it is modular, scalable, and built for long‑term extensibility.
  • Build the “Paved Road” for Production: Design and implement the core developer experience for Quince’s Data Scientists and AI Researchers, enabling them to move from “idea to production” with minimal friction and maximum reliability.
  • Drive Technical Excellence Across the Stack: Set and uphold engineering standards in CI/CD for ML, Infrastructure as Code (IaC), model versioning, experiment tracking, and deployment strategies (blue‑green, canary) — and build the tooling that makes those standards the path of least resistance.
  • Own High‑Impact System Design Decisions: Lead the technical evaluation and selection of core platform components — from inference runtimes and feature stores to orchestration frameworks — with a clear‑eyed view of build vs. buy tradeoffs.
  • Optimize Compute Performance & Cost: Design and implement GPU utilization optimizations, model batching strategies, and cloud cost controls to maximize performance per dollar across training and inference workloads.
  • Ensure Production Scalability & Reliability: Architect ML serving infrastructure that gracefully handles traffic surges, seasonal spikes, and model version transitions, with robust monitoring, alerting, and automated recovery.
  • Mentor and Elevate the Engineering Team: Provide deep technical mentorship to junior and mid‑level engineers through design reviews, code reviews, and pairing sessions — raising the collective technical bar without adding process overhead.
  • Champion Operational Excellence: Lead root‑cause analyses (RCAs) for production failures and drive systemic, permanent fixes over reactive patches. Model a culture of rigorous on‑call discipline and accountability.
Qualifications
  • 8+ years of industry experience, with at least 4+ years of focused, hands‑on work in ML Infrastructure, MLOps, or large‑scale Data Platform engineering.
  • Proven track record of designing and building MLOps platforms that support the full model lifecycle — from data ingestion and distributed training to real‑time inference and model governance.
  • Deep expertise in cloud‑native infrastructure (preferably AWS), Kubernetes (EKS), Docker, and Infrastructure as Code tools (Terraform/Pulumi).
  • Hands‑on mastery of ML frameworks such as PyTorch, TensorFlow, Kubeflow, or SageMaker, with strong opinions on building a cohesive, high‑leverage developer experience.
  • Expertise in building Feature Stores and high‑throughput data pipelines (Spark, Flink, Kafka), with a strong understanding of training/serving skew and data consistency.
  • Expert‑level knowledge of CI/CD for ML, including model versioning, experiment tracking, and deployment strategies such as blue‑green and canary rollouts.
  • Demonstrated ability to optimize GPU utilization, implement model batching, and systematically reduce cloud infrastructure costs.
  • Strong operational instincts, with a history of improving reliability through rigorous on‑call practices, proactive monitoring, and root‑cause analysis.
  • You understand the hustle of a startup and are good at handling ambiguity. You are a curious, quick learner who loves to experiment and thrives at a rapid pace.
Pay Range

$218,000-$285.000 (base) + bonus and stock

All posted ranges are reflective of base salary and may vary depending upon experience level and location. Bonus and equity may also be provided for eligible roles.

$218,000 — $285,000 USD

WHY QUINCE?

Joining Quince means being part of a mission‑driven team reshaping retail. You will work alongside talented colleagues, tackle meaningful challenges, and contribute to building a more sustainable, accessible future for customers and partners alike.

EQUAL OPPORTUNITY & HIRING INTEGRITY

Quince provides equal employment opportunities to all employees and applications for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran or military status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws.

Quince is committed to providing reasonable accommodations to qualified individuals with disabilities. If you need a reasonable accommodation to complete your application or to perform the essential functions of a role at Quince, please let us know by completing this accommodation form. We review all requests individually and will work with you to determine appropriate accommodations on a case‑by‑case basis.

Employment is contingent upon successful completion of a background check. Quince will conduct background checks in compliance with applicable federal, state, and local laws.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Engineer - ML Infra / MLOps
Staff Engineer - ML Infra / MLOps

Quince • Palo Alto (CA)

On-site
USD 218,000 - 285,000
Senior Infrastructure & Devops Engineer
Senior Infrastructure & Devops Engineer

Linuxconfig • Palo Alto (CA), Northern (KY)

Hybrid
USD 195,000 - 235,000
Infrastructure & DevOps Engineer III
Infrastructure & DevOps Engineer III

Cacheflow • Palo Alto (CA)

On-site
USD 195,000 - 235,000
Sr. Engineering Manager, MLOps Palo Alto, California, United States
Sr. Engineering Manager, MLOps Palo Alto, California, United States

Quince • Palo Alto (CA)

On-site
USD 270,000 - 300,000
Staff Software Engineer - Growth
Staff Software Engineer - Growth

Quince • Palo Alto (CA)

On-site
USD 230,000 - 275,000
Software Development Engineer III (Planning and Forecasting)
Software Development Engineer III (Planning and Forecasting)

Quince • Palo Alto (CA)

On-site
USD 205,000 - 215,000
Bonus
Equity
Senior/Staff Applied AI Researcher, Growth New Palo Alto, California, United States
Senior/Staff Applied AI Researcher, Growth New Palo Alto, California, United States

Quince • Palo Alto (CA), Northern (KY)

On-site
USD 213,000 - 285,000
Bonus
Equity
Senior Infrastructure & Devops Engineer
Senior Infrastructure & Devops Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 195,000 - 235,000
Senior/Staff Data Scientist, Storefront
Senior/Staff Data Scientist, Storefront

Quince • Palo Alto (CA)

On-site
USD 171,000 - 285,000
Bonus
Equity
Staff Software Engineer - Growth Palo Alto, California, United States
Staff Software Engineer - Growth Palo Alto, California, United States

Quince • Palo Alto (CA)

On-site
USD 230,000 - 275,000
Bonus
Equity