Staff ML Software Engineer (L6) — Platform Systems, AIMS Engineering

Netflix, Inc.

Los Gatos (CA)

On-site

USD 600,000 - 1,066,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive health plans
401(k) retirement plan with employer match
Flexible time off
Paid leave of absence programs

Job summary

Netflix, Inc. is seeking a Staff ML Software Engineer responsible for the modernization of the AIMS AI/ML stack. This role involves defining architecture, driving migrations, and ensuring systems perform optimally as complexity and traffic grow.

The ideal candidate will have significant experience with production AI/ML systems, strong Python skills, and a proven track record of improving system reliability. This position offers a competitive compensation package and a range of comprehensive benefits.

Qualifications

  • Significant experience designing and operating large-scale production AI/ML systems.
  • Hands-on experience migrating AI/ML systems across technology generations.
  • Proven track record of improving AI/ML system reliability and reducing costs.

Responsibilities

  • Define and drive the end-state architecture for the AIMS AI/ML stack.
  • Build migration tooling and shared abstractions for modernization.
  • Ensure AIMS AI/ML systems maintain performance as traffic grows.

Skills

Experience in AI/ML systems
Deep Python expertise
Software engineering fundamentals
Distributed systems background
Observability and monitoring systems

Job description

Role Overview

AI for Member Systems (AIMS) powers the recommendation, search, and personalized experience for over 300M members. The stack is mature but must evolve to support new model paradigms, tighter cost and efficiency expectations, and operational maturity at scale. This Staff ML Software Engineer will own the end‑to‑end modernization of the AIMS AI/ML stack, build observability and cost infrastructure, and define its long‑term architectural evolution.

Responsibilities
  • Define the end‑state architecture for the modernized AIMS AI/ML stack: organization, contracts, and migration path across training pipelines, AI frameworks, and data infrastructure.
  • Drive end‑to‑end migration of AIMS AI/ML systems onto a modern, Python‑native platform, coordinating across multiple AIMS teams and external platform partners, with dozens of production models in flight.
  • Build migration tooling and shared abstractions that reduce the cost of adoption for individual teams, so modernization does not require each team to solve the same problems independently.
  • Own scalability across training throughput and data pipelines, ensuring AIMS AI/ML systems stay performant as model complexity and member traffic grow.
  • Design and build observability systems that give AIMS ML practitioners deep visibility into model behavior, training pipeline health, serving latency, and data quality, making issues detectable and diagnosable before they become incidents.
  • Identify and drive cost optimization across AIMS training and serving infrastructure, developing frameworks and tooling that make compute efficiency a first‑class concern, not an afterthought.
  • Architect reliability improvements across the AIMS AI/ML stack, reducing toil, improving on‑call ergonomics, and setting the standard for operational excellence across the org.
  • Prototype and productionize GenAI‑powered tooling for anomaly detection, root cause analysis, and operational automation, applying LLM‑based systems to the problems of AI/ML reliability and cost at scale.
  • Surface systemic cost, reliability, and migration gaps by embedding with AI/ML teams across AIMS, and translating their friction into concrete engineering investments with org‑wide leverage.
  • Set technical standards for the modernized stack and raise the engineering bar across AIMS through design reviews, architectural guidance, and leading by example.
  • Own the long‑term architectural evolution of the AIMS AI/ML stack — continuously evaluating emerging infrastructure patterns, model paradigms, and platform capabilities, and translating them into a forward‑looking roadmap before they become urgent migrations.
What We’re Looking For
  • Significant experience designing, building, and operating large‑scale production AI/ML systems, including training pipelines and familiarity with model serving and online inference at high‑traffic scale.
  • Hands‑on experience migrating production AI/ML systems across technology generations; you have done this before and understand where it goes wrong.
  • Strong software engineering fundamentals with deep Python expertise and working proficiency in at least one JVM language (Scala or Java).
  • Proven track record of improving AI/ML system reliability, reducing infrastructure costs, and improving operational scalability.
  • Experience building observability and monitoring systems for AI/ML workloads; you understand what good visibility looks like across training, serving, and data pipelines.
  • Strong distributed systems background, including large‑scale batch processing and real‑time serving infrastructure.
  • Collaborate with partner teams to drive cross‑functional technical programs, setting direction, managing dependencies, and building consensus without formal authority.
  • High technical judgment: able to identify common patterns, build reusable frameworks, and make pragmatic calls on what to migrate, what to rewrite, and what to leave alone.
  • Comfortable operating without full information; you can scope a problem, define an approach, and course‑correct as you learn more.
Preferred Qualifications
  • Experience with compute and cost optimization for AI/ML workloads at scale, including capacity management and efficiency tooling.
  • Hands‑on experience building GenAI‑powered tooling for operational automation, root cause analysis, or anomaly detection in AI/ML systems.
  • Experience building developer tooling or platform abstractions that improve AI/ML practitioner velocity.
  • Applied experience in personalization domains such as recommendation systems, search, or discovery.
  • Familiarity with modern AI/ML infrastructure patterns including feature stores, model serving platforms, and experiment frameworks.
Benefits
  • Comprehensive health plans, mental health support, a 401(k) retirement plan with employer match, stock option program, disability programs, health savings and flexible spending accounts, family‑forming benefits, and life and serious injury benefits.
  • Paid leave of absence programs.
  • Full‑time hourly employees accrue 35 days annually for paid time off to be used for vacation, holidays, and sick paid time off.
  • Full‑time salaried employees are immediately entitled to flexible time off.

Compensation range: $600,000.00 – $1,066,000.00.

At Netflix, we are an equal‑opportunity employer and celebrate diversity, recognizing that diversity builds stronger teams. We approach diversity and inclusion seriously and thoughtfully. We do not discriminate on the basis of race, religion, color, ancestry, national origin, caste, sex, sexual orientation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager, Content & Business Products and Delivery F&S - GenAI Enterprise & Workforce Pro[...]
Senior Manager, Content & Business Products and Delivery F&S - GenAI Enterprise & Workforce Pro[...]

Netflix • Los Gatos (CA)

On-site
USD 350,000 - 560,000
Health plans
Mental health support
401(k) with employer match
+5
Senior Manager, Content & Business Products and Delivery F&S - GenAI Enterprise & Workforce Pro[...]
Senior Manager, Content & Business Products and Delivery F&S - GenAI Enterprise & Workforce Pro[...]

Socket.dev • Los Gatos (CA)

On-site
USD 350,000 - 560,000
Health plans
Mental health support
401(k) with employer match
+3
Manager, Data Engineering
Manager, Data Engineering

Netflix, Inc. • Los Angeles (CA)

On-site
USD 525,000 - 950,000
Health Plans
401(k) Retirement Plan
35 days annual paid leave
+1
Senior Manager, Content & Business Products and Delivery F&S - GenAI Enterprise & Workforce Pro[...]
Senior Manager, Content & Business Products and Delivery F&S - GenAI Enterprise & Workforce Pro[...]

Netflix, Inc. • Los Gatos (CA)

On-site
USD 350,000 - 560,000
Health Plans
Mental Health support
401(k) Retirement Plan with employer0
+6
Machine Learning Engineer, AI for Member Systems
Machine Learning Engineer, AI for Member Systems

Netflix • United States

On-site
USD 100,000 - 720,000
Health Plans
Mental Health Support
401(k) Retirement Plan
+7
Manager, Data Engineering
Manager, Data Engineering

Socket.dev • Los Angeles (CA)

On-site
USD 446,000 - 752,000
Research Scientist 5 - Content Promotion and Distribution
Research Scientist 5 - Content Promotion and Distribution

Netflix • Los Gatos (CA)

On-site
USD 466,000 - 750,000
Health plans
401(k) retirement plan with employer match
Flexible time off
+1
Member of Technical Staff, Agentic Systems - Games
Member of Technical Staff, Agentic Systems - Games

Netflix • Los Gatos (CA)

On-site
USD 890,000 - 1,690,000
Health plans
Stock options
401(k) matching
+2
Software Engineer 4 - DevOps, CI/CD
Software Engineer 4 - DevOps, CI/CD

Netflix • Los Gatos (CA)

On-site
USD 250,000 - 413,000
Machine Learning Scientist 4 - Content & Conversation Modeling
Machine Learning Scientist 4 - Content & Conversation Modeling

Netflix, Inc. • Seattle (WA)

On-site
USD 300,000 - 537,000
Health Plans
Mental Health support
401(k) Retirement Plan with employer match
+2