Principal Engineer, Model Development Platform

EngineersOfAI

Sunnyvale (CA)

On-site

USD 150,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

EngineersOfAI is seeking a seasoned technical leader to design and evolve the platform's architecture for reliability and scalability. You will unify the platform across disciplines and lead architectural reviews to propose balanced solutions for complex challenges.

The ideal candidate has over 10 years in creating large-scale systems, showcasing experience in various technologies including Kubernetes, Spark, and ML pipelines. This role is critical for improving performance and ensuring system observability.

Qualifications

  • 10+ years of experience designing and building large-scale distributed systems, ML/AI infrastructure, or developer platforms.
  • At least 3 years as a staff or principal-level engineer.
  • Proven ability to design systems spanning web platforms and ML pipelines.

Responsibilities

  • Design and evolve the platform's architecture for reliability and scalability.
  • Unify the platform across various disciplines.
  • Lead architectural reviews and propose balanced solutions.

Skills

Technical Leadership at Scale
Architectural Depth & Breadth
Reliability and performance
Hands-On Systems Design

Tools

Kubernetes
Spark
Ray
Airflow
MLflow

Job description

Responsibilities
  • System architecture & reliability - Design and evolve the platform's overall architecture for reliability, observability, and scalability. Set performance, latency, and availability targets, and drive the engineering standards to meet them.
  • Cross-domain technical leadership - Unify the platform across disciplines, from front-end UIs and distributed training to Spark data pipelines and optimization-based experiment scheduling, ensuring systems interoperate cleanly.
  • Hands-on problem solving - Dive into the hardest challenges across subteams, lead architectural reviews, and propose pragmatic solutions that balance innovation with operational simplicity.
  • Experimentation & scheduling systems - Build systems that optimize how models are tested in simulation and on-road, using techniques like linear programming and heuristic optimization to balance hardware, safety, and research priorities while improving throughput and turnaround.
  • Data & compute infrastructure - Architect pipelines that ingest, transform, and enrich petabytes of fleet sensor data, and drive efficient compute use across GPU, CPU, cloud, and edge for both prototyping and large-scale training.
  • Strategic collaboration - Partner with Product, Research, and Operations to align architecture with user needs and co-own the platform's long-term roadmap.
About You

Essential

  • Technical Leadership at Scale – 10+ years of experience designing and building large-scale distributed systems, ML/AI infrastructure, full stack web application, or developer platforms, including at least 3 years as a staff or principal-level engineer.
  • Architectural Depth & Breadth – Proven ability to design systems spanning web platforms, ML pipelines, and large-scale compute orchestration (e.g., Spark, Ray, Kubernetes, Airflow, MLflow).
  • Reliability and performance – Experience driving platform reliability improvements, defining SLAs/SLOs, and building self-healing and observable systems that operate at “four nines” availability or better.
  • Hands-On Systems Design
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Engineer - Model-Based Platform
Sr. Engineer - Model-Based Platform

Stellantis • Auburn Hills (MI)

On-site
USD 140,000 - 190,000
Engineer V - Software
Engineer V - Software

Worky • Town of Florida (NY)

On-site
USD 160,000 - 210,000
Sr. Engineer - Model-Based Platform
Sr. Engineer - Model-Based Platform

Stellantis NV • Auburn (AL)

On-site
USD 120,000 - 150,000
Principal ML Platform Engineer - Scalable, Reliable Systems
Principal ML Platform Engineer - Scalable, Reliable Systems

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Staff Platform Machine Learning Engineer – Engine
Staff Platform Machine Learning Engineer – Engine

Jobtailor • New York (NY)

On-site
USD 140,000 - 190,000
Platform Engineering Manager
Platform Engineering Manager

Jobtailor • California (MO)

On-site
USD 180,000 - 240,000
Principal Platform Architect
Principal Platform Architect

TetraScience • United States

On-site
USD 230,000 - 320,000
Equity
Unlimited PTO
Life Insurance
+1
Senior Consultant, AI/ML Engineer
Senior Consultant, AI/ML Engineer

Hollstadt Consulting • Minnesota

On-site
USD 140,000 - 200,000
AI Principal Engineer
AI Principal Engineer

Accellor • Mountain View (CA)

On-site
USD 180,000 - 280,000
AI Platform Architect
AI Platform Architect

Cerebras • Austin (TX)

On-site