Software Engineer, ML Serving Platform

DoorDash, Inc.

San Francisco, Sunnyvale, Seattle (CA, CA, WA)

On-site

USD 140,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

401(k) with employer matching
Paid parental leave
Wellness benefits
Medical, dental, and vision benefits

Job summary

DoorDash, Inc. in San Francisco is seeking an experienced software engineer to help build the next generation of our ML serving platform.

You will work on connecting request routing, feature retrieval, and model inference on CPU/GPU infrastructure, addressing latency and scale challenges. You will own projects from design to production, collaborating with modelers and platform teams, and contributing to self-serve tooling and open source integrations.

Qualifications

  • 2+ years of software engineering experience building production services.
  • Proficiency in one of Java, Kotlin, Go, C++, or Python.
  • Understanding distributed systems, concurrency, networking, timeouts, and performance trade-offs.
  • Degree in CS or related field, or equivalent practical experience.

Responsibilities

  • Build and maintain production services for the ML serving platform.
  • Improve latency, reliability, and resource efficiency.
  • Collaborate with internal teams to understand modeling requirements.
  • Own design, testing, rollout and on-call production support.

Skills

Java
Kotlin
Go
C++
Python

Education

Bachelor's degree in Computer Science or related field

Job description

About the Team

DoorDash’s ML Serving Platform delivers tens of millions of predictions per second, powering search, recommendations, advertising, delivery estimates, and logistics across DoorDash, Wolt, and Deliveroo. Our customers are modelers and engineering teams across our internal business verticals. We build self-serve infrastructure that empowers them to bring new models into production, adopt open source software and models, and expand what they can accomplish with machine learning at scale. We work at the intersection of a rapidly evolving open source ecosystem and demanding production workloads. With active customer demand and growing modeling ambitions, we’re advancing the infrastructure behind today’s predictions while building the capabilities that enable the next generation of ML innovation.

About the Role

You’ll help build the next generation of our ML serving platform, connecting request routing and online feature retrieval with model inference on CPU and GPU infrastructure. You’ll tackle challenging infrastructure problems involving latency, reliability, resource efficiency, and scale, and make those capabilities accessible through self-serve tools and workflows. Working alongside experienced platform engineers, you’ll own defined projects from technical design and implementation through testing, rollout, and production support. You’ll partner directly with internal teams to understand emerging modeling requirements, remove adoption barriers, and turn advances in open source technology into measurable production improvements. You’re excited about this opportunity because you will… Empower modelers through self-serve infrastructure. Build tools, APIs, and workflows that help internal teams deploy, configure, validate, and operate models independently, accelerating ML adoption across business verticals. Bring open source innovation into production. Evaluate and integrate evolving inference frameworks and enable open source models, translating promising capabilities into reliable, efficient services at scale. Solve demanding inference infrastructure problems. Improve the systems that route prediction requests, retrieve online features, and execute models on a platform serving tens of millions of predictions per second. Help evolve our disaggregated serving architecture. Build modular components that allow routing, feature retrieval, and model execution to evolve and scale independently as workloads and modeling requirements change. Improve Kubernetes deployment and autoscaling. Help workloads respond to changing traffic while meeting latency and availability requirements and using CPU and GPU resources efficiently. Make adoption and rollout easier. Build integrations, validation, and migration tooling that help teams adopt new serving capabilities and our unified global platform with confidence. Own performance and reliability in production. Use benchmarking, profiling, metrics, and tracing to identify bottlenecks; participate in on-call and improve automation and runbooks to make the platform easier to operate. We’re excited about you because… You have 2+ years of software engineering experience building and maintaining production services or infrastructure. You have strong computer science fundamentals and proficiency in a backend or systems programming language such as Java, Kotlin, Go, C++, or Python. You understand distributed systems fundamentals, including concurrency, networking, timeouts, failure handling, and performance trade-offs. You enjoy challenging infrastructure problems and can independently turn a defined problem into a technical design, tested implementation, and safe production rollout. You have experience debugging production systems and using operational data to improve reliability, performance, or cost. You care about the engineers using your platform and collaborate effectively with customers and partners to make complex infrastructure easier to adopt. You hold a degree in Computer Science or a related field, or have equivalent practical experience.

Nice to Have Experience

Nice to Have Experience with ML inference infrastructure, online feature retrieval, or other latency-sensitive distributed services. Experience operating containerized workloads on Kubernetes, including deployment, resource management, or autoscaling. Experience building self-serve developer platforms, APIs, or automation that helps other engineers move faster. Experience integrating open source infrastructure or models and validating their behavior under production workloads. Familiarity with CPU/GPU performance profiling, inference runtimes, or model deployment workflows.

Compensation

The successful candidate’s starting pay will fall within the pay range listed below and is determined based on job-related factors including, but not limited to, skills, experience, qualifications, work location, and market conditions. Base salary is localized according to an employee’s work location. Ranges are market-dependent and may be modified in the future. In addition to base salary, the compensation for this role includes opportunities for equity grants.

DoorDash cares about you and your overall well-being. That’s why we offer a comprehensive benefits package to all regular employees, which includes a 401(k) plan with employer matching, 16 weeks of paid parental leave, wellness benefits, commuter benefits match, paid time off and paid sick leave in compliance with applicable laws (e.g. Colorado Healthy Families and Workplaces Act). DoorDash also offers medical, dental, and vision benefits, 11 paid holidays, disability and basic life insurance, family-forming assistance, and a mental health program, among others. See below for paid time off details: For salaried roles: flexible paid time off/vacation, plus 80 hours of paid sick time per year. F

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Machine Learning Platform - Gen AI
Software Engineer, Machine Learning Platform - Gen AI

DoorDash, Inc. • San Francisco (CA), Sunnyvale (CA), Seattle (WA)

On-site
USD 180,000 - 260,000
401(k) plan with employer matching
Parental leave (16 weeks)
Wellness benefits
+2
Software Engineer, Machine Learning Platform - MDX
Software Engineer, Machine Learning Platform - MDX

Aimlroles • San Francisco (CA)

Hybrid
USD 131,000 - 192,000
Software Engineer, Machine Learning Platform - MDX
Software Engineer, Machine Learning Platform - MDX

DoorDash USA • San Francisco (CA)

Hybrid
USD 131,000 - 192,000
Software Engineer, Machine Learning Platform - Gen AI
Software Engineer, Machine Learning Platform - Gen AI

Doordashusa • San Francisco (CA), Sunnyvale (CA), Seattle (WA)

On-site
USD 131,000 - 192,000
401(k) plan with employer matching
Paid parental leave (16 weeks)
Wellness benefits
Senior Software Engineer, Machine Learning Infrastructure - Generative AI
Senior Software Engineer, Machine Learning Infrastructure - Generative AI

DoorDash • Sunnyvale (CA)

On-site
USD 137,100 - 201,600
401(k) plan with employer matching
16 weeks of paid parental leave
Medical, dental, and vision benefits
+1
Senior Software Engineer, Machine Learning Infrastructure - Generative AI
Senior Software Engineer, Machine Learning Infrastructure - Generative AI

DoorDash • Seattle (WA)

On-site
USD 137,100 - 201,600
401(k) plan with employer matching
16 weeks of paid parental leave
Comprehensive wellness benefits
Software Engineer, Machine Learning (All Levels / All Teams)
Software Engineer, Machine Learning (All Levels / All Teams)

DoorDash, Inc. • Sunnyvale (CA), San Francisco (CA), Seattle (WA), New York (NY)

On-site
USD 137,000 - 299,000
401(k) plan with employer matching
16 weeks of paid parental leave
Wellness benefits
+13
Software Engineer, ML Serving Platform
Software Engineer, ML Serving Platform

DoorDash USA • San Francisco (CA), Sunnyvale (CA), Seattle (WA)

On-site
USD 131,000 - 192,000
Equity grants
401(k) with employer matching
Paid parental leave (16 weeks)
+4
Machine Learning Intern (PhD) - Summer 2027
Machine Learning Intern (PhD) - Summer 2027

DoorDash, Inc. • San Francisco (CA), Sunnyvale (CA), New York (NY), Seattle (WA)

On-site
USD 60,000 - 90,000
401(k) with employer matching
Paid parental leave
Medical, dental, and vision benefits
+3
Software Engineer, Full Stack - Developer Insights
Software Engineer, Full Stack - Developer Insights

DoorDash, Inc. • San Francisco (CA), Los Angeles (CA), New York (NY), Seattle (WA)

On-site
USD 131,000 - 192,000
401(k) with employer matching
Parental leave (16 weeks)
Health, dental & vision benefits
+3