Software Engineer – ML Platform

Avride Inc.

Austin (TX)

On-site

USD 130,000 - 175,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Avride Inc. in Austin, Texas, is seeking an ML Platform Engineer to own core ML stack components such as workflow orchestration, distributed execution, and resource governance.

You will help ML teams run experiments and train models at scale on Kubernetes, with a focus on reliability and developer experience. You will build abstractions and services that enable scalable, cost-efficient, and fast training workloads, collaborating with ML teams to debug issues and drive platform improvements.

Qualifications

  • Proficiency in Python or Go; C++ is a plus.
  • Track record of designing and building scalable, maintainable systems and services.
  • Experience operating production services end-to-end: APIs, reliability practices, observability.
  • Deep knowledge of Kubernetes: scheduling, resource management, controllers, and pod lifecycle under pressure.
  • Solid Linux and systems debugging skills: performance investigation, networking, storage/IO.
  • Ability to troubleshoot complex production issues across logs, metrics, and traces and drive them to resolution.

Responsibilities

  • Build and scale ML compute platform on Kubernetes, using Argo Workflows for training, evaluation, and data processing orchestration.
  • Design and implement core platform capabilities, including a Ray-based internal SDK for distributed execution, and multi-tenant resource governance.
  • Improve end-to-end training throughput and platform efficiency by optimizing data access patterns, caching, and removing bottlenecks in storage, network, and resource contention.
  • Work directly with ML teams to debug complex workload issues, drive root-cause analysis, and turn recurring problems into platform-level fixes.
  • Evaluate, integrate and extend open-source tooling (Argo Workflows, Ray, Kubernetes ecosystem) to meet evolving platform needs.

Skills

Python
Go
C++

Tools

Kubernetes
Argo Workflows
Ray
MLflow

Job description

The ML Platform team at Avride builds the infrastructure that powers large-scale ML training and data processing for autonomous driving. We sit between Cloud Platform and ML engineers, turning low-level compute, storage, and networking primitives into an ML platform that teams actually use - scalable orchestration, distributed compute, and production-grade tooling for the full model lifecycle.

About the role

As an ML Platform Engineer at Avride, you'll own critical pieces of the ML stack: workflow orchestration, distributed execution, resource governance, performance.You will shape how ML teams across the company run experiments and train models at scale. You will build the abstractions and services that make training workloads reliable, cost-efficient, and fast, helping ML teams run at scale on Kubernetes with strong reliability and excellent developer experience.

What you will do
  • Build and scale our ML compute platform on Kubernetes, using Argo Workflows for training, evaluation, and data processing orchestration
  • Design and implement core platform capabilities, including a Ray-based internal SDK for distributed execution, and multi-tenant resource governance - scheduling, priorities, quotas, and policy enforcement across GPU, CPU, memory, and IO
  • Improve end-to-end training throughput and platform efficiency by optimizing data access patterns, caching, and removing bottlenecks in storage, network, and resource contention
  • Work directly with ML teams to debug complex workload issues, drive root-cause analysis, and turn recurring problems into platform-level fixes
  • Evaluate, integrate and extend open-source tooling (Argo Workflows, Ray, Kubernetes ecosystem) to meet evolving platform needs
What you will need
  • Strong proficiency in Python or Go; C++ is a plus
  • Track record of designing and building scalable, maintainable systems and services
  • Experience operating production services end-to-end: APIs, reliability practices, observability
  • Deep knowledge of Kubernetes: how scheduling, resource management, controllers, and pod lifecycle actually behave under pressure
  • Solid Linux and systems debugging skills: performance investigation, networking, storage/IO
  • Ability to troubleshoot complex production issues across logs, metrics, and traces and drive them to resolution
Nice to have
  • Experience with Argo Workflows, Ray, MLflow, or comparable distributed ML tooling
  • Hands-on experience building or operating large-scale ML training systems: GPU scheduling, distributed training, training data pipelines
  • Track record of optimizing resource usage and performance in distributed environments

Candidates are required to be authorized to work in the U.S. The employer is not offering relocation, sponsorship, and remote work options are not available.

Avride is an equal opportunity employer and committed to providing reasonable accommodations to qualified applicants and employees with disabilities to ensure they have equal access to employment opportunities. Avride complies with the Americans with Disabilities Act (ADA), if you need a reasonable accommodation to assist with the application or hiring process, or to perform the essential functions of a job, please email jobs@avride.ai .

Avride complies with the Americans with Disabilities Act (ADA). Can you perform the essential functions of the position for which you are applying, with or without reasonable accommodation? Select...

This is not a remote role. Are you able to work on site at our facility in North Austin, Monday-Friday?

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Platform Engineer: Scalable Kubernetes & Ray
ML Platform Engineer: Scalable Kubernetes & Ray

Avride Inc. • Austin (TX)

On-site
USD 130,000 - 175,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Avride • Austin (TX)

On-site
USD 100,000 - 130,000
Machine Learning Engineer II - Autonomous Driving Training Infrastructure
Machine Learning Engineer II - Autonomous Driving Training Infrastructure

Voiceflow • United States

On-site
USD 160,000 - 210,000
Comprehensive healthcare suite
Rich retirement benefits
Flexible vacation policy
+1
Machine Learning Engineer II Software
Machine Learning Engineer II Software

Front Door Defense • United States

On-site
USD 160,000 - 210,000
Comprehensive healthcare suite
Health Savings Accounts
Generous paid parental leave
+2
Software Engineer – Simulation Backend
Software Engineer – Simulation Backend

Avride • Austin (TX)

On-site
USD 100,000 - 130,000
Machine Learning Engineer II - Autonomous Driving Training Infrastructure
Machine Learning Engineer II - Autonomous Driving Training Infrastructure

Maymobility • United States

On-site
USD 160,000 - 210,000
Comprehensive healthcare suite
Rich retirement benefits
Flexible vacation policy
+1
Senior ML Inference Engineer – Platform
Senior ML Inference Engineer – Platform

Jobtailor • California (MO)

On-site
USD 140,000 - 190,000
Senior ML Platform Engineer (Autonomous Driving)
Senior ML Platform Engineer (Autonomous Driving)

42dot Inc. • Sunnyvale (CA)

On-site
USD 133,000 - 254,000
Senior ML Platform Engineer (Autonomous Driving)
Senior ML Platform Engineer (Autonomous Driving)

42dot • San Francisco (CA)

On-site
USD 133,000 - 254,000
Forward Deployed AI Engineer
Forward Deployed AI Engineer

Avathon • Pleasanton (CA)

On-site
USD 130,000 - 180,000
Equity in the form of stock options
Comprehensive health, dental, and vision coverage
401(k) participation
+2