ML Platform Engineer: Scalable Kubernetes & Ray

Avride Inc.

Austin (TX)

On-site

USD 130,000 - 175,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Avride Inc. in Austin, Texas, is seeking an ML Platform Engineer to own core ML stack components such as workflow orchestration, distributed execution, and resource governance.

You will help ML teams run experiments and train models at scale on Kubernetes, with a focus on reliability and developer experience. You will build abstractions and services that enable scalable, cost-efficient, and fast training workloads, collaborating with ML teams to debug issues and drive platform improvements.

Qualifications

  • Proficiency in Python or Go; C++ is a plus.
  • Track record of designing and building scalable, maintainable systems and services.
  • Experience operating production services end-to-end: APIs, reliability practices, observability.
  • Deep knowledge of Kubernetes: scheduling, resource management, controllers, and pod lifecycle under pressure.
  • Solid Linux and systems debugging skills: performance investigation, networking, storage/IO.
  • Ability to troubleshoot complex production issues across logs, metrics, and traces and drive them to resolution.

Responsibilities

  • Build and scale ML compute platform on Kubernetes, using Argo Workflows for training, evaluation, and data processing orchestration.
  • Design and implement core platform capabilities, including a Ray-based internal SDK for distributed execution, and multi-tenant resource governance.
  • Improve end-to-end training throughput and platform efficiency by optimizing data access patterns, caching, and removing bottlenecks in storage, network, and resource contention.
  • Work directly with ML teams to debug complex workload issues, drive root-cause analysis, and turn recurring problems into platform-level fixes.
  • Evaluate, integrate and extend open-source tooling (Argo Workflows, Ray, Kubernetes ecosystem) to meet evolving platform needs.

Skills

Python
Go
C++

Tools

Kubernetes
Argo Workflows
Ray
MLflow

Job description

Avride Inc. in Austin, Texas, is seeking an ML Platform Engineer to own core ML stack components such as workflow orchestration, distributed execution, and resource governance.

You will help ML teams run experiments and train models at scale on Kubernetes, with a focus on reliability and developer experience. You will build abstractions and services that enable scalable, cost-efficient, and fast training workloads, collaborating with ML teams to debug issues and drive platform improvements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer – ML Platform
Software Engineer – ML Platform

Avride Inc. • Austin (TX)

On-site
USD 130,000 - 175,000
Lead ML Platform Engineer: Training & Inference at Scale
Lead ML Platform Engineer: Training & Inference at Scale

Paramount • Burbank (CA)

On-site
USD 157,000 - 235,000
Benefits package
On-site & virtual events
Generous PTO
Principal ML Platform Engineer - Scalable, Reliable Systems
Principal ML Platform Engineer - Scalable, Reliable Systems

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
ML Platform Engineer & Research Innovator
ML Platform Engineer & Research Innovator

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
ML Platform Engineer: Scalable AI Infra & Tools
ML Platform Engineer: Scalable AI Infra & Tools

synthesia • United States

On-site
USD 100,000 - 140,000
Forward Deployed Engineer - AI/ML Platforms
Forward Deployed Engineer - AI/ML Platforms

anyscale • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior ML Platform Engineer - Cloud, Kubernetes & CI/CD
Senior ML Platform Engineer - Cloud, Kubernetes & CI/CD

Jobtailor • New York (NY)

On-site
USD 140,000 - 190,000
ML Platform Engineer: Scalable AI Infrastructure
ML Platform Engineer: Scalable AI Infrastructure

Bjak • Germany (OH)

On-site
USD 120,000 - 180,000
Software Engineer (SE / Sr SE), Data & ML Platform
Software Engineer (SE / Sr SE), Data & ML Platform

Plus 2 • Santa Clara (CA)

On-site
USD 150,000 - 200,000
ML Platform Engineer: Scalable AI Infra & Reliability
ML Platform Engineer: Scalable AI Infra & Reliability

Bjak • Town of Poland (NY)

On-site
USD 140,000 - 220,000