ML Platform Engineer: Scale & Debug GPU Server Fleet

Annapurna Labs (U.S.) Inc.

Austin (TX)

On-site

USD 120,000 - 180,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
RSUs

Job summary

Annapurna Labs (U.S.) Inc. is seeking a Platform Development Engineer to own a fleet of ML server platforms, ensuring health, sellability, and strong customer experience. You will lead end-to-end testing, drive automation, and coordinate with hardware and software engineering teams to deliver reliable, scalable solutions.

The role emphasizes data-driven diagnostics, building robust tooling, and collaborating across teams to resolve issues that span hardware and software boundaries.

Qualifications

  • 2+ years of non-internship professional software development experience.
  • 1+ years of designing or architecting scalable systems.
  • 1+ years of administrative experience in networking, storage, operating systems and hands-on systems engineering experience.
  • Knowledge of systems engineering fundamentals (networking, storage, operating systems).
  • Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby.
  • Experience with Linux/Unix.
  • Experience debugging and systems analysis to identify and quickly resolve or mitigate issues.
  • Bachelor's degree in Computer Science, Computer Engineering, or Electrical Engineering.

Responsibilities

  • Member of a team responsible for system remediation, operational excellence, and customer experience on bleeding edge ML products.
  • Utilize data to root cause hardware failures and identify live trends on the most complex systems in AWS.
  • Implement and improve system level testing across the product lifecycle.
  • Develop software which can be maintained, improved upon, documented, tested, and reused.
  • Dive deep on issues at the intersection of hardware and software.

Skills

C++
Python
Java
Linux
Systems analysis

Education

Bachelor's degree in CS/CE/EE
Master's degree (preferred)

Tools

PowerShell
Ruby

Job description

Annapurna Labs (U.S.) Inc. is seeking a Platform Development Engineer to own a fleet of ML server platforms, ensuring health, sellability, and strong customer experience. You will lead end-to-end testing, drive automation, and coordinate with hardware and software engineering teams to deliver reliable, scalable solutions.

The role emphasizes data-driven diagnostics, building robust tooling, and collaborating across teams to resolve issues that span hardware and software boundaries.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Hardware Systems Engineer: Fleet Automation
ML Hardware Systems Engineer: Fleet Automation

JobCubby • Austin (TX), Northern (KY)

Hybrid
USD 136,000 - 184,000
Health insurance
401(k) matching
RSUs
+1
ML Hardware Platform Engineer
ML Hardware Platform Engineer

Socket.dev • Austin (TX)

On-site
USD 136,000 - 184,000
Machine Learning Hardware Platform Engineer
Machine Learning Hardware Platform Engineer

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 136,000 - 184,000
Health insurance
RSUs / restricted stock units
401(k) match
+2
ML Hardware Systems Engineer – Fleet Automation & Debugging
ML Hardware Systems Engineer – Fleet Automation & Debugging

Amazon • Austin (TX)

On-site
USD 136,000 - 184,000
Health insurance
401(k) matching
Paid time off
+1
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
ML Platform Engineer — Scale GPU-Driven Research Infra
ML Platform Engineer — Scale GPU-Driven Research Infra

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations
AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 136,000 - 184,000
Health insurance
RSUs / restricted stock units
401(k) match
+2
ML Platform Engineer: Build GPU-Scale Infra & Research
ML Platform Engineer: Build GPU-Scale Infra & Research

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
Staff ML Platform Engineer: Scale GPUs & Production
Staff ML Platform Engineer: Scale GPUs & Production

JobCubby • San Jose (CA), Northern (KY)

Hybrid
USD 212,000 - 307,000