Senior ML Platform Engineer: Distributed GPU Training Infra

Menlo Ventures

Seattle (WA)

On-site

USD 189,000 - 245,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Xaira Therapeutics is seeking a Senior Software Engineer to join our Platform team to design, build, and deploy the AI infrastructure that powers our world-class research. You’ll work with AI Scientists and engineers to enable thousands of GPUs for training and inference of biology foundation models in Seattle.

This role covers MLOps for GPU clusters and backend API work, with opportunities to dive into storage, telemetry, and experiment tracking.

Qualifications

  • Degree in Computer Science, ML, computational biology or related field.
  • 5+ years of industry experience building and deploying ML systems.
  • Experience leading technical projects and cross-functional execution.
  • Strong programming skills in Python.

Responsibilities

  • Develop and improve our model training system, dispatching distributed training jobs to clusters across multiple clouds.
  • Deploy storage subsystems that improve dataset management and throughput for training datasets.
  • Build evaluation infrastructure that enables easy execution and tracking.
  • Build base tooling for integrating model training with telemetry, experiment tracking, and checkpointing.
  • You’ll get to go in the lab; this can be a perk for the right candidate.

Skills

Python
Machine Learning
Leadership
Cross-functional collaboration
Problem-solving

Education

Bachelor's degree in CS/ML/Computational Biology or related field

Tools

Terraform
Ansible
Torch
Jax

Job description

Xaira Therapeutics is seeking a Senior Software Engineer to join our Platform team to design, build, and deploy the AI infrastructure that powers our world-class research. You’ll work with AI Scientists and engineers to enable thousands of GPUs for training and inference of biology foundation models in Seattle.

This role covers MLOps for GPU clusters and backend API work, with opportunities to dive into storage, telemetry, and experiment tracking.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, ML Platform
Senior Software Engineer, ML Platform

Menlo Ventures • Seattle (WA)

On-site
USD 189,000 - 245,000
Senior ML Infra Engineer for Distributed GPU Training
Senior ML Infra Engineer for Distributed GPU Training

Genesis Molecular AI • City of Utica (NY)

On-site
USD 150,000 - 190,000
Competitive compensation with salary +
Staff ML Infra Engineer: Distributed Training & Inference
Staff ML Infra Engineer: Distributed Training & Inference

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior ML Platform Engineer – Distributed GPU & Game AI
Senior ML Platform Engineer – Distributed GPU & Game AI

WorkGenius Group • Santa Monica (CA)

On-site
USD 124,000 - 207,000
Senior AI Infra Engineer for Large-Scale ML Platforms
Senior AI Infra Engineer for Large-Scale ML Platforms

NVIDIA • Seattle (WA)

On-site
USD 224,000 - 357,000
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Principal AI Inference Platform Engineer - GPU/K8s
Principal AI Inference Platform Engineer - GPU/K8s

Lila Sciences • Cambridge (ME)

On-site
USD 192,000 - 272,000
Medical, dental, and vision
Life and disability insurance
Flexible time off
+4
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2