Senior ML Platform Engineer for Distributed GPU Training

Xairatherapeutics

Seattle (WA)

On-site

USD 205,000 - 325,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Bonus

Job summary

Xaira Therapeutics seeks a Senior Software Engineer to design, build, and deploy the AI infrastructure powering our research teams. You will enable thousands of GPUs for training and inference of biological foundation models across multi-cloud clusters.

Ideal candidates have 5+ years of experience deploying ML systems, strong Python skills, and a track record of technical leadership. The role emphasizes backend tooling and scalable pipelines in a collaborative, interdisciplinary environment.

Qualifications

  • 5+ years building and deploying ML systems in production.
  • Experience leading technical projects and cross-functional execution.
  • Strong Python programming skills.
  • Experience with infrastructure/ops tools such as Terraform and Ansible.
  • Experience with deep learning frameworks such as Torch and Jax.

Responsibilities

  • Develop and improve our model training system, dispatching distributed training jobs to clusters across multiple clouds.
  • Deploy storage subsystems to improve dataset management and throughput for training datasets.
  • Build evaluation infrastructure for easy execution and tracking.
  • Build base tooling to integrate model training with telemetry, experiment tracking, and checkpointing.
  • Collaborate with AI Scientists and other engineers; knowledge of biology not required to start, but possible lab exposure is a perk.

Skills

Python programming
ML systems production
Project leadership
Slurm/Kubernetes approach
Problem solving
Collaboration

Education

Degree in Computer Science / ML / Computational Biology or related field

Tools

Terraform
Ansible
Torch
Jax

Job description

Xaira Therapeutics seeks a Senior Software Engineer to design, build, and deploy the AI infrastructure powering our research teams. You will enable thousands of GPUs for training and inference of biological foundation models across multi-cloud clusters.

Ideal candidates have 5+ years of experience deploying ML systems, strong Python skills, and a track record of technical leadership. The role emphasizes backend tooling and scalable pipelines in a collaborative, interdisciplinary environment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Platform Engineer: Distributed GPU Training Infra
Senior ML Platform Engineer: Distributed GPU Training Infra

Menlo Ventures • Seattle (WA)

On-site
USD 189,000 - 245,000
Senior ML Platform Engineer - Equity & Biotech Impact
Senior ML Platform Engineer - Equity & Biotech Impact

Xaira Therapeutics • Seattle (WA)

On-site
USD 205,000 - 275,000
Senior ML Platform Engineer for Infra in Drug Discovery
Senior ML Platform Engineer for Infra in Drug Discovery

Xaira • Seattle (WA)

On-site
USD 205,000 - 275,000
Senior Software Engineer, ML Platform
Senior Software Engineer, ML Platform

Xaira Therapeutics • Seattle (WA)

On-site
USD 205,000 - 275,000
Senior Software Engineer, ML Platform
Senior Software Engineer, ML Platform

Menlo Ventures • Seattle (WA)

On-site
USD 189,000 - 245,000
Senior Software Engineer, ML Platform
Senior Software Engineer, ML Platform

Xairatherapeutics • Seattle (WA)

On-site
USD 205,000 - 325,000
Equity
Bonus
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior ML Infra Engineer for Distributed GPU Training
Senior ML Infra Engineer for Distributed GPU Training

Genesis Molecular AI • City of Utica (NY)

On-site
USD 150,000 - 190,000
Competitive compensation with salary +
Senior ML Systems Engineer – Distributed Training
Senior ML Systems Engineer – Distributed Training

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1
Remote Senior Training Infrastructure Engineer—Multi-GPU AI
Remote Senior Training Infrastructure Engineer—Multi-GPU AI

Luma AI • San Francisco (CA)

Hybrid
USD 187,500 - 395,000