Distributed ML Systems Engineer

Achira

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Achira is building best-in-class foundation models for simulation in drug discovery. We seek a Software Engineer passionate about distributed computing to architect and build the infrastructure for ML data generation, model training, and fine-tuning across large-scale distributed systems.

You will ensure compute clusters are efficient, observable, cost-effective, and reliable, pushing the boundaries of ML development with emphasis on distributed systems, performance optimization, and cloud cost

Qualifications

  • Experience building or operating distributed compute systems.
  • Strong understanding of parallel computing and resource management.
  • Experience profiling and optimizing distributed workloads.
  • Hands-on with cloud platforms and cluster orchestration.
  • Familiarity with ML frameworks and MLOps principles.

Responsibilities

  • Architect, build, and optimize distributed compute infrastructure for ML data processing, training, and fine-tuning.
  • Improve cluster observability, scheduling, and resource utilization.

Skills

Distributed computing experience
Parallel computing
Cloud platforms (AWS/GCP/Azure)
Kubernetes or Slurm
ML frameworks (PyTorch/TensorFlow/JAX)

Education

Bachelor's degree in CS or related field

Tools

Ray
Dask
Celery
Kubernetes
Spark
Slurm

Job description

Achira is building best-in-class foundation models for simulation in drug discovery. We seek a Software Engineer passionate about distributed computing to architect and build the infrastructure for ML data generation, model training, and fine-tuning across large-scale distributed systems.

You will ensure compute clusters are efficient, observable, cost-effective, and reliable, pushing the boundaries of ML development with emphasis on distributed systems, performance optimization, and cloud cost

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SWE - Distributed
SWE - Distributed

Achira • San Francisco (CA)

On-site
USD 150,000 - 230,000
Machine Learning Research Engineer (MLRE) - Workflows/Systems
Machine Learning Research Engineer (MLRE) - Workflows/Systems

Achira • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Systems Architect for Scalable Drug Discovery Workflows
ML Systems Architect for Scalable Drug Discovery Workflows

Achira • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Machine Learning Research Engineer (MLRE) - Research
Machine Learning Research Engineer (MLRE) - Research

Achira • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Research Engineer (MLRE) - GPUs
Machine Learning Research Engineer (MLRE) - GPUs

Achira • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior ML Systems Engineer, Distributed Infrastructure
Senior ML Systems Engineer, Distributed Infrastructure

Field AI • Seattle (WA)

On-site
USD 110,000 - 160,000
ML Research Scientist - Atomistic Modeling for Drugs
ML Research Scientist - Atomistic Modeling for Drugs

Achira • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
ML Research Scientist (MLRS) - Representation Learning for Molecular AI
ML Research Scientist (MLRS) - Representation Learning for Molecular AI

Achira • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior ML Systems Engineer – Distributed Training
Senior ML Systems Engineer – Distributed Training

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1
Staff ML Systems Engineer: Distributed AI for Robotics
Staff ML Systems Engineer: Distributed AI for Robotics

Lever, Inc. • Seattle (WA)

On-site
USD 195,000 - 230,000
Equity participation
Comprehensive benefits