SWE - Distributed

Achira

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Achira is building best-in-class foundation models for simulation in drug discovery. We seek a Software Engineer passionate about distributed computing to architect and build the infrastructure for ML data generation, model training, and fine-tuning across large-scale distributed systems.

You will ensure compute clusters are efficient, observable, cost-effective, and reliable, pushing the boundaries of ML development with emphasis on distributed systems, performance optimization, and cloud cost

Qualifications

  • Experience building or operating distributed compute systems.
  • Strong understanding of parallel computing and resource management.
  • Experience profiling and optimizing distributed workloads.
  • Hands-on with cloud platforms and cluster orchestration.
  • Familiarity with ML frameworks and MLOps principles.

Responsibilities

  • Architect, build, and optimize distributed compute infrastructure for ML data processing, training, and fine-tuning.
  • Improve cluster observability, scheduling, and resource utilization.

Skills

Distributed computing experience
Parallel computing
Cloud platforms (AWS/GCP/Azure)
Kubernetes or Slurm
ML frameworks (PyTorch/TensorFlow/JAX)

Education

Bachelor's degree in CS or related field

Tools

Ray
Dask
Celery
Kubernetes
Spark
Slurm

Job description

Why Achira
  • Join a world-class team of scientists, ML researchers, and engineers working together to reshape the future of drug discovery.

  • Work on cutting edge ML infrastructure at frontier scale: massive compute, massive data, and massive ambition.

  • Own impactful work end-to-end — from ideation to architecture to deployment on large-scale infrastructure.

  • Work in an environment that rewards rigor, speed, and a builder’s mindset.

About the Role

Achira is building best-in-class foundation models to solve the most challenging problems in simulation for drug discovery and beyond. Atomistic Foundation simulation models (FSMs) as world models of the physical microcosm span machine learning interaction potentials (MLIPs), neural network potentials (NNPs), and diverse classes of generative models.

We're seeking a Software Engineer passionate about distributed computing and its applications in machine learning. You'll have the opportunity to architect and build from the ground up the infrastructure for our ML data generation pipelines, model training, and fine-tuning workflows across large-scale distributed systems.

Your expertise will ensure our compute clusters are efficient, observable, cost-effective, and reliable—helping us push the boundaries of ML development. If you're passionate about distributed systems, performance optimization, and cloud cost efficiency, we'd love to hear from you.

You'll be empowered to eat, breathe, and think about the orchestration of complex workloads on multiple vendors scattered anywhere on the planet. Achira is a company which lives and breaths on computation, facile access at the lowest cost for our uniquely suited workloads is a mission critical endeavor.

What You’ll Do
  • Architect & Build: Design, implement, and optimize distributed compute infrastructure for ML data processing, training, and fine-tuning.

  • Optimize & Monitor: Improve cluster observability, scheduling, and resource utilization (CPU/GPU/TPU).

  • Compute Efficiency: Research and implement cost-efficient compute solutions (spot instances, auto-scaling, multi-cloud strategies).

  • Tooling: Develop tools for monitoring, debugging, and performance tuning of large-scale ML workloads.

  • Collaboration: Collaborate with ML engineers to accelerate training pipelines and reduce bottlenecks.

  • Innovation: Stay current with emerging technologies in distributed computing (e.g., Ray, Kubernetes, Spark, Slurm) and apply them strategically.

About You
  • You are excited about and have lots of experience in building or working with distributed computing frameworks (e.g., Ray, Dask, Celery)

  • You have a good grasp of parallel computing, job scheduling, and resource management.

  • You're comfortable identifying and resolving performance issues in distributed systems (profiling, bottlenecks, network overhead)

  • You've implemented solutions using cloud compute platforms (AWS, GCP, Azure) and cluster orchestration (Kubernetes, Slurm)

  • You are familiar with popular ML frameworks (PyTorch, TensorFlow, or JAX) and MLOps best practices such as model deployment and GPU performance monitoring


Eligibility

In compliance with United States federal law, all persons hired will be required to verify identity and eligibility to work in the United States and to provide required employment eligibility verification documentation upon hire.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Research Engineer (MLRE) - GPUs
Machine Learning Research Engineer (MLRE) - GPUs

Achira • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Research Engineer (MLRE) - Workflows/Systems
Machine Learning Research Engineer (MLRE) - Workflows/Systems

Achira • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Research Engineer (MLRE) - Research
Machine Learning Research Engineer (MLRE) - Research

Achira • San Francisco (CA)

On-site
USD 120,000 - 160,000
Distributed ML Systems Engineer
Distributed ML Systems Engineer

Achira • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Research Scientist (MLRS) - Representation Learning for Molecular AI
ML Research Scientist (MLRS) - Representation Learning for Molecular AI

Achira • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
Senior ML Systems Engineer, Frameworks & Tooling
Senior ML Systems Engineer, Frameworks & Tooling

Cohere • United States

Remote
USD 150,000 - 230,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
AI Engineer in Synthetic Data
AI Engineer in Synthetic Data

Logical Intelligence • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Research Scientist (MLRS) - Generative AI
ML Research Scientist (MLRS) - Generative AI

Achira • San Francisco (CA)

On-site
USD 110,000 - 150,000