Staff Research Engineer, Scientific Computing and ML/Physics Infrastructure

Socket.dev

Cambridge (MA)

On-site

USD 224,000 - 294,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits package

Job summary

Lila Sciences is hiring a Research Engineer to bridge research and production for ML/physics infrastructure. You will turn research tools into scalable, reliable systems supporting model training, molecular simulation, and data workflows.

You’ll collaborate with scientists to optimize GPU use, distributed execution, and workflow reliability across clusters, containers, and cloud environments, delivering reusable services for researchers and AI agents.

Qualifications

  • Strong software engineering skills in Python and ML/scientific computing experience.
  • Experience building or operating distributed systems for research or data-intensive workloads.
  • Knowledge of GPU computing, profiling, and failure modes in large-scale workloads.
  • Experience with PyTorch, JAX, CUDA-aware workflows or related frameworks.

Responsibilities

  • Turn research tools into scalable, efficient systems.
  • Collaborate with scientists to deploy experiments and workflows.
  • Build and support ML/physics infrastructure for training and simulation.
  • Ensure workflows run reliably across clusters and environments.
  • Improve GPU utilization and distributed execution.

Skills

Python
ML frameworks
Distributed systems
GPU computing
Linux/Containers
Kubernetes/Ray
Debugging
Research-to-prod

Tools

PyTorch
JAX
CUDA
Docker
Kubernetes
Slurm
Ray
Argo

Job description

Your Impact at LILA

Lila Sciences is seeking a Research Engineer, Scientific Computing and ML/Physics Infrastructure to help turn promising research tools into robust, scalable systems. This role bridges research and production: you will work with scientists and ML researchers who can prototype useful tools, then help make those tools efficient, distributed, fault tolerant, and usable across Lila's compute environments.

The Molecular Intelligence team is building ML and physics-based infrastructure for drug discovery, including biophysics workflows, computational chemistry tools, cofolding models, low-data learning systems, simulation workflows, and agent-usable scientific pipelines. We need an engineer who can improve code quality, architecture, GPU efficiency, cluster portability, and operational reliability without slowing down research velocity.

What You'll Be Building
  • Take research tools, prototypes, and scientific workflows developed by scientists or academic-style researchers and make them scalable, efficient, and maintainable.
  • Collaborate directly with computational biophysics, computational chemistry, and machine learning scientists to turn research workflows into scalable agent-usable systems.
  • Build and support ML and physics infrastructure for model training, molecular simulation, data processing, and agent-executed scientific workflows.
  • Ensure workflows run reliably across multiple clusters and compute environments.
  • Improve GPU utilization, distributed execution, throughput, fault tolerance, and reproducibility for ML and scientific workloads.
  • Architect larger-scale systems around research code, including job orchestration, retry behavior, monitoring, artifact handling, and workflow traceability.
  • Optimize ML, physics, and pipeline code for performance and scalability.
  • Maintain development and execution environments across local, cloud, and GPU-based systems.
  • Package scientific tools into reusable services, workflows, or APIs that can be used by researchers, pipelines, and AI agents.
  • Partner with research, platform, and infrastructure teams to bridge exploratory scientific work with reliable engineering systems.
  • Document systems clearly and establish pragmatic engineering patterns for research teams.
What You'll Need to Succeed
  • Strong software engineering skills in Python and experience working with ML, scientific computing, or simulation codebases.
  • Experience building, scaling, or operating distributed systems for research, ML, physics, simulation, or data-intensive workloads.
  • Practical knowledge of GPU computing, performance profiling, distributed execution, and failure modes in large-scale workloads.
  • Experience with PyTorch, JAX, CUDA-aware workflows, or related ML/scientific computing frameworks.
  • Practical knowledge of Linux, Docker or containers, dependency management, and reproducible development environments.
  • Experience with orchestration, scheduling, or distributed execution systems such as Kubernetes, Slurm, Ray, Flyte, Argo, or similar tools.
  • Ability to take prototype-quality research code and improve its architecture, scalability, reliability, and maintainability.
  • Strong debugging skills across code, environments, infrastructure, data pipelines, and compute clusters.
  • Ability to work directly with researchers, understand ambiguous technical needs, and convert them into robust engineering solutions.
Bonus Points For
  • Familiarity with chemistry, computational biophysics, molecular simulation, computational chemistry, cheminformatics, or drug discovery workflows.
  • Experience with cloud GPU infrastructure, multi-cluster execution, or hybrid compute environments.
  • Experience building tools for LLM agents or automated research workflows.
  • Experience with workflow observability, checkpointing, retries, and fault-tolerant scientific workloads.
  • Experience with CI, testing, packaging, and release practices for research software.
  • Comfort supporting fast-moving research teams without over-engineering exploratory work.
Compensation

We offer competitive base compensation with bonus potential and generous early-stage equity. Your final offer will reflect your background, expertise, and expected impact.

U.S. Benefits.

Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office based employees; and a company subsidized lunch program.

International Benefits.

Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions; international salaries are set to local market.

Expected Base Salary Range

$224,000 – $294,000 USD

About LILA

Lila Sciences is building Scientific Superintelligence™ to solve humankind's greatest challenges. We believe science is the most inspiring frontier for AI. Rather than hard-coding expert knowledge into tools, LILA builds systems that can learn for themselves.

LILA combines advanced AI models with proprietary AI Science Factory™ instruments into an operating system for science that executes the entire scientific method autonomously, accelerating discovery at unprecedented speed, scale, and impact across medicine, materials, and energy. Learn more at www.lila.ai.

Guided by our core values of truth, trust, curiosity, grit, and velocity, we move with startup speed while tackling problems of historic importance. If this sounds like an environment you'd love to work in, even if you don't meet every qualification listed above, we encourage you to apply.

We’re All In

Lila Sciences is committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status.

Information you provide during your application process will be handled in accordance with our Candidate Privacy Policy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, App
Senior Software Engineer, App

jobr.pro • San Francisco (CA), Cambridge (MA)

On-site
USD 144,000 - 240,000
Medical, dental, and vision coverage
Flexible time off
Paid parental leave
+3
Scientist II/Senior Scientist, Computational Biophysics
Scientist II/Senior Scientist, Computational Biophysics

Lilasciences • Cambridge (MA), San Francisco (CA)

On-site
USD 141,000 - 218,000
Staff Software Engineer — AI-Driven Scientific Platform
Staff Software Engineer — AI-Driven Scientific Platform

Lila Sciences • San Francisco (CA), Cambridge (MA)

On-site
Senior Software Engineer, Scientific System of Record
Senior Software Engineer, Scientific System of Record

Lila Sciences • Cambridge (MA)

On-site
USD 144,000 - 240,000
Medical, dental, and vision coverage
Flexible time off
Educational assistance program
+2
Staff Software Engineer, Scientific System of Record
Staff Software Engineer, Scientific System of Record

Lila Sciences • San Francisco (CA), Cambridge (MA)

On-site
USD 144,000 - 288,000
Medical, dental, and vision coverage
Flexible time off
Paid parental leave
+3
Staff ML Engineer, Life Sciences AI
Staff ML Engineer, Life Sciences AI

Lilasciences • San Francisco (CA)

On-site
USD 162,000 - 201,000
Medical, dental, and vision coverage
Flexible time off
Educational assistance program
Sr Principal/Principal Software Engineer, App
Sr Principal/Principal Software Engineer, App

Lila Sciences • San Francisco (CA), Cambridge (MA)

On-site
USD 204,000 - 348,000
Medical, dental, and vision coverage
Flexible time off
Paid parental leave
+1
Scientist II/Senior Scientist, Computational Chemistry, Drug Discovery
Scientist II/Senior Scientist, Computational Chemistry, Drug Discovery

Lilasciences • Cambridge (MA), San Francisco (CA)

On-site
USD 141,000 - 218,000
Equity
Medical, dental, vision
Paid time off
+1
Staff Software Engineer, Lab Software
Staff Software Engineer, Lab Software

Lila Sciences • Cambridge (MA)

On-site
USD 192,000 - 256,000
Senior Principal AI Platform Engineer – Equity
Senior Principal AI Platform Engineer – Equity

Lila Sciences • San Francisco (CA), Cambridge (MA)

On-site