Research Engineer - ML Systems & Data Pipelines

constellation

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health, dental, vision insurance
Relocation assistance
Office in SF Mission District
Team dinners & off-sites
Parental leave

Job summary

Constellation in San Francisco is seeking a Research Engineer to sit between data and models, orchestrating and optimizing training on long-horizon multimodal sequences while building pipelines to turn a growing corpus into learnable material for models.

You will write research code, align data across streams, and ensure reproducibility as the team experiments and scales the workflow with engineering and research colleagues.

Qualifications

  • 3+ years building ML systems or research infrastructure, including distributed training runs on 100+ GPUs.
  • Expert-level Python and deep knowledge of PyTorch internals: DDP/FSDP, mixed precision, gradient accumulation, profiling tools.
  • Track record of building or maintaining research libraries others depend on.
  • Experience with data pipelines over large unstructured and multimodal datasets and familiarity with columnar/streaming formats.
  • Hands-on experience with experiment tracking and dataset versioning tooling.
  • Ability to read papers, reimplement components, and evaluate loss curves critically.
  • Strong debugging skills across shards, collate functions, and training runs.
  • You thrive in a high-bandwidth, collaborative environment and care about the impact on people.

Responsibilities

  • Orchestrate and optimize training: distributed config (DDP/FSDP), mixed precision, checkpoints and profiling.
  • Build high-throughput data loading over TB-scale multimodal data: sharding, prefetching, caching and format choices.
  • Wrangle complex data: align/synchronize multi-stream time series; manage dataset versioning.
  • Derive rich features from raw recordings and version datasets for reproducibility.
  • Write clean, tested research libraries for datasets, models, transforms and metrics.
  • Make runs traceable: code version, config, dataset version and environment.
  • Build evaluation harnesses that surface regressions on new checkpoints.
  • Turn research prototypes into repeatable pipelines while preserving exploratory flexibility.

Skills

ML systems
Python
PyTorch internals
Distributed training
Experiment tracking
Data pipelines
Research libraries

Tools

torch_geometric
torchaudio
torcheeg
torch_brain
neuralsets

Job description

Constellation in San Francisco is seeking a Research Engineer to sit between data and models, orchestrating and optimizing training on long-horizon multimodal sequences while building pipelines to turn a growing corpus into learnable material for models.

You will write research code, align data across streams, and ensure reproducibility as the team experiments and scales the workflow with engineering and research colleagues.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer, ML Systems & Multimodal Pipelines
Research Engineer, ML Systems & Multimodal Pipelines

Breakout Ventures • San Francisco (CA)

On-site
USD 180,000 - 250,000
Health insurance
Relocation assistance
Workspace stipend
+1
Senior ML Systems Engineer: Production Pipelines
Senior ML Systems Engineer: Production Pipelines

MakerMaker • San Francisco (CA)

On-site
USD 180,000 - 260,000
Staff ML Data Engineer — Frontier Models & Data Pipelines
Staff ML Data Engineer — Frontier Models & Data Pipelines

Icehouseventures • San Francisco (CA)

On-site
USD 227,000 - 313,000
Equity compensation
Bonus incentive
Hybrid on-site in San Francisco
Staff ML Data Engineer: Frontier Models Data Pipelines
Staff ML Data Engineer: Frontier Models Data Pipelines

Procore • San Francisco (CA)

Hybrid
USD 227,000 - 313,000
Equity compensation
Bonus incentive compensation
Hybrid work model (3 days onsite)
Senior ML Engineer - Data Pipelines for LLMs
Senior ML Engineer - Data Pipelines for LLMs

Cisco Systems, Inc. • San Jose (CA)

Hybrid
USD 203,000 - 259,000
Medical, dental, vision insurance
401(k) with Cisco matching
Paid parental leave
Senior ML Systems Engineer - Production Pipelines
Senior ML Systems Engineer - Production Pipelines

MakerMaker.AI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 270,000
ML Engineer: Research-to-Production Pipelines
ML Engineer: Research-to-Production Pipelines

Deepgram • United States

On-site
USD 180,000 - 240,000
Senior ML Infrastructure Engineer — Frontier RL & LLM Training
Senior ML Infrastructure Engineer — Frontier RL & LLM Training

Preference Model • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health, vision, dental benefits
401K match
Lunch provided onsite
+2
ML Research Engineer, Data San Francisco, CA · On-site
ML Research Engineer, Data San Francisco, CA · On-site

Weave Robotics, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Research Data Platform Engineer
Research Data Platform Engineer

Anthropic Limited • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 405,000