Staff Engineer, Research Infrastructure & ML Platform

Simile

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity grants
Health, dental, vision
Flexible time off

Job summary

Simile, the simulation company in San Francisco, seeks a Member of Technical Staff in Research Infrastructure to build and own the ML platform researchers rely on—from data schemas to production services serving millions of interdependent agent calls.

You will optimize training and inference pipelines, lead data architecture for scale, and own the GPU cluster. Ideal candidates have deep ML/system experience, Python, PyTorch/JAX, and a track record shipping production ML platforms.

Qualifications

  • Strong systems and ML proficiency with Python.
  • Experience with modern ML frameworks (PyTorch/JAX).
  • Production ML platforms or MLOps experience.
  • Experience architecting and debugging production distributed systems.
  • Own deployment pipelines from data to serving.
  • Research- and data-literate with ability to navigate ML frontier.
  • Self-directed and pragmatic with team collaboration.
  • Strong quantitative foundation in CS/Math/Stats.

Responsibilities

  • Build the ML platform our researchers rely on, from data schemas to production serving.
  • Make training and data pipelines fast, profiling GPU memory and FLOPs.
  • Make serving fast and cost-effective for population-scale simulations.
  • Lead data-architecture redesign to handle high-volume simulations.
  • Own the multi-node GPU cluster health, scaling, and autoscaling.
  • Engineer evaluation tooling with rigorous statistical approaches.
  • Push state of the art by translating academic work into production.

Skills

Python
PyTorch
JAX
Distributed systems
MLOps
GPU
Research & data literacy
Self-directed
Teamwork

Education

B.S. in Computer Science or related field

Tools

NVIDIA GPUs
CUDA
NCCL
InfiniBand
NVLink
Triton
TensorRT-LLM
vLLM

Job description

Simile, the simulation company in San Francisco, seeks a Member of Technical Staff in Research Infrastructure to build and own the ML platform researchers rely on—from data schemas to production services serving millions of interdependent agent calls.

You will optimize training and inference pipelines, lead data architecture for scale, and own the GPU cluster. Ideal candidates have deep ML/system experience, Python, PyTorch/JAX, and a track record shipping production ML platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Research Scientist — ML & Behavioral Simulation
Staff Research Scientist — ML & Behavioral Simulation

Simile • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health & Wellness
Equity
Flexible time off
Research Infrastructure - Member of Technical Staff
Research Infrastructure - Member of Technical Staff

Simile • San Francisco (CA)

On-site
USD 200,000 - 400,000
Equity grants
Health, dental, vision
Flexible time off
Research Engineer: High-Performance ML Infrastructure
Research Engineer: High-Performance ML Infrastructure

Fleet AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff ML Platform Engineer: Scalable Graph & Pipelines
Staff ML Platform Engineer: Scalable Graph & Pipelines

Tensec • Springfield (VA)

On-site
USD 230,000 - 322,000
Medical, dental, and vision insurance
401(k) program with employer match
Generous time off and parental leave
Staff ML Systems Engineer - Scalable AI Infrastructure
Staff ML Systems Engineer - Scalable AI Infrastructure

Meta • Menlo Park (CA)

On-site
USD 183,000 - 257,000
Sr. Platform Engineer, ML Infrastructure
Sr. Platform Engineer, ML Infrastructure

Insilico Search Partners • Cambridge (MA)

On-site
USD 140,000 - 210,000
Member of Technical Staff - ML Infra
Member of Technical Staff - ML Infra

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Machine Learning Research Engineer
Machine Learning Research Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Systems ML Engineer — AI Infrastructure
Systems ML Engineer — AI Infrastructure

Meta • Seattle (WA)

On-site
USD 184,000 - 257,000
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000