Systems ML Engineer: Performance, Cloud/Edge, Equity

Breakout Ventures

Cambridge (MA)

On-site

USD 180,000 - 270,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Lunch subsidy
Health insurance
401K with matching

Job summary

Transfyr, a Cambridge, MA startup building physical AI for science, seeks a Systems ML Engineer to optimize training and inference across cloud and edge.

You will profile bottlenecks, implement kernel fusion, design data pipelines, develop custom GPU kernels, and ensure secure, reliable production deployments.

Join a fast-moving team solving hard problems at the intersection of science, perception, and robotics, with strong collaboration and growth opportunities.

Qualifications

  • Demonstrated expertise in ML systems engineering and production deployment.
  • Experience profiling and optimizing large-scale models for training and inference.
  • Ability to design and implement high-performance GPU kernels (CUDA/Triton).
  • Strong cross-functional collaboration with research, perception, and operations.

Responsibilities

  • Profile and optimize data loading, gradient compute, and communication to reduce step time.
  • Improve distributed training pipelines with PyTorch Distributed.
  • Develop high-performance GPU kernels in CUDA or Triton.
  • Design and optimize data loading and inference pipelines for multimodal lab data.
  • Manage deployment across cloud and edge while controlling costs.
  • Debug and resolve performance bottlenecks and production issues.
  • Collaborate with research and perception teams to ensure reliable data flow.

Skills

ML systems engineering
Profiling & optimization
Cross-functional collaboration
Ambiguity tolerance

Tools

PyTorch Distributed
CUDA
Triton
Nsight
Cloud & edge infrastructure

Job description

Transfyr, a Cambridge, MA startup building physical AI for science, seeks a Systems ML Engineer to optimize training and inference across cloud and edge.

You will profile bottlenecks, implement kernel fusion, design data pipelines, develop custom GPU kernels, and ensure secure, reliable production deployments.

Join a fast-moving team solving hard problems at the intersection of science, perception, and robotics, with strong collaboration and growth opportunities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems ML Engineer — High-Performance AI (Cambridge)
Systems ML Engineer — High-Performance AI (Cambridge)

Transfyr Bio • Cambridge (MA)

On-site
USD 150,000 - 230,000
Low-cost health insurance
HSA
401K with matching
+1
Member of the Technical Staff - Systems ML Engineer
Member of the Technical Staff - Systems ML Engineer

Transfyr Bio • Cambridge (MA)

On-site
USD 150,000 - 230,000
Low-cost health insurance
HSA
401K with matching
+1
Member of the Technical Staff - Systems ML Engineer
Member of the Technical Staff - Systems ML Engineer

Breakout Ventures • Cambridge (MA)

On-site
USD 180,000 - 270,000
Equity
Lunch subsidy
Health insurance
+1
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Senior ML Performance Engineer — Ultra-Fast Inference + Equity
Senior ML Performance Engineer — Ultra-Fast Inference + Equity

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
AI Performance Engineer: Scale ML Workloads at Datacenter
AI Performance Engineer: Scale ML Workloads at Datacenter

Applied Intuition • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
AI/ML Engineer for Real-World Science Systems
AI/ML Engineer for Real-World Science Systems

Transfyr Bio • Cambridge (MA)

On-site
USD 150,000 - 190,000
Competitive compensation
Full benefits
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Remote-Ready ML Systems Engineer: High-Throughput Training
Remote-Ready ML Systems Engineer: High-Throughput Training

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 160,000