Systems ML Engineer: Cloud & Edge Performance

S27a

Cambridge (MA)

On-site

USD 170,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Full benefits

Job summary

Transfyr in Cambridge, MA is seeking a Systems ML Engineer to optimize and deploy large-scale ML models across training and production. You will own performance, profile bottlenecks, and implement kernel-level improvements across cloud and edge environments.

You will collaborate with research and perception teams to push training efficiency, manage deployments with monitoring and rollback, and ensure robustness of multimodal data pipelines in real lab settings.

Qualifications

  • Proven ML systems engineering in production environments.
  • Experience with distributed training and model serving.
  • Cloud & edge infrastructure proficiency (AWS, Kubernetes).
  • Understanding of ML stack from hardware to deployment.
  • Strong cross-stack debugging: GPU, data loading, memory, and communication.
  • Experience with CUDA/kernel tuning for performance.

Responsibilities

  • Profile and optimize performance using Nsight, PyTorch Profiler to identify bottlenecks and implement kernel-level improvements.
  • Improve efficiency of distributed training pipelines with PyTorch Distributed.
  • Develop and maintain high-performance GPU kernels in CUDA/Triton for critical workloads.
  • Design and optimize data loading and inference pipelines for multimodal lab data.
  • Manage deployment across cloud and edge in active lab environments with attention to security and cost.
  • Debug and resolve production bottlenecks, resource issues, and failures across the stack.
  • Collaborate with research and perception teams to move data through pipelines reliably.
  • Implement monitoring, versioning, and rollback in deployments to protect updates.

Skills

ML systems engineering
Distributed training
Cloud & edge infra
Full ML stack
Cross-stack debugging
Algorithm optimization

Tools

CUDA
Triton
PyTorch
Kubernetes
Terraform
NCCL

Job description

Transfyr in Cambridge, MA is seeking a Systems ML Engineer to optimize and deploy large-scale ML models across training and production. You will own performance, profile bottlenecks, and implement kernel-level improvements across cloud and edge environments.

You will collaborate with research and perception teams to push training efficiency, manage deployments with monitoring and rollback, and ensure robustness of multimodal data pipelines in real lab settings.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems ML Engineer - Edge/Cloud Performance Architect
Systems ML Engineer - Edge/Cloud Performance Architect

Transfyr Bio • Cambridge (MA)

On-site
USD 160,000 - 230,000
Systems ML Engineer: Edge & Cloud Performance
Systems ML Engineer: Edge & Cloud Performance

Breakout Ventures • Cambridge (MA)

On-site
USD 170,000 - 260,000
ML Performance Engineer: Scale Training & Throughput
ML Performance Engineer: Scale Training & Throughput

Applied Intuition • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Senior ML Engineer - Production AI for Edge & Cloud
Senior ML Engineer - Production AI for Edge & Cloud

webAI • Washington

On-site
USD 150,000 - 210,000
Health benefits
401(k) match
Wellness stipend
+2
Senior ML Ops Engineer — Edge & Cloud Pipelines
Senior ML Ops Engineer — Edge & Cloud Pipelines

Compunnel, Inc. • Westbrook (ME)

On-site
USD 120,000 - 160,000
ML Systems Engineer: Optimizing Training & GPU Kernels
ML Systems Engineer: Optimizing Training & GPU Kernels

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
Remote ML Performance Engineer: Optimize Training Inference
Remote ML Performance Engineer: Optimize Training Inference

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Remote-Ready ML Systems Engineer: High-Throughput Training
Remote-Ready ML Systems Engineer: High-Throughput Training

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 160,000
Senior ML Systems Engineer – Production & LLMOps
Senior ML Systems Engineer – Production & LLMOps

Bain & Company • Chicago (IL)

Hybrid
USD 155,000 - 169,000
Medical insurance
Dental insurance
Vision plan
+2