Remote ML Performance Engineer — Throughput Optimizer

Bright Vision Technologies

Renton (WA)

On-site

USD 100,000 - 150,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking an ML Performance Engineer to optimize large neural network workloads across training and inference. The role covers end-to-end performance from kernels to distributed systems, requiring deep GPU understanding and data-driven optimization.

You will work with product, design, and engineering teams to translate vague requirements into robust, scalable solutions. The ideal candidate has 6+ years in ML performance, strong Python/C++, and hands-on GPU

Qualifications

  • 6+ years in performance engineering, ML systems, or HPC.
  • Proficiency in Python and C++.
  • Experience optimizing DL workloads on modern GPUs.
  • Strong understanding of distributed training and inference.
  • Experience with CPU/GPU profiling and distributed systems.
  • Familiarity with model compression and accuracy implications.
  • Strong memory hierarchy and parallelism knowledge.
  • Excellent measurement, debugging, and analytical reasoning.
  • Strong communication and collaboration skills.

Responsibilities

  • Profile and optimize end-to-end AI training and inference for throughput, latency, and cost.
  • Identify bottlenecks across data loading, compute, communication, and memory.
  • Tune quantization, sparsity, pruning to reduce footprint and accelerate inference.
  • Optimize distributed training with tensor/pipeline parallelism and sharding.
  • Tune attention and memory optimizations for large models.
  • Drive compiler-level optimizations using Triton, XLA, TorchInductor, TVM.
  • Improve data pipelines, storage patterns, and throughput.
  • Build and maintain benchmarks and regression tests.
  • Collaborate with ML and platform teams to standardize pipelines.
  • Advise on hardware selection and scheduling for cost efficiency.
  • Document performance tuning findings and share with teams.
  • Stay current with AI systems research and production translate.

Skills

Python
C++
GPU optimization
Profiling tools
Distributed training
ML systems
HPC
Measurement & debugging
Communication & collaboration

Education

Bachelors or Masters in CS/CE or related field

Tools

Triton
XLA
TorchInductor
TVM
CUTLASS

Job description

Bright Vision Technologies is seeking an ML Performance Engineer to optimize large neural network workloads across training and inference. The role covers end-to-end performance from kernels to distributed systems, requiring deep GPU understanding and data-driven optimization.

You will work with product, design, and engineering teams to translate vague requirements into robust, scalable solutions. The ideal candidate has 6+ years in ML performance, strong Python/C++, and hands-on GPU

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Performance Engineer - GPU & ML Systems
Remote AI Performance Engineer - GPU & ML Systems

Bright Vision Technologies • Farmington Hills (MI)

On-site
USD 75,000 - 100,000
Remote work
Senior Remote Model Optimization Engineer
Senior Remote Model Optimization Engineer

Bright Vision Technologies • Round Rock (TX)

On-site
USD 150,000 - 175,000
Remote AI Optimization Engineer for High-Throughput ML
Remote AI Optimization Engineer for High-Throughput ML

Visa Hunt • United States

On-site
USD 85,000 - 115,000
Remote ML Systems Engineer Scalable Inference Platforms
Remote ML Systems Engineer Scalable Inference Platforms

Bright Vision Technologies • Round Rock (TX)

On-site
USD 145,000 - 165,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Remote AI Performance Engineer — Optimize LLM Pipelines
Remote AI Performance Engineer — Optimize LLM Pipelines

Bright Vision Technologies • Glastonbury (CT)

On-site
USD 90,000 - 150,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
AI Performance Engineer: Scale ML Throughput & Efficiency
AI Performance Engineer: Scale ML Throughput & Efficiency

Decisive Point • Sunnyvale (CA)

Hybrid
USD 180,000 - 240,000
Remote ML Infrastructure Engineer: Scale AI Inference
Remote ML Infrastructure Engineer: Scale AI Inference

JobCubby • Redwood City (CA)

Hybrid
USD 105,000 - 143,000
Remote AI Systems Engineer - Scale GPU ML Infra
Remote AI Systems Engineer - Scale GPU ML Infra

Bright-Vision-Technologies • United States

Remote
USD 90,000 - 100,000