Remote ML Performance Engineer: Throughput & Tuning Expert

Bright Vision Technologies

Plymouth (MN)

Remote

USD 100,000 - 150,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking an AI Performance Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems.

The role spans the full stack from low-level kernel optimization to distributed system tuning, requiring deep understanding of GPU architecture, model parallelism, memory management, and compiler-level optimization.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or related field.
  • Six or more years of experience in performance engineering, ML systems, or HPC.
  • Strong proficiency in Python and C++.
  • Hands-on experience optimizing deep learning workloads on modern GPUs.
  • Deep understanding of distributed training and inference techniques.
  • Experience with profiling tools across CPU, GPU, and distributed systems.

Responsibilities

  • Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost.
  • Identify and eliminate bottlenecks across data loading, model compute, communication, and memory.
  • Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference.
  • Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
  • Tune attention implementations using FlashAttention, paged attention, and related techniques.
  • Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving.
  • Drive compiler-level optimizations using Triton, XLA, TorchInductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains.
  • Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training.
  • Build and maintain rigorous benchmark suites and regression frameworks across workloads.
  • Collaborate with ML and platform engineering teams to embed best practices in standard pipelines.
  • Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies.
  • Evaluate new hardware and software offerings, and advise on adoption.
  • Document performance tuning playbooks and share findings broadly across engineering teams.
  • Stay current with AI systems research and translate advances into production improvements.

Skills

Python
C++
GPU optimization
Distributed training
Profiling tools
Memory systems
Debugging/Analytics

Education

Bachelor’s or Master’s degree in CS/CE or related field

Tools

Triton
XLA
TorchInductor
TVM
CUTLASS
TensorRT-LLM/DeepSpeed

Job description

Bright Vision Technologies is seeking an AI Performance Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems.

The role spans the full stack from low-level kernel optimization to distributed system tuning, requiring deep understanding of GPU architecture, model parallelism, memory management, and compiler-level optimization.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Performance Engineer - Max Throughput
Remote AI Performance Engineer - Max Throughput

Bazeta • Northern (KY)

Hybrid
USD 90,000 - 110,000
Remote Model Optimization Engineer — AI Performance
Remote Model Optimization Engineer — AI Performance

Bright Vision Technologies • United States

Remote
USD 150,000 - 175,000
Remote AI Performance Optimization Engineer
Remote AI Performance Optimization Engineer

NEPSE Trading • Northern (KY)

Hybrid
USD 90,000 - 110,000
Remote ML Systems Engineer: High-Performance Inference & Scale
Remote ML Systems Engineer: High-Performance Inference & Scale

Bright Vision Technologies • United States

Remote
USD 145,000 - 165,000
Remote AI Performance Engineer - Scale ML Inference
Remote AI Performance Engineer - Scale ML Inference

United States Digital Space LLC • United States

Remote
USD 75,000 - 100,000
Remote ML Infra Architect: Scalable GPU Training
Remote ML Infra Architect: Scalable GPU Training

Bright Vision Technologies • Plymouth (MN)

Remote
USD 100,000 - 150,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
AI Performance Engineer: Scale ML Throughput & Efficiency
AI Performance Engineer: Scale ML Throughput & Efficiency

Decisive Point • Sunnyvale (CA)

Hybrid
USD 180,000 - 240,000
AI Performance Engineer: Scale ML Throughput & Training
AI Performance Engineer: Scale ML Throughput & Training

InvestedintheMission • Sunnyvale (CA)

On-site
USD 180,000 - 260,000