Remote ML Performance Engineer – Scale AI Throughput

Bright Vision Technologies

Bothell (WA)

On-site

USD 100,000 - 150,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bright Vision Technologies is seeking an ML Performance Engineer for a 100% remote U.S. role that focuses on maximizing throughput, minimizing latency, and reducing costs across training and inference for large neural networks.

The position spans from low-level kernel work to distributed system tuning with a strong emphasis on measurable production impact. You will work with product and platform teams to drive end-to-end performance improvements, applying profiler tools and compiler backends

Qualifications

  • Bachelor’s or Master’s degree in CS, CE, or related field.
  • 6+ years in performance engineering, ML systems, or HPC.
  • Strong proficiency in Python and C++.
  • Experience optimizing deep learning workloads on modern GPUs.
  • Deep understanding of distributed training and inference.
  • Experience with profiling tools across CPU, GPU, and distributed systems.
  • Familiarity with model compression techniques and accuracy implications.
  • Strong memory hierarchy, communication primitives, and parallelism knowledge.
  • Excellent measurement, debugging, and analytical reasoning skills.
  • Strong communication and collaboration skills.

Responsibilities

  • Profile and optimize AI training and inference pipelines for throughput, latency, and cost.
  • Identify bottlenecks across data loading, model compute, communication, and memory.
  • Implement and tune quantization, sparsity, and pruning strategies.
  • Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
  • Tune attention implementations using FlashAttention and related techniques.
  • Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving.
  • Drive compiler-level optimizations using Triton, XLA, TorchInductor, or TVM.
  • Optimize data pipelines, sharding strategies, and storage access patterns.
  • Build and maintain rigorous benchmark suites and regression frameworks.
  • Collaborate with ML and platform teams to embed best practices.
  • Drive cost-efficiency through model architecture, hardware selection, and scheduling.
  • Evaluate new hardware and software offerings and advise on adoption.
  • Document performance tuning playbooks and share findings widely.
  • Stay current with AI systems research and translate advances into production improvements.

Skills

Python
C++
GPU optimization
Distributed training
Profiling tools
Model compression
Memory hierarchy

Education

Bachelor’s or Master’s in CS/CE

Tools

Triton
CUTLASS
TensorRT
DeepSpeed

Job description

Bright Vision Technologies is seeking an ML Performance Engineer for a 100% remote U.S. role that focuses on maximizing throughput, minimizing latency, and reducing costs across training and inference for large neural networks.

The position spans from low-level kernel work to distributed system tuning with a strong emphasis on measurable production impact. You will work with product and platform teams to drive end-to-end performance improvements, applying profiler tools and compiler backends

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Performance Engineer - GPU & ML Systems
Remote AI Performance Engineer - GPU & ML Systems

Bright Vision Technologies • Farmington Hills (MI)

On-site
USD 75,000 - 100,000
Remote work
Senior AI Performance Engineer — Remote
Senior AI Performance Engineer — Remote

Bright Vision Technologies • Nashua (NH)

On-site
USD 100,000 - 150,000
Remote AI Performance Engineer — Systems & Optimization
Remote AI Performance Engineer — Systems & Optimization

Bright Vision Technologies • Reston (VA)

Remote
USD 100,000 - 150,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Senior ML Infrastructure Engineer – Remote
Senior ML Infrastructure Engineer – Remote

Bright Vision Technologies • Redwood City (CA), San Mateo (CA)

On-site
USD 105,000 - 143,000
Remote ML Infrastructure Engineer: Scale AI Inference
Remote ML Infrastructure Engineer: Scale AI Inference

JobCubby • Redwood City (CA)

Hybrid
USD 105,000 - 143,000
Remote ML Systems Engineer Scalable Inference Platforms
Remote ML Systems Engineer Scalable Inference Platforms

Bright Vision Technologies • Round Rock (TX)

On-site
USD 145,000 - 165,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Remote ML Platform Engineer — Scalable Inference Systems
Remote ML Platform Engineer — Scalable Inference Systems

Bright Vision Technologies • Nashua (NH)

On-site
USD 100,000 - 160,000
Senior ML Model Serving Engineer (Remote)
Senior ML Model Serving Engineer (Remote)

NEPSE Trading • Northern (KY)

Hybrid
USD 74,000 - 98,000