ML Systems Engineer: GPU-Accelerated Training & Inference

voltai-com

Edison (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
Professional Growth
Visa Sponsorship

Job summary

Voltai is hiring in California to build and optimize high-performance ML pipelines and inference systems for LLMs in enterprise environments. You will design CUDA kernels, implement low-precision techniques, and Architect distributed training across multiple GPUs and nodes.

The role emphasizes production-ready research translate, collaboration with researchers and infra teams, and tuning for real-world hardware deployments in a fast-paced startup setting.

Qualifications

  • Expertise in systems-level programming for ML workloads.
  • Experience optimizing transformer kernels and GPU throughput.
  • Proficiency with GPU-accelerated inference frameworks.
  • Ability to translate research to production-grade systems.
  • Strong performance monitoring and bottleneck analysis.

Responsibilities

  • Design and maintain high-performance ML pipelines for training, evaluation, and inference of LLMs and retrieval-augmented systems, with a focus on hardware efficiency and throughput
  • Optimize core transformer operations at the kernel level, designing and tuning custom kernels and low-level implementations for GPU-accelerated workloads
  • Implement and integrate low-precision computation techniques to reduce memory footprint and accelerate inference with minimal accuracy degradation
  • Build and maintain inference engines for on premises deployments
  • Architect distributed training and inference systems
  • Collaborate closely with researchers and infra teams to bring cutting-edge model innovations into production
  • Interface directly with enterprise hardware environments, tuning performance based on real-world deployment constraints

Skills

C/C++/Rust
CUDA kernel design
Low-precision compute
Inference systems
Distributed training
Research to production
Performance monitoring

Tools

CUDA
vLLM
SGLang
TensorRT
GGUF/GGML

Job description

Voltai is hiring in California to build and optimize high-performance ML pipelines and inference systems for LLMs in enterprise environments. You will design CUDA kernels, implement low-precision techniques, and Architect distributed training across multiple GPUs and nodes.

The role emphasizes production-ready research translate, collaboration with researchers and infra teams, and tuning for real-world hardware deployments in a fast-paced startup setting.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Systems Engineer: Inference & GPU-Driven Distributed Workloads
ML Systems Engineer: Inference & GPU-Driven Distributed Workloads

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior ML Ops Engineer: Productionize LLMs at Scale
Senior ML Ops Engineer: Productionize LLMs at Scale

voltai-com • Edison (CA)

On-site
USD 150,000 - 230,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Senior ML Inference Engineer: High-Performance GPU Systems
Senior ML Inference Engineer: High-Performance GPU Systems

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
ML Systems Engineer: High-Performance Distributed Training
ML Systems Engineer: High-Performance Distributed Training

Motional • San Francisco (CA)

Hybrid
USD 144,000 - 192,000
Medical
Dental
Vision
+4
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Machine Learning Systems Engineer
Machine Learning Systems Engineer

voltai-com • Edison (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Member of Technical Staff, MLSys
Member of Technical Staff, MLSys

Bake AI • San Mateo (CA), Northern (KY)

On-site
USD 180,000 - 240,000