Machine Learning Systems Engineer

voltai-com

Edison (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
Professional Growth
Visa Sponsorship

Job summary

Voltai is hiring in California to build and optimize high-performance ML pipelines and inference systems for LLMs in enterprise environments. You will design CUDA kernels, implement low-precision techniques, and Architect distributed training across multiple GPUs and nodes.

The role emphasizes production-ready research translate, collaboration with researchers and infra teams, and tuning for real-world hardware deployments in a fast-paced startup setting.

Qualifications

  • Expertise in systems-level programming for ML workloads.
  • Experience optimizing transformer kernels and GPU throughput.
  • Proficiency with GPU-accelerated inference frameworks.
  • Ability to translate research to production-grade systems.
  • Strong performance monitoring and bottleneck analysis.

Responsibilities

  • Design and maintain high-performance ML pipelines for training, evaluation, and inference of LLMs and retrieval-augmented systems, with a focus on hardware efficiency and throughput
  • Optimize core transformer operations at the kernel level, designing and tuning custom kernels and low-level implementations for GPU-accelerated workloads
  • Implement and integrate low-precision computation techniques to reduce memory footprint and accelerate inference with minimal accuracy degradation
  • Build and maintain inference engines for on premises deployments
  • Architect distributed training and inference systems
  • Collaborate closely with researchers and infra teams to bring cutting-edge model innovations into production
  • Interface directly with enterprise hardware environments, tuning performance based on real-world deployment constraints

Skills

C/C++/Rust
CUDA kernel design
Low-precision compute
Inference systems
Distributed training
Research to production
Performance monitoring

Tools

CUDA
vLLM
SGLang
TensorRT
GGUF/GGML

Job description

About Voltai

Voltai is the leading AI company building agentic systems and frontier foundation models for semiconductor and electronics design. Backed by Sequoia Capital, we’re putting AI in the hands of hardware engineers in over 70% of the world’s largest semiconductor and electronics companies to have effortless control over their next-generation chip and board designs, powering the future of automotive, industrial automation, consumer electronics, IoT, and semiconductor manufacturing.

About the Team

Our founding team consists of IOI/IPhO olympiad medalists, Stanford professors, ex-CTO of Synopsys, and our business leadership has scaled revenue in their previous companies to over $1.5bn. At Voltai, we are combining the world’s best talent in the intersection of software and hardware.

Key Responsibilities
  • Design and maintain high-performance ML pipelines for training, evaluation, and inference of LLMs and retrieval-augmented systems, with a focus on hardware efficiency and throughput
  • Optimize core transformer operations at the kernel level, designing and tuning custom kernels and low-level implementations for GPU-accelerated workloads
  • Implement and integrate low-precision computation techniques to reduce memory footprint and accelerate inference with minimal accuracy degradation
  • Build and maintain inference engines for on premises deployments
  • Architect distributed training and inference systems
  • Collaborate closely with researchers and infra teams to bring cutting-edge model innovations into production
  • Interface directly with enterprise hardware environments, tuning performance based on real-world deployment constraints
Required Skill Sets
  • Languages: Expertise in C, C++, or Rust
  • Design and Optimize CUDA Kernels for LLMs: Develop and fine-tune custom CUDA kernels to accelerate core transformer operations
  • Implement Low-Precision Computation Techniques: Apply quantization methods like AWQ and GPTQ to reduce model size and inference latency. Ensure minimal accuracy loss while maximizing throughput on GPU architectures with familiarity with concepts like GGUF and GGML
  • Develop and Maintain High-Performance Inference Systems: Build, improve, and maintain inference engines such as vLLM, SGLang, and TensorRT with a focus on low-latency and high throughput
  • Architect Distributed Training and Inference Solutions: Design systems that support model parallelism (tensor, pipeline, expert etc) to enable efficient training and inference across multiple GPUs and nodes
  • Integrate Research into Production Systems: Translate cutting-edge research findings into robust, production-ready systems. Ensure that innovations in model architectures and optimization techniques are effectively deployed.
  • Monitor and Optimize System Performance: Implement monitoring tools to track system metrics, identify bottlenecks, and optimize performance
Bonus Points
  • Some background in hardware/electronics, gained through professional, academic, or personal projects
  • Contributions to open-source initiatives
  • Notable awards or publications in leading journals/conferences
  • Experience thriving in a fast-paced, hyper-growth startup environment
Our Benefits
  • Unlimited PTO: Recharge when you need it, no questions asked.
  • Comprehensive Health Coverage: Medical, dental, and vision insurance for you and your dependents.
  • Free Meals and Snacks: Daily lunches, dinners, and snacks in the office.
  • Professional Growth: We invest in your continuous learning and offer opportunities to expand your skills.
  • Visa Sponsorship: We welcome global talent and provide visa sponsorship to support qualified candidates.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Research Engineer
Machine Learning Research Engineer

voltai-com • Edison (CA)

On-site
USD 180,000 - 250,000
Unlimited PTO
Health coverage
Meals & snacks
+2
Machine Learning Research Scientist
Machine Learning Research Scientist

voltai-com • Edison (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Machine Learning Operations
Machine Learning Operations

voltai-com • Edison (CA)

On-site
USD 150,000 - 230,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Software Engineering - Data Engineer
Software Engineering - Data Engineer

voltai-com • Edison (CA)

On-site
USD 140,000 - 220,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Software Engineer
Software Engineer

voltai-com • Edison (CA)

On-site
USD 140,000 - 210,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Software Engineer - Backend Engineer
Software Engineer - Backend Engineer

voltai-com • Edison (CA)

On-site
USD 180,000 - 240,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Engineering Manager
Engineering Manager

voltai-com • Edison (CA)

On-site
USD 150,000 - 230,000
Unlimited PTO
Comprehensive Health Coverage
Free meals and snacks
+2
Applied AI Engineer
Applied AI Engineer

voltai-com • Edison (CA)

On-site
USD 140,000 - 210,000
Unlimited PTO
Health coverage
Free meals & snacks
+2
Software Engineering - Infrastructure
Software Engineering - Infrastructure

voltai-com • Edison (CA)

On-site
USD 130,000 - 170,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2
Strategic Projects Lead
Strategic Projects Lead

voltai-com • Edison (CA)

On-site
USD 120,000 - 180,000
Unlimited PTO
Comprehensive Health Coverage
Free Meals and Snacks
+2