High-Performance Inference Engineer (ML Systems)

Garuda Ventures

Hermosa Beach (CA)

On-site

USD 100,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Garuda Ventures is looking for engineers to bridge the gap between ML research and high-performance inference. You will work on our inference engine and model conversion toolkit, implementing new model architectures and building a wide range of features.

This role is perfect for those passionate about writing high-performance code and eager to learn. If you have experience in areas such as JAX, Rust systems programming, and benchmarking, we want to hear from you!

Qualifications

  • Passion for reading ML research papers.
  • High-performance code writing proficiency.
  • Basic engineering skills with shipping code experience.

Responsibilities

  • Work on inference engine and model conversion toolkit.
  • Implement new model architectures and support new modalities.
  • Write optimized kernels and build new features.

Skills

JAX / Equinox / Pallas stack
Rust systems programming
Writing Metal / Vulkan kernels
Neural codecs and voice model architectures
Trellis-based quantization approaches
Advanced speculative decoding methods
Understanding of Transformer / SSM / Diffusion / Vision language models
Benchmarking inference performance and model quality
Strong linear algebra and optimization methods

Job description

Garuda Ventures is looking for engineers to bridge the gap between ML research and high-performance inference. You will work on our inference engine and model conversion toolkit, implementing new model architectures and building a wide range of features.

This role is perfect for those passionate about writing high-performance code and eager to learn. If you have experience in areas such as JAX, Rust systems programming, and benchmarking, we want to hear from you!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference engineer
Inference engineer

Garuda Ventures • Hermosa Beach (CA)

On-site
USD 100,000 - 140,000
Remote ML Systems Engineer – High-Performance Inference
Remote ML Systems Engineer – High-Performance Inference

Triwill Group • United States

Remote
USD 145,000 - 165,000
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2
GenAI Inference Engineer — High-Performance ML Systems
GenAI Inference Engineer — High-Performance ML Systems

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 142,000 - 205,000
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

Relha LLC • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2