Inference engineer

Garuda Ventures

Hermosa Beach (CA)

On-site

USD 100,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Garuda Ventures is looking for engineers to bridge the gap between ML research and high-performance inference. You will work on our inference engine and model conversion toolkit, implementing new model architectures and building a wide range of features.

This role is perfect for those passionate about writing high-performance code and eager to learn. If you have experience in areas such as JAX, Rust systems programming, and benchmarking, we want to hear from you!

Qualifications

  • Passion for reading ML research papers.
  • High-performance code writing proficiency.
  • Basic engineering skills with shipping code experience.

Responsibilities

  • Work on inference engine and model conversion toolkit.
  • Implement new model architectures and support new modalities.
  • Write optimized kernels and build new features.

Skills

JAX / Equinox / Pallas stack
Rust systems programming
Writing Metal / Vulkan kernels
Neural codecs and voice model architectures
Trellis-based quantization approaches
Advanced speculative decoding methods
Understanding of Transformer / SSM / Diffusion / Vision language models
Benchmarking inference performance and model quality
Strong linear algebra and optimization methods

Job description

We're looking for engineers who can bridge the gap between ML research and high-performance inference.

You'll work across our inference engine and model conversion toolkit, implementing new model architectures, supporting new modalities, writing optimized kernels, and building a wide range of features such as function calling and batch decoding.

This role is ideal for someone who reads papers for fun, enjoys writing high-performance code, and gets excited about constant learning.

Nobody knows everything. We'd rather you know one area deeply than everything superficially. If you're good at least in a couple of these areas, you're a great fit:

  • JAX / Equinox / Pallas stack
  • Rust systems programming with a focus on developer experience
  • Writing Metal / Vulkan kernels
  • Neural codecs and voice model architectures
  • Trellis-based quantization approaches
  • Advanced speculative decoding methods, such as EAGLE
  • Deep understanding of Transformer / SSM / Diffusion / Vision language models
  • Benchmarking inference performance and model quality
  • Strong linear algebra, optimization methods, and probability theory

And of course, basic engineering skills, we will ship a lot of code

We welcome applications from students and early-career engineers. If you've participated in projects that demonstrate systems thinking and ML understanding, we want to hear from you!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

High-Performance Inference Engineer (ML Systems)
High-Performance Inference Engineer (ML Systems)

Garuda Ventures • Hermosa Beach (CA)

On-site
USD 100,000 - 140,000
Inference Engineer
Inference Engineer

techire.® • San Francisco (CA)

On-site
USD 140,000 - 210,000
Medical insurance (including dental &视
Dental insurance
Vision insurance
+4
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 220,000 - 485,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits