Inference Infrastructure Engineer, Serving

Elorian AI

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Parental leave

Job summary

Elorian AI in Palo Alto, CA, seeks an infrastructure engineer to design, optimize, and scale systems that serve large multimodal models. You will drive faster, cheaper, and more reliable inference, enabling researchers to push model capabilities while reducing bottlenecks.

You will own infra for deployment and evaluation at scale, collaborating with researchers to enable high-performance inference across architectures, focusing on low latency, high throughput, and robust observability.

Qualifications

  • 3+ years building low-latency inference serving systems for large models.
  • Knowledge of quantization, batching, speculative decoding, KV cache management.
  • Experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang.
  • Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline).
  • Experience with autoscaling and load balancing for production ML services.
  • Track record of GPU cost optimization at scale.

Responsibilities

  • Build low-latency, high-throughput inference serving systems for large multimodal models.
  • Design latency, throughput, and efficiency improvements (quantization, batching, speculative decoding, KV cache).
  • Optimize codebase and GPU fleet to maximize FLOPs, bandwidth, and memory usage.
  • Implement multi-GPU and multi-node model parallelism for serving (tensor or pipeline).
  • Build autoscaling and load balancing for production ML services.
  • Establish reliability, observability, and reproducibility across the inference stack.
  • Collaborate with researchers to enable high-performance inference for novel architectures.

Skills

Low-latency serving
Inference optimization
Multi-GPU parallelism
Autoscaling
GPU cost optimization
Serving frameworks

Tools

vLLM
TensorRT-LLM
Triton
SGLang

Job description

We are a well-funded, early-stage AI lab focused on building the next generation of frontier multimodal AI models. Founded by former DeepMind researchers, including Andrew Dai, who was previously a leader on Gemini. Our team currently consists of 20 world-class scientists and engineers. We recently raised $55M in seed funding from Striker Ventures, Menlo Ventures, Altimeter Capital, and NVIDIA. We are tackling some of the hardest problems in artificial intelligence, and we are growing fast.

About the Role

We're looking for an infrastructure engineer to design, optimize, and scale the systems that serve our large multimodal models. Your work will make inference faster, more cost-effective, and more reliable, so our teams can focus on advancing model capabilities rather than managing bottlenecks.

Our focus is on performant, efficient inference, both to power real-world applications and to accelerate research. This role owns the infrastructure that ensures every deployment and evaluation runs smoothly at scale for our visual foundation models.

What You Will Do

  • Build low-latency, high-throughput inference serving systems for our large multimodal models
  • Design and implement techniques that improve latency, throughput, and efficiency, including quantization, batching, speculative decoding, and KV cache management
  • Optimize our codebase and GPU fleet to fully utilize hardware FLOPs, bandwidth, and memory
  • Implement multi-GPU and multi-node model parallelism for serving (tensor or pipeline parallel)
  • Build autoscaling and load balancing for production ML services
  • Establish standards for reliability, observability, and reproducibility across the inference stack
  • Collaborate with researchers to enable high-performance inference for novel architectures

Skills and Qualifications

  • 3+ years of experience building low-latency, high-throughput inference serving systems for large models
  • Strong knowledge of inference optimization techniques (quantization, batching, speculative decoding, KV cache management)
  • Hands-on experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang
  • Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline parallel)
  • Experience with autoscaling and load balancing for production ML services
  • A track record of GPU cost optimization at scale

Preferred qualifications (strong candidates may have some, not all):

  • Experience serving multimodal (vision + language) models
  • Contributions to open-source ML or systems infrastructure projects (e.g., vLLM, SGLang, TensorRT-LLM, Triton)
  • A bias for action and comfort working across stacks and teams in an early-stage environment

Logistics

  • Location: This role is based on-site in Palo Alto, California.
  • Compensation: Depending on background, skills, and experience, the expected annual base salary range for this position is $200,000 - $400,000 USD, plus equity and benefits.
  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • Benefits: We offer health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Elorian AI is an equal opportunity employer. We are committed to building a diverse team and inclusive environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options
Software Engineer, Inference - Multi Modal
Software Engineer, Inference - Multi Modal

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Inference Engineer
Inference Engineer

techire.® • San Francisco (CA)

On-site
USD 140,000 - 210,000
Medical insurance (including dental &视
Dental insurance
Vision insurance
+4
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Inference Engineer
Inference Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Office Bellevue
Competitive compensation
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Software Engineer, Model Inference
Software Engineer, Model Inference

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 490,000
Distributed Systems Engineer, Real-Time Inference at Scale
Distributed Systems Engineer, Real-Time Inference at Scale

adaption • San Francisco (CA)

On-site