Inference Infra Engineer: Scale Low-Latency AI Serving

Elorian

Palo Alto (CA)

On-site

USD 200,000 - 400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Parental leave
Relocation support

Job summary

Elorian is seeking an infrastructure engineer in Palo Alto to design, optimize, and scale systems for large multimodal models. You will focus on fast, cost-efficient inference and reliable deployments across production environments.

You will implement multi-GPU and multi-node model parallelism, autoscaling, and observability standards, collaborating with researchers to advance novel architectures and achieve high-performance inference.

Qualifications

  • 3+ years building low-latency inference serving systems for large models
  • Strong knowledge of inference optimization techniques (quantization, batching, speculative decoding, KV cache)
  • Hands-on experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang
  • Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline)
  • Strong systems programming skills; C++/CUDA a plus alongside Python
  • Experience with autoscaling and load balancing for production ML services
  • A track record of GPU cost optimization at scale

Responsibilities

  • Build low-latency, high-throughput inference serving systems for large multimodal models
  • Design and implement latency, throughput, and efficiency improvements (quantization, batching, speculative decoding, KV cache management)
  • Optimize codebase and GPU fleet for full hardware utilization
  • Implement multi-GPU/multi-node model parallelism for serving (tensor/pipeline)
  • Build autoscaling and load balancing for production ML services
  • Establish reliability, observability, and reproducibility standards across the inference stack
  • Collaborate with researchers to enable high-performance inference for novel architectures

Skills

Low-latency inference
Quantization
Multi-GPU/Node parallelism
C++/CUDA
Autoscaling & load balancing
GPU cost optimization

Tools

vLLM
TensorRT-LLM
Triton
SGLang

Job description

Elorian is seeking an infrastructure engineer in Palo Alto to design, optimize, and scale systems for large multimodal models. You will focus on fast, cost-efficient inference and reliable deployments across production environments.

You will implement multi-GPU and multi-node model parallelism, autoscaling, and observability standards, collaborating with researchers to advance novel architectures and achieve high-performance inference.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
AI Inference Engineer - Scalable, Low-Latency Systems
AI Inference Engineer - Scalable, Low-Latency Systems

SPACE EXPLORATION TECHNOLOGIES CORP • Palo Alto (CA), Northern (KY)

Hybrid
USD 135,000 - 210,000
401(k)
Medical, vision and dental coverage
Paid parental leave
+4
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
AI Inference Systems Engineer (High-Throughput, Low-Latency)
AI Inference Systems Engineer (High-Throughput, Low-Latency)

SpaceX • Palo Alto (CA)

On-site
USD 135,000 - 210,000
Stock options
Excellent medical coverage
401(k) plan
Inference Performance Engineer — Scale AI Serving
Inference Performance Engineer — Scale AI Serving

Adaption Labs • United States

Hybrid
USD 140,000 - 210,000
Flexible work
Adaption Passport travel stipend
Lunch stipend
+1
Inference Infrastructure Engineer for Large-Scale AI
Inference Infrastructure Engineer for Large-Scale AI

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Inference Infrastructure Engineer
AI Inference Infrastructure Engineer

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Distributed LLM Inference Engineer - Scale HighThroughput AI
Distributed LLM Inference Engineer - Scale HighThroughput AI

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6
AI Infrastructure Intern — Inference at Scale
AI Infrastructure Intern — Inference at Scale

DeepInfra • Palo Alto (CA)

On-site