Low-Latency Inference Infrastructure Engineer

Elorian

Palo Alto, Northern (CA, KY)

Hybrid

USD 200,000 - 400,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
Relocation support

Job summary

Elorian AI in Palo Alto is seeking an infrastructure engineer to design, optimize, and scale the systems that serve our large multimodal models. Your work will make inference faster, more cost-efficient, and more reliable, so our teams can focus on advancing model capabilities rather than managing bottlenecks.

Our focus is on performant, efficient inference, both to power real-world applications and to accelerate research.

Qualifications

  • 3+ years of experience building low-latency, high-throughput inference serving systems for large models.
  • Strong knowledge of inference optimization techniques (quantization, batching, speculative decoding, KV cache management).
  • Hands-on experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang.
  • Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline parallel).
  • Strong systems programming skills; C++/CUDA a plus alongside Python.
  • Experience with autoscaling and load balancing for production ML services.
  • A track record of GPU cost optimization at scale.

Responsibilities

  • Build low-latency, high-throughput inference serving systems for our large multimodal models.
  • Design and implement techniques that improve latency, throughput, and efficiency, including quantization, batching, speculative decoding, and KV cache management.
  • Optimize our codebase and GPU fleet to fully utilize hardware FLOPs, bandwidth, and memory.
  • Implement multi-GPU and multi-node model parallelism for serving (tensor or pipeline parallel).
  • Build autoscaling and load balancing for production ML services.
  • Establish standards for reliability, observability, and reproducibility across the inference stack.
  • Collaborate with researchers to enable high-performance inference for novel architectures.

Skills

Low-latency inference
Inference optimization
Serving frameworks
Multi-GPU parallelism
Systems programming (C++/CUDA)
Autoscaling & load balancing
GPU cost optimization

Tools

vLLM
TensorRT-LLM
Triton
SGLang

Job description

Elorian AI in Palo Alto is seeking an infrastructure engineer to design, optimize, and scale the systems that serve our large multimodal models. Your work will make inference faster, more cost-efficient, and more reliable, so our teams can focus on advancing model capabilities rather than managing bottlenecks.

Our focus is on performant, efficient inference, both to power real-world applications and to accelerate research.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Inference Systems Engineer: Optimize AI Serving & Latency
Inference Systems Engineer: Optimize AI Serving & Latency

Emploive • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Flexible work
Adaption Passport
Lunch Stipend
+1
AI Inference Systems Engineer (High-Throughput, Low-Latency)
AI Inference Systems Engineer (High-Throughput, Low-Latency)

SpaceX • Palo Alto (CA)

On-site
USD 135,000 - 210,000
Stock options
Excellent medical coverage
401(k) plan
AI Inference Engineer - Scalable, Low-Latency Systems
AI Inference Engineer - Scalable, Low-Latency Systems

SPACE EXPLORATION TECHNOLOGIES CORP • Palo Alto (CA), Northern (KY)

Hybrid
USD 135,000 - 210,000
401(k)
Medical, vision and dental coverage
Paid parental leave
+4
Inference Systems Engineer — High-Performance AI Serving
Inference Systems Engineer — High-Performance AI Serving

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
ML Infrastructure & Platform Engineer
ML Infrastructure & Platform Engineer

Odyssey • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Inference Engineer: High-Throughput ML Serving
Inference Engineer: High-Throughput ML Serving

Adaption Labs, Inc. • San Francisco (CA)

Hybrid
USD 180,000 - 230,000
Flexible work: In-person collaboration
Adaption Passport
Lunch Stipend
+1