Inference Infra Architect for Multimodal ML Systems

Elorian AI

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Parental leave

Job summary

Elorian AI in Palo Alto, CA, seeks an infrastructure engineer to design, optimize, and scale systems that serve large multimodal models. You will drive faster, cheaper, and more reliable inference, enabling researchers to push model capabilities while reducing bottlenecks.

You will own infra for deployment and evaluation at scale, collaborating with researchers to enable high-performance inference across architectures, focusing on low latency, high throughput, and robust observability.

Qualifications

  • 3+ years building low-latency inference serving systems for large models.
  • Knowledge of quantization, batching, speculative decoding, KV cache management.
  • Experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang.
  • Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline).
  • Experience with autoscaling and load balancing for production ML services.
  • Track record of GPU cost optimization at scale.

Responsibilities

  • Build low-latency, high-throughput inference serving systems for large multimodal models.
  • Design latency, throughput, and efficiency improvements (quantization, batching, speculative decoding, KV cache).
  • Optimize codebase and GPU fleet to maximize FLOPs, bandwidth, and memory usage.
  • Implement multi-GPU and multi-node model parallelism for serving (tensor or pipeline).
  • Build autoscaling and load balancing for production ML services.
  • Establish reliability, observability, and reproducibility across the inference stack.
  • Collaborate with researchers to enable high-performance inference for novel architectures.

Skills

Low-latency serving
Inference optimization
Multi-GPU parallelism
Autoscaling
GPU cost optimization
Serving frameworks

Tools

vLLM
TensorRT-LLM
Triton
SGLang

Job description

Elorian AI in Palo Alto, CA, seeks an infrastructure engineer to design, optimize, and scale systems that serve large multimodal models. You will drive faster, cheaper, and more reliable inference, enabling researchers to push model capabilities while reducing bottlenecks.

You will own infra for deployment and evaluation at scale, collaborating with researchers to enable high-performance inference across architectures, focusing on low latency, high throughput, and robust observability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Low-Latency Inference Infrastructure Engineer
Low-Latency Inference Infrastructure Engineer

Elorian • Palo Alto (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Inference Systems Engineer for Scalable Multimodal AI
Inference Systems Engineer for Scalable Multimodal AI

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Engineering Manager: Inference Infrastructure Leader
Engineering Manager: Inference Infrastructure Leader

EngineersOfAI • New York (NY), Northern (KY)

Hybrid
USD 230,000 - 360,000
Senior ML Infrastructure Engineer - Online Inference
Senior ML Infrastructure Engineer - Online Inference

Unity Enterprise • Northern (KY)

Hybrid
USD 210,000 - 273,000
Comprehensive health insurance
Employee stock ownership
Generous vacation and personal days
+1
Inference Systems Engineer: Optimize AI Serving & Latency
Inference Systems Engineer: Optimize AI Serving & Latency

Emploive • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Flexible work
Adaption Passport
Lunch Stipend
+1
Inference Engineer: High-Throughput ML Serving
Inference Engineer: High-Throughput ML Serving

Adaption Labs, Inc. • San Francisco (CA)

Hybrid
USD 180,000 - 230,000
Flexible work: In-person collaboration
Adaption Passport
Lunch Stipend
+1
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon Inc. • Boston (MA)

On-site
USD 167,000 - 226,000