Inference Infra Architect for Multimodal ML Systems

Elorian AI

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Parental leave

Job summary

Elorian AI in Palo Alto, CA, seeks an infrastructure engineer to design, optimize, and scale systems that serve large multimodal models. You will drive faster, cheaper, and more reliable inference, enabling researchers to push model capabilities while reducing bottlenecks.

You will own infra for deployment and evaluation at scale, collaborating with researchers to enable high-performance inference across architectures, focusing on low latency, high throughput, and robust observability.

Qualifications

  • 3+ years building low-latency inference serving systems for large models.
  • Knowledge of quantization, batching, speculative decoding, KV cache management.
  • Experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang.
  • Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline).
  • Experience with autoscaling and load balancing for production ML services.
  • Track record of GPU cost optimization at scale.

Responsibilities

  • Build low-latency, high-throughput inference serving systems for large multimodal models.
  • Design latency, throughput, and efficiency improvements (quantization, batching, speculative decoding, KV cache).
  • Optimize codebase and GPU fleet to maximize FLOPs, bandwidth, and memory usage.
  • Implement multi-GPU and multi-node model parallelism for serving (tensor or pipeline).
  • Build autoscaling and load balancing for production ML services.
  • Establish reliability, observability, and reproducibility across the inference stack.
  • Collaborate with researchers to enable high-performance inference for novel architectures.

Skills

Low-latency serving
Inference optimization
Multi-GPU parallelism
Autoscaling
GPU cost optimization
Serving frameworks

Tools

vLLM
TensorRT-LLM
Triton
SGLang

Job description

Elorian AI in Palo Alto, CA, seeks an infrastructure engineer to design, optimize, and scale systems that serve large multimodal models. You will drive faster, cheaper, and more reliable inference, enabling researchers to push model capabilities while reducing bottlenecks.

You will own infra for deployment and evaluation at scale, collaborating with researchers to enable high-performance inference across architectures, focusing on low latency, high throughput, and robust observability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Realtime Multimodal Inference Architect
Realtime Multimodal Inference Architect

techire.® • San Francisco (CA)

On-site
USD 140,000 - 210,000
Medical insurance (including dental &视
Dental insurance
Vision insurance
+4
Inference Systems Engineer for Scalable Multimodal AI
Inference Systems Engineer for Scalable Multimodal AI

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Senior Inference Systems Engineer — Low-Latency ML Serving
Senior Inference Systems Engineer — Low-Latency ML Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Socket.dev • Boston (MA)

On-site
USD 168,000 - 227,000
AI Inference Infrastructure Engineer
AI Inference Infrastructure Engineer

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
ML Infra Engineer - Ray + PyTorch for Multimodal AI
ML Infra Engineer - Ray + PyTorch for Multimodal AI

Bonfirevc • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Senior AI Infra Engineer — Real-Time Multimodal Inference
Senior AI Infra Engineer — Real-Time Multimodal Inference

Ambient AI, Inc. • Redwood City (CA)

Hybrid
USD 190,000 - 270,000
Stock options
Health, dental, vision
401(k)
+1
Head of ML Systems & Inference
Head of ML Systems & Inference

Doist • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Distributed LLM Inference & Optimization Engineer
Distributed LLM Inference & Optimization Engineer

Together AI • San Francisco (CA)

On-site
USD 160,000 - 230,000
Startup equity
Health insurance
Competitive benefits