LLM Inference Engineer: Scale & Optimize Production

Comunidade Metodista

Palo Alto (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Hippocratic AI is seeking an experienced LLM Inference Engineer to optimize its LLM serving infrastructure in Palo Alto, CA. You will design multi-node architectures, implement multi-LoRA serving, and apply quantization techniques to reduce footprint while maintaining quality.

The role emphasizes practical production deployment, benchmarking, and latency-focused optimizations across GPUs, with a strong emphasis on Python, C++, and CUDA.

Qualifications

  • 2+ years of experience optimizing LLM inference systems at scale.
  • Proven expertise with distributed serving architectures for large language models.
  • Hands-on experience implementing quantization techniques for transformer models.
  • Strong understanding of modern inference optimization methods, including latency optimization strategies.
  • Proficiency in Python and C++.
  • Experience with CUDA programming and GPU optimization.

Responsibilities

  • Design and implement multi-node serving architectures for distributed LLM inference.
  • Optimize multi-LoRA serving systems.
  • Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality.
  • Implement speculative decoding and other latency optimization strategies.
  • Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases.
  • Continuously benchmark and improve system performance across various deployment scenarios and GPU types.

Skills

Python
C++
CUDA
Distributed systems
Quantization
Latency optimization

Tools

vLLM
TensorRT-LLM
SGLang
Custom CUDA kernels

Job description

Hippocratic AI is seeking an experienced LLM Inference Engineer to optimize its LLM serving infrastructure in Palo Alto, CA. You will design multi-node architectures, implement multi-LoRA serving, and apply quantization techniques to reduce footprint while maintaining quality.

The role emphasizes practical production deployment, benchmarking, and latency-focused optimizations across GPUs, with a strong emphasis on Python, C++, and CUDA.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Inference Architect (On-Site Menlo Park)
Senior LLM Inference Architect (On-Site Menlo Park)

Hippocratic AI Inc. • Menlo Park (CA)

On-site
USD 180,000 - 280,000
Equity
Health insurance
Competitive compensation
LLM Inference Engineer (Mid, Senior, Staff)
LLM Inference Engineer (Mid, Senior, Staff)

Hippocratic AI Inc. • Menlo Park (CA)

On-site
USD 180,000 - 280,000
Fast, Scalable LLM Inference Engineer — Healthcare AI
Fast, Scalable LLM Inference Engineer — Healthcare AI

Hippocratic AI Inc. • Menlo Park (CA)

On-site
USD 180,000 - 280,000
AI Infrastructure Engineer: GPU ML Inference & Scale
AI Infrastructure Engineer: GPU ML Inference & Scale

Next Frontier Capital • Albuquerque (NM)

On-site
USD 130,000 - 210,000
LLM Inference Engineer at Hippocratic AI Palo Alto, CA
LLM Inference Engineer at Hippocratic AI Palo Alto, CA

Comunidade Metodista • Palo Alto (CA)

On-site
USD 150,000 - 210,000
LLM Inference Engineer: Multi-GPU KV Cache & Throughput
LLM Inference Engineer: Multi-GPU KV Cache & Throughput

Triune Infomatics Inc • San Jose (CA)

Hybrid
USD 170,000 - 210,000
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Equity
Insurance for dependents
Winter Break
LLM/VLM Inference Optimization Research Engineer
LLM/VLM Inference Optimization Research Engineer

Bytedance • San Jose (CA)

On-site
USD 244,800 - 450,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+1
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Xapply • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive compensation and equity
Medical, dental, vision insurance (US)
Flexible PTO including Winter Break
+4
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1