LLM Inference Engineer (Mid, Senior, Staff)

Hippocratic AI Inc.

Menlo Park (CA)

On-site

USD 180,000 - 280,000

Full time

45 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Hippocratic AI Inc. is seeking an LLM Inference Engineer to own the serving infrastructure enabling fast, reliable healthcare AI for patient conversations. You will optimize end-to-end latency (<100 ms) and cost, shaping the roadmap for scalable inference across production deployments.

By day 90 you will ship measurable improvements to the inference stack, validate gains, and outline a performance optimization plan. At 12 months you will deploy advanced serving architectures and contribute techniques that become core infra capabilities.

Qualifications

  • Experience with CUDA programming and GPU optimization.
  • Proficiency in Python and C++ for production systems.
  • Hands-on experience implementing quantization techniques for transformer models.

Responsibilities

  • Design and implement multi-node serving architectures for distributed LLM inference.
  • Optimize multi-LoRA serving systems.
  • Apply advanced quantization techniques to reduce model footprint while preserving quality.
  • Implement speculative decoding and other latency optimization strategies.
  • Develop disaggregated serving solutions with optimized caching for prefill and decoding phases.
  • Continuously benchmark and improve system performance across deployment scenarios and GPU types.

Skills

CUDA programming
GPU optimization
Python
C++
Quantization
Speculative decoding
Distributed serving
Open-source frameworks

Tools

vLLM
SGLang
TensorRT-LLM

Job description

  • As HAI’s LLM Inference Engineer, you will own the serving infrastructure that determines whether our breakthrough healthcare AI reaches patients efficiently and reliably
  • You’ll optimize the systems that translate raw model capability into sub-100ms responses—making the difference between conversational experiences that feel natural and those that feel broken
  • This role exists because inference optimization at scale is where research meets reality: your work directly determines latency, cost, and availability for millions of patient conversations across healthcare systems
  • Own your first major outcome: By day 90, you will have shipped a measurable improvement to our inference serving stack (reduced latency, improved throughput, or optimized cost per inference), validated the gains across our production deployment scenarios, and established the performance optimization roadmap that will guide infrastructure investment
  • Drive lasting impact: At 12 months, you will have designed and deployed advanced serving architectures (disaggregated inference, optimized caching, speculative decoding) that meaningfully improve patient experience and operational efficiency, contributed novel optimization techniques that become part of our core infrastructure, and made our serving stack a durable competitive advantage in healthcare AI deployment
  • You’ll work alongside systems engineers, ML researchers, and infrastructure experts who are obsessed with making AI systems fast, reliable, and cost-effective. This is a team that values deep technical rigor, continuous benchmarking, and solving hard systems problems that have real impact on patient experience and business unit economics
  • Design and implement multi-node serving architectures for distributed LLM inference
  • Optimize multi-LoRA serving systems
  • Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality
  • Implement speculative decoding and other latency optimization strategies
  • Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases
  • Continuously benchmark and improve system performance across various deployment scenarios and GPU types
  • Experience with CUDA programming and GPU optimization
  • Proficiency in Python and C++
  • Hands-on experience implementing quantization techniques for transformer models
  • Experience optimizing LLM inference systems at scale
  • Speculative decoding techniques with draft models
  • Proven expertise with distributed serving architectures for large language models
  • Eagle speculative decoding approaches
  • Strong understanding of modern inference optimization methods, including:
  • Track record of deploying inference systems in production environments
  • Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM
  • Deep understanding of performance optimization systems
  • Experience with custom CUDA kernels
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)

Bitus Labs • Irvine (CA)

On-site
USD 140,000 - 190,000
Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise • Fort Collins (CO)

On-site
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements
Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud, Inc • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallel Wireless • United States

On-site
USD 180,000 - 280,000