Fast, Scalable LLM Inference Engineer — Healthcare AI

Hippocratic AI Inc.

Menlo Park (CA)

On-site

USD 180,000 - 280,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Hippocratic AI Inc. is seeking an LLM Inference Engineer to own the serving infrastructure enabling fast, reliable healthcare AI for patient conversations. You will optimize end-to-end latency (<100 ms) and cost, shaping the roadmap for scalable inference across production deployments.

By day 90 you will ship measurable improvements to the inference stack, validate gains, and outline a performance optimization plan. At 12 months you will deploy advanced serving architectures and contribute techniques that become core infra capabilities.

Qualifications

  • Experience with CUDA programming and GPU optimization.
  • Proficiency in Python and C++ for production systems.
  • Hands-on experience implementing quantization techniques for transformer models.

Responsibilities

  • Design and implement multi-node serving architectures for distributed LLM inference.
  • Optimize multi-LoRA serving systems.
  • Apply advanced quantization techniques to reduce model footprint while preserving quality.
  • Implement speculative decoding and other latency optimization strategies.
  • Develop disaggregated serving solutions with optimized caching for prefill and decoding phases.
  • Continuously benchmark and improve system performance across deployment scenarios and GPU types.

Skills

CUDA programming
GPU optimization
Python
C++
Quantization
Speculative decoding
Distributed serving
Open-source frameworks

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Hippocratic AI Inc. is seeking an LLM Inference Engineer to own the serving infrastructure enabling fast, reliable healthcare AI for patient conversations. You will optimize end-to-end latency (<100 ms) and cost, shaping the roadmap for scalable inference across production deployments.

By day 90 you will ship measurable improvements to the inference stack, validate gains, and outline a performance optimization plan. At 12 months you will deploy advanced serving architectures and contribute techniques that become core infra capabilities.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Inference Architect (On-Site Menlo Park)
Senior LLM Inference Architect (On-Site Menlo Park)

Hippocratic AI Inc. • Menlo Park (CA)

On-site
USD 180,000 - 280,000
Equity
Health insurance
Competitive compensation
LLM Inference Engineer (Mid, Sr, Staff)
LLM Inference Engineer (Mid, Sr, Staff)

Hippocratic AI Inc. • Menlo Park (CA)

On-site
USD 180,000 - 280,000
Equity
Health insurance
Competitive compensation
LLM Inference Engineer at Hippocratic AI Palo Alto, CA
LLM Inference Engineer at Hippocratic AI Palo Alto, CA

Comunidade Metodista • Palo Alto (CA)

On-site
USD 150,000 - 210,000
LLM Inference Engineer: Scale & Optimize Production
LLM Inference Engineer: Scale & Optimize Production

Comunidade Metodista • Palo Alto (CA)

On-site
USD 150,000 - 210,000
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Hippocratic AI Inc. • Menlo Park (CA), Northern (KY)

Hybrid
USD 220,000 - 320,000
LLM Inference Engineer (Mid, Senior, Staff)
LLM Inference Engineer (Mid, Senior, Staff)

Hippocratic AI Inc. • Menlo Park (CA)

On-site
USD 180,000 - 280,000
Healthcare AI Engineer (LLM, Remote)
Healthcare AI Engineer (LLM, Remote)

DeepScribe • Northern (KY)

Hybrid
USD 150,000 - 250,000
Meaningful equity
Flexible PTO
Work from home stipend
+1
Healthcare AI Forward Deployed Engineer Resident
Healthcare AI Forward Deployed Engineer Resident

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 150,000 - 210,000
Remote AI Engineer - Healthcare LLMs & Real-Time AI
Remote AI Engineer - Healthcare LLMs & Real-Time AI

SupportFinity™ • United States

On-site
USD 150,000 - 250,000
$150,000 to $250,000 annual salary.
Meaningful equity stake in the company
Flexible PTO
+2
LLM Inference Engineer — High-Performance AI Serving
LLM Inference Engineer — High-Performance AI Serving

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000