Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Hippocratic AI is seeking an experienced LLM Inference Engineer to optimize its LLM serving infrastructure in Palo Alto, CA. You will design multi-node architectures, implement multi-LoRA serving, and apply quantization techniques to reduce footprint while maintaining quality.
The role emphasizes practical production deployment, benchmarking, and latency-focused optimizations across GPUs, with a strong emphasis on Python, C++, and CUDA.
Hippocratic AI is seeking an experienced LLM Inference Engineer to optimize its LLM serving infrastructure in Palo Alto, CA. You will design multi-node architectures, implement multi-LoRA serving, and apply quantization techniques to reduce footprint while maintaining quality.
The role emphasizes practical production deployment, benchmarking, and latency-focused optimizations across GPUs, with a strong emphasis on Python, C++, and CUDA.