Stand out for this role — generate a tailored resume and cover letter in about a minute.
Hippocratic AI Inc. is seeking an LLM Inference Engineer to own the serving infrastructure enabling fast, reliable healthcare AI for patient conversations. You will optimize end-to-end latency (<100 ms) and cost, shaping the roadmap for scalable inference across production deployments.
By day 90 you will ship measurable improvements to the inference stack, validate gains, and outline a performance optimization plan. At 12 months you will deploy advanced serving architectures and contribute techniques that become core infra capabilities.