Get more replies from employers
Send a job-specific resume in minutes.
DevHub is building scalable LLM infrastructure to power large-scale inference workloads. You will work with cross-functional teams to improve reliability, latency, and efficiency of distributed AI systems in a fast-growing environment.
We seek a senior backend/infrastructure engineer with 8+ years' experience in distributed systems, scalable APIs, and cloud-native infrastructure. Expertise in ML infrastructure, GPU orchestration, and SOA is essential; PyTorch and vLLM experience is a plus.
Build and optimize LLM infrastructure to power large-scale inference workloads for both partner and self-hosted models. Collaborate with cross-functional teams to improve the reliability, latency, and efficiency of distributed AI workloads.