Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Runpod is hiring a seasoned systems engineer to optimize LLM inference performance in a remote-first environment. You will define measurement pipelines for throughput, latency, and cost per token, and build repeatable tooling for rigorous benchmarking.
You will profile the serving stack from scheduler and memory to GPU kernels, diagnose bottlenecks, and implement optimizations for large models on single and multi-node GPU deployments. Join a fast-moving AI infrastructure team.
Runpod is hiring a seasoned systems engineer to optimize LLM inference performance in a remote-first environment. You will define measurement pipelines for throughput, latency, and cost per token, and build repeatable tooling for rigorous benchmarking.
You will profile the serving stack from scheduler and memory to GPU kernels, diagnose bottlenecks, and implement optimizations for large models on single and multi-node GPU deployments. Join a fast-moving AI infrastructure team.