We are looking for a Performance Engineer, Inference to understand and improve the systems that
serve our foundation models. Inference is a tightly coupled system spanning model execution, serving
runtimes, distributed systems, accelerators, scheduling, memory, and reliability. You will measure the
system end to end, identify the highest-leverage performance gaps, and work across teams to close
them while preserving correctness.
Responsibilities
- Run cross-layer performance investigations across throughput, latency, memory efficiency, reliability, and cost.
- Build profiling, benchmarking, and observability tools that make inference performance measurable and explainable.
- Identify bottlenecks across model servers, batching and scheduling, distributed execution, memory systems, and accelerators.
- Partner with model, platform, and infrastructure teams to prioritize and land high-impact optimizations.
- Validate that performance improvements preserve model quality and numerical correctness.
Minimum Qualifications
- Hands-on experience profiling and optimizing ML systems or other performance-critical production systems.
- Production experience with at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, or an equivalent serving runtime.
- Experience serving or operating large models across accelerator-backed infrastructure, including multi-GPU or multi-accelerator systems.
- Strong Python skills and the ability to read, instrument, and modify large production codebases.
- Solid understanding of transformer inference, distributed systems, latency/throughput trade-offs, and accelerator performance fundamentals.
Preferred Qualifications
- Experience with large-scale or multi-node inference, including tensor, pipeline, data, or expert parallelism.
- Experience with GPUs, TPUs, NPUs, or other ML accelerators and associated profiling tools.
- Experience with quantization, low-precision inference, KV-cache optimization, speculative decoding, or long-context serving.
- Experience contributing to or modifying inference runtimes, kernels, compilers, or distributed serving components.
- Experience optimizing inference for constrained or on-device environments.