An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Perplexity is building and running the inference engine behind every Perplexity query, deploying dozens of model architectures at scale with tight latency and cost budgets. You will join a Rust/Python/CUDA stack to advance the serving runtime, GPU kernels, and inference performance.
We seek someone with deep GPU programming experience, production distributed systems know-how, and the ability to read papers, implement kernels, and debug incidents rapidly in a fast-moving environment.
Perplexity is building and running the inference engine behind every Perplexity query, deploying dozens of model architectures at scale with tight latency and cost budgets. You will join a Rust/Python/CUDA stack to advance the serving runtime, GPU kernels, and inference performance.
We seek someone with deep GPU programming experience, production distributed systems know-how, and the ability to read papers, implement kernels, and debug incidents rapidly in a fast-moving environment.