Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Perplexity in New York is seeking an experienced engineer to join our team and own ML inference workloads, building a Rust/Python-based serving runtime and CUDA kernel work that scales to production. You will work with transformer-based models, enabling efficient weight loading, scheduling, and KV-cache management in our in-house infrastructure.
You will read research papers, implement kernels, and diagnose production incidents in a fast-moving environment, collaborating across languages and
We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join us.
Compensation Range: $220K - $485K