An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Perplexity in New York is seeking an experienced engineer to join our team and own ML inference workloads, building a Rust/Python-based serving runtime and CUDA kernel work that scales to production. You will work with transformer-based models, enabling efficient weight loading, scheduling, and KV-cache management in our in-house infrastructure.
You will read research papers, implement kernels, and diagnose production incidents in a fast-moving environment, collaborating across languages and
Perplexity in New York is seeking an experienced engineer to join our team and own ML inference workloads, building a Rust/Python-based serving runtime and CUDA kernel work that scales to production. You will work with transformer-based models, enabling efficient weight loading, scheduling, and KV-cache management in our in-house infrastructure.
You will read research papers, implement kernels, and diagnose production incidents in a fast-moving environment, collaborating across languages and