An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Perplexity is building and running the inference engine behind every Perplexity query, deploying dozens of model architectures at scale with tight latency and cost budgets. We are seeking another engineer to join our team to advance transformer-based retrieval, text generation, and multimodal models in production.
You will work on a Rust-native serving runtime, CUDA/CuTe kernel migration, and ongoing performance optimizations to support growing traffic while maintaining reliability.
Perplexity is building and running the inference engine behind every Perplexity query, deploying dozens of model architectures at scale with tight latency and cost budgets. We are seeking another engineer to join our team to advance transformer-based retrieval, text generation, and multimodal models in production.
You will work on a Rust-native serving runtime, CUDA/CuTe kernel migration, and ongoing performance optimizations to support growing traffic while maintaining reliability.