Get more replies from employers
Send a job-specific resume in minutes.
Perplexity is seeking an engineer to join our inference stack. You will work on GPU-accelerated ML inference, supporting transformer models, caching, and low-latency serving. You\'ll collaborate across Rust, Python, CUDA, CuTe DSL, and deploy production distributed systems under real load with a focus on performance and reliability.
You will contribute to a stack built around PyTorch, CUDA kernels, and modern ML tooling, delivering scalable inference in a fast-paced environment.
Perplexity is seeking an engineer to join our inference stack. You will work on GPU-accelerated ML inference, supporting transformer models, caching, and low-latency serving. You\'ll collaborate across Rust, Python, CUDA, CuTe DSL, and deploy production distributed systems under real load with a focus on performance and reliability.
You will contribute to a stack built around PyTorch, CUDA kernels, and modern ML tooling, delivering scalable inference in a fast-paced environment.