Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience in software engineering with a focus on ML inference, familiarity with deep learning frameworks, and a strong understanding of GPU architectures. Compensation ranges from $220K to $485K.
Qualifications
3+ years of experience in ML inference or high-performance systems.
Familiar with a deep learning framework like PyTorch or TensorFlow.
Understanding of GPU architectures and optimization techniques.
Responsibilities
Support transformer-based models in inference infrastructure.
Migrate GPU kernels to CuTe DSL.
Develop Rust-based inference server and optimize performance.
Build dashboards and automated remediation for production reliability.
Skills
GPU programming
Performance optimization
Distributed systems
Rust
Python
CUDA
LLM architectures
Tools
PyTorch
CUDA-GDB
Kubernetes
Job description
Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience in software engineering with a focus on ML inference, familiarity with deep learning frameworks, and a strong understanding of GPU architectures. Compensation ranges from $220K to $485K.