AI Inference Engineer – High-Performance GPU Systems

Perplexity

California (MO)

On-site

USD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Perplexity is seeking an engineer to join our inference stack. You will work on GPU-accelerated ML inference, supporting transformer models, caching, and low-latency serving. You\'ll collaborate across Rust, Python, CUDA, CuTe DSL, and deploy production distributed systems under real load with a focus on performance and reliability.

You will contribute to a stack built around PyTorch, CUDA kernels, and modern ML tooling, delivering scalable inference in a fast-paced environment.

Qualifications

  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).

Responsibilities

  • Develop models support for transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.
  • GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.
  • Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
  • Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
  • Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents.

Skills

GPU programming
LLM architectures
Distributed systems
Self-directed
Rust/Python/CUDA

Education

Bachelor's degree in Computer Science or related field

Tools

PyTorch
CUDA
Kubernetes
Nsight Compute/Systems

Job description

Perplexity is seeking an engineer to join our inference stack. You will work on GPU-accelerated ML inference, supporting transformer models, caching, and low-latency serving. You\'ll collaborate across Rust, Python, CUDA, CuTe DSL, and deploy production distributed systems under real load with a focus on performance and reliability.

You will contribute to a stack built around PyTorch, CUDA kernels, and modern ML tooling, delivering scalable inference in a fast-paced environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • California (MO)

On-site
USD 120,000 - 170,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 220,000 - 485,000
Senior AI Inference Engineer - GPU, Rust & CUDA
Senior AI Inference Engineer - GPU, Rust & CUDA

Perplexity • San Francisco (CA)

On-site
USD 220,000 - 485,000
Member of Technical Staff (Software Engineer, Inference & Training Platform) Perplexity AI San [...]
Member of Technical Staff (Software Engineer, Inference & Training Platform) Perplexity AI San [...]

Neura Market • San Francisco (CA)

On-site
USD 180,000 - 260,000
Hybrid AI Inference Engineer — Kernel & Performance
Hybrid AI Inference Engineer — Kernel & Performance

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Senior GPU Inference Engineer for Real-Time AI
Senior GPU Inference Engineer for Real-Time AI

Cerebras • United States

On-site
USD 150,000 - 210,000
Staff Software Engineer — AI Acceleration & Infrastructure
Staff Software Engineer — AI Acceleration & Infrastructure

Perplexity AI Inc. • New York (NY)

Hybrid
USD 120,000 - 180,000
Staff Inference Systems Engineer — High-Throughput AI
Staff Inference Systems Engineer — High-Throughput AI

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2