AI Inference Engineer — High-Performance GPU Systems

Perplexity

Palo Alto (CA)

On-site

USD 220,000 - 485,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Perplexity is building and running the inference engine behind every Perplexity query, deploying dozens of model architectures at scale with tight latency and cost budgets. You will join a Rust/Python/CUDA stack to advance the serving runtime, GPU kernels, and inference performance.

We seek someone with deep GPU programming experience, production distributed systems know-how, and the ability to read papers, implement kernels, and debug incidents rapidly in a fast-moving environment.

Qualifications

  • 3+ years of professional software engineering in ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • Understanding of common LLM architectures and inference optimization techniques (quantization, speculative decoding, prefill-decode disaggregation).

Responsibilities

  • Develop new models and support transformer-based retrieval, text-generation, and multimodal models in inference infrastructure.
  • Migrate GPU kernels to CuTe DSL for portability and performance.
  • Build Rust-native serving runtime to handle growing traffic and reduce Python pains.
  • Profile and fix bottlenecks from network to batching and GPU kernel interleaving.
  • Develop dashboards, alerts, and remediation for production incidents.

Skills

GPU programming
LLM architectures
Distributed systems
Rust Python CUDA
End-to-end ownership
Self-directed

Tools

Nsight Compute
CUDA-GDB
PTX/SASS
Kubernetes

Job description

Perplexity is building and running the inference engine behind every Perplexity query, deploying dozens of model architectures at scale with tight latency and cost budgets. You will join a Rust/Python/CUDA stack to advance the serving runtime, GPU kernels, and inference performance.

We seek someone with deep GPU programming experience, production distributed systems know-how, and the ability to read papers, implement kernels, and debug incidents rapidly in a fast-moving environment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Inference Engineer: GPU Performance & Rust Stack
AI Inference Engineer: GPU Performance & Rust Stack

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI Inference Engineer (GPU/Rust/CUDA)
AI Inference Engineer (GPU/Rust/CUDA)

Perplexity • New York (NY)

On-site
USD 220,000 - 485,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 485,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • New York (NY)

On-site
USD 220,000 - 485,000
Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Strategic Finance Lead for GPU Compute & Capacity Planning
Strategic Finance Lead for GPU Compute & Capacity Planning

Perplexity AI • United States

On-site
USD 140,000 - 210,000
Staff AI Systems Engineer – Model Serving & Infra
Staff AI Systems Engineer – Model Serving & Infra

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 405,000
Staff Inference Systems Engineer — High-Throughput AI
Staff Inference Systems Engineer — High-Throughput AI

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000