AI Inference Engineer: GPU Performance & Rust Stack

Perplexity

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Perplexity is building and running the inference engine behind every Perplexity query, deploying dozens of model architectures at scale with tight latency and cost budgets. We are seeking another engineer to join our team to advance transformer-based retrieval, text generation, and multimodal models in production.

You will work on a Rust-native serving runtime, CUDA/CuTe kernel migration, and ongoing performance optimizations to support growing traffic while maintaining reliability.

Qualifications

  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.

Responsibilities

  • Develop new model support for transformer-based retrieval, text-generation, and multimodal models in inference infrastructure.
  • Migrate GPU kernels to CuTe DSL and ensure portability to future hardware.
  • Build Rust-native serving runtime to handle growing traffic and reduce Python-related pains.
  • Profile, optimize, and interleave GPU kernels for performance across end-to-end pipelines.
  • Improve reliability with dashboards, alerts, and automated remediation for production incidents.

Skills

GPU programming
Performance tuning
Distributed systems
Cross-language development

Tools

CUDA
CuTe DSL
Triton
CUTLASS

Job description

Perplexity is building and running the inference engine behind every Perplexity query, deploying dozens of model architectures at scale with tight latency and cost budgets. We are seeking another engineer to join our team to advance transformer-based retrieval, text generation, and multimodal models in production.

You will work on a Rust-native serving runtime, CUDA/CuTe kernel migration, and ongoing performance optimizations to support growing traffic while maintaining reliability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Engineer – High-Performance GPU Systems
AI Inference Engineer – High-Performance GPU Systems

Perplexity • California (MO)

On-site
USD 120,000 - 170,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • California (MO)

On-site
USD 120,000 - 170,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI Inference Platform Engineer
AI Inference Platform Engineer

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 140,000 - 210,000
Equity
Medical, dental and vision insurance (
Flexible PTO including Winter Break
+4
AI Inference & HPC Engineer – Performance & APIs
AI Inference & HPC Engineer – Performance & APIs

Topaz Labs • Emeryville (CA)

On-site
USD 90,000 - 150,000
Full medical/dental/vision coverage
15 days PTO
5 personal days + holidays
+2
Senior AI Systems Engineer — Scalable Model Serving
Senior AI Systems Engineer — Scalable Model Serving

Apply • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
GenAI Performance Engineer: C++, Python, Rust Expert
GenAI Performance Engineer: C++, Python, Rust Expert

Obsidian • New York (NY)

On-site
USD 110,000 - 170,000
AI Inference Performance & Scale Engineer
AI Inference Performance & Scale Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 210,000
Benefits at a glance
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Staff Software Engineer — AI Acceleration & Infrastructure
Staff Software Engineer — AI Acceleration & Infrastructure

Perplexity AI Inc. • New York (NY)

Hybrid
USD 120,000 - 180,000