AI Inference Engineer (GPU/Rust/CUDA)

Perplexity

New York (NY)

On-site

USD 220,000 - 485,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Perplexity in New York is seeking an experienced engineer to join our team and own ML inference workloads, building a Rust/Python-based serving runtime and CUDA kernel work that scales to production. You will work with transformer-based models, enabling efficient weight loading, scheduling, and KV-cache management in our in-house infrastructure.

You will read research papers, implement kernels, and diagnose production incidents in a fast-moving environment, collaborating across languages and

Qualifications

  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).

Responsibilities

  • New models support for transformer-based retrieval, text-generation, and multimodal models in the inference infrastructure.
  • Port CUDA kernels to CuTe DSL and migrate GPU kernels for portability.
  • Develop a Rust-native serving runtime to handle growing traffic and reduce Python pains.
  • Profile and optimize performance across network, batching, and GPU kernel interleaving.
  • Build dashboards, alerts, and automated remediation to catch regressions.

Skills

ML inference
High-performance systems
Distributed systems
Rust
Python
CUDA

Tools

PyTorch
JAX
TensorFlow

Job description

Perplexity in New York is seeking an experienced engineer to join our team and own ML inference workloads, building a Rust/Python-based serving runtime and CUDA kernel work that scales to production. You will work with transformer-based models, enabling efficient weight loading, scheduling, and KV-cache management in our in-house infrastructure.

You will read research papers, implement kernels, and diagnose production incidents in a fast-moving environment, collaborating across languages and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Inference Engineer: GPU Performance & Rust Stack
AI Inference Engineer: GPU Performance & Rust Stack

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI Inference Engineer — High-Performance GPU Systems
AI Inference Engineer — High-Performance GPU Systems

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 485,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 485,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • New York (NY)

On-site
USD 220,000 - 485,000
Rust ML Systems Engineer: Real-Time GPU Inference
Rust ML Systems Engineer: Real-Time GPU Inference

Archetype AI Inc. • San Mateo (CA)

On-site
USD 180,000 - 280,000
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Staff AI Systems Engineer – Model Serving & Infra
Staff AI Systems Engineer – Model Serving & Infra

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 405,000
Staff Inference Runtime Architect (Rust/Python)
Staff Inference Runtime Architect (Rust/Python)

Anthropic • New York (NY)

Hybrid
USD 405,000 - 485,000
Generous vacation
Parental leave
Flexible working hours
+1