Member of Technical Staff (AI Inference Engineer)

Perplexity

New York (NY)

On-site

USD 220,000 - 485,000

Full time

4 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Perplexity in New York is seeking an experienced engineer to join our team and own ML inference workloads, building a Rust/Python-based serving runtime and CUDA kernel work that scales to production. You will work with transformer-based models, enabling efficient weight loading, scheduling, and KV-cache management in our in-house infrastructure.

You will read research papers, implement kernels, and diagnose production incidents in a fast-moving environment, collaborating across languages and

Qualifications

  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).

Responsibilities

  • New models support for transformer-based retrieval, text-generation, and multimodal models in the inference infrastructure.
  • Port CUDA kernels to CuTe DSL and migrate GPU kernels for portability.
  • Develop a Rust-native serving runtime to handle growing traffic and reduce Python pains.
  • Profile and optimize performance across network, batching, and GPU kernel interleaving.
  • Build dashboards, alerts, and automated remediation to catch regressions.

Skills

ML inference
High-performance systems
Distributed systems
Rust
Python
CUDA

Tools

PyTorch
JAX
TensorFlow

Job description

We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join us.

What you will work on
Examples Of Real Work The Team Does
  • New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.
  • GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.
  • Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
  • Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
  • Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents.
Who we're looking for
  • Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus.
  • You understand modern LLM architectures and are able to bring them up reliably in a production environment.
  • You've built and operated production distributed systems under real load - ideally performance-critical ones.
  • Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels.
  • You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday.
  • Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you.
Good if you touched any of
  • ML compilers and framework internals: PyTorch internals, torch.compile, custom operators.
  • Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism.
  • Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving.
  • Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis.
  • Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads.
Qualifications
  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).

Compensation Range: $220K - $485K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
AI Inference Engineer: GPU Performance & Rust Stack
AI Inference Engineer: GPU Performance & Rust Stack

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff (Performance Optimization)
Member of Technical Staff (Performance Optimization)

Fireworks AI • United States

On-site
USD 150,000 - 260,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Kindredventures • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1