Performance Engineer: Optimize GPU Inference at Scale

TechTwitter

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Morph is hiring a performance engineer to make the entire system faster, cheaper, and more reliable. You will work directly with the founders on problems that determine how efficiently frontier-scale models can be served, in a small team with enormous compute and immediate production impact.

You’ll trace latency and throughput from the API layer down to individual kernels, optimize batching, routing, quantization, and distributed execution, and build benchmarks and observability to make

Qualifications

  • Experience optimizing complex production systems.
  • Ability to turn profiling data into engineering decisions.
  • Strong Python and GPU performance knowledge.

Responsibilities

  • Find the gap between theoretical hardware performance and production performance.
  • Trace latency and throughput regressions from the API layer down to individual kernels.
  • Optimize batching, scheduling, routing, quantization, and distributed execution.
  • Build benchmarks and observability that make bottlenecks obvious.
  • Work with NVLink and RoCE.
  • Validate that every optimization preserves model quality and correctness.

Skills

Python
CuTEdsl
GPU performance
Profiling
Model serving

Tools

NVLink
RoCE

Job description

Morph is hiring a performance engineer to make the entire system faster, cheaper, and more reliable. You will work directly with the founders on problems that determine how efficiently frontier-scale models can be served, in a small team with enormous compute and immediate production impact.

You’ll trace latency and throughput from the API layer down to individual kernels, optimize batching, routing, quantization, and distributed execution, and build benchmarks and observability to make

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Morph is hiring Member of Technical Staff
Morph is hiring Member of Technical Staff

TechTwitter • San Francisco (CA), Northern (KY)

On-site
USD 180,000 - 240,000
Performance Engineer: GPU Kernel & Inference Optimize
Performance Engineer: GPU Kernel & Inference Optimize

WORLD LABS • San Francisco (CA)

On-site
USD 200,000 - 300,000
GPU Performance Engineer: Scale ML Inference & Systems
GPU Performance Engineer: Scale ML Inference & Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
Performance Engineer — GPU & System Optimization
Performance Engineer — GPU & System Optimization

PDT Partners • New York (NY)

Hybrid
USD 90,000 - 130,000
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Machine Learning Engineer, GPU Performance
Machine Learning Engineer, GPU Performance

Brahma Consulting Group • San Francisco (CA)

On-site
USD 150,000 - 210,000
Performance Engineer, Inference Engine - Flexible Hours
Performance Engineer, Inference Engine - Flexible Hours

Anthropic • San Francisco (CA), New York (NY)

On-site
USD 350,000 - 850,000
Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Performance Engineer
Performance Engineer

Tangier Partners • Boston (MA)

On-site
USD 150,000 - 210,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000