Inference Runtime Performance Engineer — GPU Kernels

iFrame Corporation

San Francisco (CA)

Remote

USD 220,000 - 360,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

iFrame Corporation is hiring for a role on the runtime team to own end-to-end performance for multiple model families on our managed inference runtime. You will write fused CUDA/Triton kernels and work across tokenizer through KV cache and decoding, targeting H100/H200 and other accelerators.

With 5+ years in systems-level performance and a track record of speed-ups, you will design long-context primitives and publish external write-ups quarterly while sharing on-call duties with SRE and

Qualifications

  • Five-plus years of systems-level performance work, with GPU kernel experience.
  • Strong CUDA and at least one of Triton, CUTLASS, or HIP.
  • Ability to read PTX and SASS when profiling demands it.
  • Track record of measurable speed-ups in production or open source.

Responsibilities

  • Profile real customer workloads with Nsight and internal tracer to turn bottlenecks into benchmarks.
  • Write fused, attention-aware kernels in CUDA and Triton (and HIP) for H100, H200, B200, B300.
  • Own one model family end-to-end from tokenizer to KV cache and decode paths on all accelerators.
  • Design and ship next round of long-context primitives: paged KV, ring attention, sliding-window cache eviction.
  • Co-author external write-ups each quarter: blog, paper, or kernel release.
  • Carry runtime on-call rotation alongside SRE and customer engineering ~one week every six weeks.

Skills

CUDA
Triton
HIP
Kernels
Transformer internals

Tools

CUTLASS
PTX/SASS
NVIDIA tools

Job description

iFrame Corporation is hiring for a role on the runtime team to own end-to-end performance for multiple model families on our managed inference runtime. You will write fused CUDA/Triton kernels and work across tokenizer through KV cache and decoding, targeting H100/H200 and other accelerators.

With 5+ years in systems-level performance and a track record of speed-ups, you will design long-context primitives and publish external write-ups quarterly while sharing on-call duties with SRE and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Kernel Engineer: GPU Performance & Inference
Kernel Engineer: GPU Performance & Inference

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Kernel Engineer
Kernel Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Real-Time GPU Optimization Engineer - Inference
Real-Time GPU Optimization Engineer - Inference

techire ai • San Francisco (CA)

On-site
USD 230,000 - 300,000
Senior CUDA Kernel Engineer for High-Performance Inference
Senior CUDA Kernel Engineer for High-Performance Inference

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
+3
Performance Engineer: GPU Kernel & Inference Optimize
Performance Engineer: GPU Kernel & Inference Optimize

WORLD LABS • San Francisco (CA)

On-site
USD 200,000 - 300,000
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000