Inference Runtime Performance Engineer — GPU Kernels

iFrame Corporation

San Francisco (CA)

Remote

USD 220,000 - 360,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

iFrame Corporation is hiring for a role on the runtime team to own end-to-end performance for multiple model families on our managed inference runtime. You will write fused CUDA/Triton kernels and work across tokenizer through KV cache and decoding, targeting H100/H200 and other accelerators.

With 5+ years in systems-level performance and a track record of speed-ups, you will design long-context primitives and publish external write-ups quarterly while sharing on-call duties with SRE and

Qualifications

  • Five-plus years of systems-level performance work, with GPU kernel experience.
  • Strong CUDA and at least one of Triton, CUTLASS, or HIP.
  • Ability to read PTX and SASS when profiling demands it.
  • Track record of measurable speed-ups in production or open source.

Responsibilities

  • Profile real customer workloads with Nsight and internal tracer to turn bottlenecks into benchmarks.
  • Write fused, attention-aware kernels in CUDA and Triton (and HIP) for H100, H200, B200, B300.
  • Own one model family end-to-end from tokenizer to KV cache and decode paths on all accelerators.
  • Design and ship next round of long-context primitives: paged KV, ring attention, sliding-window cache eviction.
  • Co-author external write-ups each quarter: blog, paper, or kernel release.
  • Carry runtime on-call rotation alongside SRE and customer engineering ~one week every six weeks.

Skills

CUDA
Triton
HIP
Kernels
Transformer internals

Tools

CUTLASS
PTX/SASS
NVIDIA tools

Job description

iFrame Corporation is hiring for a role on the runtime team to own end-to-end performance for multiple model families on our managed inference runtime. You will write fused CUDA/Triton kernels and work across tokenizer through KV cache and decoding, targeting H100/H200 and other accelerators.

With 5+ years in systems-level performance and a track record of speed-ups, you will design long-context primitives and publish external write-ups quarterly while sharing on-call duties with SRE and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • United States

On-site
USD 182,000 - 242,000
Medical Insurance
401(k) Matching
Paid Parental Leave
+3
Senior GPU Kernel Engineer: Inference Performance
Senior GPU Kernel Engineer: Inference Performance

Neura Market • Sunnyvale (CA), Northern (KY)

Hybrid
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Tuition Reimbursement
+2
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
Senior GPU Kernel Engineer for High-Performance AI Inference
Senior GPU Kernel Engineer for High-Performance AI Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
GPU Kernel Engineer — Fast ML Training & Inference
GPU Kernel Engineer — Fast ML Training & Inference

Tilde Research • Palo Alto (CA)

On-site
USD 150,000 - 260,000
GPU Kernel Architect for High-Performance AI Inference
GPU Kernel Architect for High-Performance AI Inference

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 120,000 - 180,000
Senior GPU Kernel Engineer: Inference Throughput
Senior GPU Kernel Engineer: Inference Throughput

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with employer match
+3
Senior GPU Kernel Architect & Optimizer
Senior GPU Kernel Architect & Optimizer

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical–dental–vision insurance
401(k) with employer match
Flexible PTO
+3