Real-Time GPU Optimization Engineer - Inference

techire ai

San Francisco (CA)

On-site

USD 230,000 - 300,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

techire ai is seeking a GPU Optimisation Engineer for real-time inference in production AI workloads. The role focuses on pushing GPU performance to sub-50ms latency under high concurrency, close to the metal across kernel and runtime layers.

You will profile and optimize large generative models, write custom CUDA/Triton kernels, and collaborate with research to deliver production-ready inference at scale. SF relocation/visa sponsorship available.

Qualifications

  • Strong experience with CUDA and/or Triton.
  • Deep understanding of GPU execution (memory hierarchy, scheduling, occupancy, concurrency).
  • Experience optimising inference latency and throughput for large generative models.
  • Familiarity with attention kernels, decoding paths, or LLM‑style runtimes.
  • Comfort profiling with low‑level GPU tooling.

Responsibilities

  • Profiling GPU bottlenecks across memory bandwidth, kernel fusion, quantisation, and scheduling.
  • Writing and tuning custom CUDA / Triton kernels for performance‑critical paths.
  • Improving attention, decoding, and KV cache efficiency in inference runtimes.
  • Modifying and extending vLLM‑style systems to better suit real‑time workloads.
  • Optimising models to fit GPU memory constraints without degrading output quality.
  • Benchmarking across NVIDIA GPUs (with exposure to AMD and other accelerators over time).
  • Partnering directly with research to turn new model ideas into fast, production‑ready inference.

Skills

CUDA
Triton
GPU architecture
Profiling tooling
Low-level optimization

Job description

techire ai is seeking a GPU Optimisation Engineer for real-time inference in production AI workloads. The role focuses on pushing GPU performance to sub-50ms latency under high concurrency, close to the metal across kernel and runtime layers.

You will profile and optimize large generative models, write custom CUDA/Triton kernels, and collaborate with research to deliver production-ready inference at scale. SF relocation/visa sponsorship available.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Optimization Engineer
GPU Optimization Engineer

techire ai • San Francisco (CA)

On-site
USD 230,000 - 300,000
Senior GPU Inference Engine Engineer
Senior GPU Inference Engine Engineer

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
+2
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior GPU Inference Engineer for Real-Time AI
Senior GPU Inference Engineer for Real-Time AI

Cerebras • United States

On-site
USD 150,000 - 210,000
Staff GPU Inference Engineer — Real-Time AI Systems
Staff GPU Inference Engineer — Real-Time AI Systems

Cerebras • United States

Remote
USD 150,000 - 230,000
Senior Inference Performance Engineer - GPU & CUDA
Senior Inference Performance Engineer - GPU & CUDA

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,000 - 239,000
Equity compensation
Remote work
Software Engineer – AI Inference Engine
Software Engineer – AI Inference Engine

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
+2