Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs, Inc.

San Francisco (CA)

On-site

USD 150,000 - 350,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Gimlet Labs, Inc. is seeking a Member of Technical Staff focused on optimizing GPU and accelerator kernels for AI workloads. This role involves analyzing and tuning performance across diverse execution platforms and collaborating with compilers to ensure clean integration. Ideal candidates possess a strong foundation in software engineering and experience with performance-critical systems. Compensation ranges from $150K to $350K, reflecting the market demand for expertise in GPU execution and performance optimization.

Qualifications

  • Strong software engineering fundamentals.
  • Experience working on performance-critical systems close to hardware.
  • Comfort reasoning about low-level execution behavior, memory hierarchies, and performance tradeoffs.

Responsibilities

  • Design, implement, and optimize GPU and accelerator kernels for AI workloads.
  • Analyze and tune performance across the GPU execution stack.
  • Work with compilers and runtimes to ensure kernels integrate cleanly.

Skills

Software engineering fundamentals
Performance-critical systems
Low-level execution behavior
Memory hierarchies
Performance tradeoffs

Tools

CUDA
Triton
CUTLASS
Profiling tools

Job description

About Us

Gimlet Labs is building the first heterogeneous neocloud for AI workloads.

As AI systems scale, the industry is hitting fundamental limits in power, capacity, and cost with today’s homogeneous, vertically integrated infrastructure. Gimlet addresses this by decoupling AI workloads from the underlying hardware. Our platform intelligently partitions workloads into components and orchestrates each component to hardware that best fits its performance and efficiency needs. This approach enables heterogeneous systems across multi-vendor and multi-generation hardware, including the latest emerging accelerators. These systems unlock step‑function improvements in performance and cost efficiency at scale.

On top of this foundation, Gimlet is building a production‑grade neocloud for agentic workloads. Customers use Gimlet to deploy and manage their workloads through stable, production‑ready APIs, without having to reason about hardware selection, placement, or low‑level performance optimization.

Gimlet works with foundation labs, hyperscalers, and AI native companies to power real production workloads built to scale to gigawatt‑class AI datacenters.

Mission

Gimlet Labs is seeking a Member of Technical Staff focused on kernels and GPU performance. In this role, you will work close to accelerators and execution hardware to extract maximum performance from AI workloads across diverse and rapidly evolving platforms. You will analyze low‑level execution behavior, design and optimize kernels, and ensure performance is reliable across both established and emerging hardware.

This role is ideal for engineers who enjoy deep performance work, reasoning about hardware tradeoffs, and turning theoretical peak performance into real‑world results.

Responsibilities
  • Design, implement, and optimize GPU and accelerator kernels for AI workloads
  • Analyze and tune performance across the GPU execution stack, including memory access patterns, synchronization, and instruction scheduling
  • Work with compilers and runtimes to ensure kernels integrate cleanly and perform well in end‑to‑end systems
  • Bring up and optimize execution on new or emerging accelerators
  • Profile, benchmark, and debug performance issues across kernels, runtimes, and hardware
  • Ensure performance optimizations are robust, correct, and production‑ready at scale
Qualifications
  • Strong software engineering fundamentals
  • Experience working on performance‑critical systems close to hardware
  • Comfort reasoning about low‑level execution behavior, memory hierarchies, and performance tradeoffs
Preferred Qualifications
  • Experience with CUDA, Triton, CUTLASS, or other accelerator programming models
  • Deep understanding of GPU execution models (warps/wavefronts, blocks, grids)
  • Experience optimizing memory access patterns (coalescing, shared memory, cache behavior)
  • Familiarity with occupancy, latency hiding, and instruction‑level parallelism
  • Experience using profiling and performance analysis tools
  • Familiarity with multi‑GPU or distributed execution is a plus

Compensation Range: $150K - $350K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff GPU Kernel & Performance Engineer
Staff GPU Kernel & Performance Engineer

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 150,000 - 350,000
Founding GPU Kernel Engineer
Founding GPU Kernel Engineer

SF Tensor • San Francisco (CA)

On-site
USD 285,000 - 315,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - GPU Performance Engineer
Member of Technical Staff - GPU Performance Engineer

Liquid AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive base salary with equity
100% medical, dental, and vision premiums
401(k) matching up to 4%
+2
Member of Technical Staff - Compilers
Member of Technical Staff - Compilers

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 150,000
Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff GPU Performance / Kernel Engineer
Staff GPU Performance / Kernel Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 140,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+2
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000