Member of Technical Staff - GPU Performance Engineer

Liquid AI

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive base salary with equity
100% medical, dental, and vision premiums
401(k) matching up to 4%
Unlimited PTO
Company-wide Refill Days

Job summary

Liquid AI is seeking a talented individual to design and optimize custom CUDA kernels for high-performance AI systems. Located in San Francisco or Boston, the ideal candidate will have strong C/C++ skills and experience in GPU architecture optimization. Responsibilities include profiling workflows, integrating kernels with PyTorch, and collaborating with research to deliver performance improvements. Competitive salary, health coverage, and unlimited PTO are offered.

Qualifications

  • Experience in writing high-performance GPU kernels.
  • Strong understanding of performance methodology in GPU architecture.
  • Proficiency in C/C++ and experience with low-level profiling tools.

Responsibilities

  • Design and ship custom CUDA kernels.
  • Profile and optimize training and inference workflows.
  • Integrate kernels into PyTorch pipelines.

Skills

CUDA kernel development
GPU architecture performance
Low-level profiling
C/C++ programming

Tools

Nsight Systems
Nsight Compute
PyTorch

Job description

About Liquid AI

Spun out of MIT CSAIL, we build general‑purpose AI systems that run efficiently across deployment targets, from data center accelerators to on‑device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there.

The Opportunity

Our models and workflows require performance work that generic frameworks don’t solve. You’ll design and ship custom CUDA kernels, profile at the hardware level, and integrate research ideas into production code that delivers measurable speedups in real pipelines (training, post‑training, and inference). Our team is small, fast‑moving, and high‑ownership. We're looking for someone who finds joy in memory hierarchies, tensor cores, and profiler output.

While San Francisco and Boston are preferred, we are open to other locations.

What We’re Looking For

We need someone who:

  • Works profiler‑first: You use tools like Nsight Systems / Nsight Compute to find bottlenecks, validate hypotheses, and iterate until improvements show up in end‑to‑end benchmarks.
  • Bridges theory and practice: You can translate ideas from papers into implementations that are robust, testable, and performant.
  • Executes independently: Given an ambiguous bottleneck, you can drive from profiling to kernel/integration changes to benchmarked results to maintained ownership.
  • Cares about the details: Memory hierarchy, occupancy, launch configurations, tensor core utilization, bandwidth vs compute limits.
The Work
  • Write high‑performance GPU kernels for our novel model architectures.
  • Integrate kernels into PyTorch pipelines (custom ops, extensions, dispatch, benchmarking).
  • Profile and optimize training and inference workflows to eliminate bottlenecks.
  • Build correctness tests and numerics checks.
  • Build/maintain performance benchmarks and guardrails to prevent regressions.
  • Collaborate closely with researchers to turn promising ideas into shipped speedups.
Desired Experience
  • Authored custom CUDA kernels (not only calling cuDNN/cuBLAS).
  • Strong understanding of GPU architecture and performance: memory hierarchy, warps, shared memory/register pressure, bandwidth vs compute limits.
  • Proficiency with low‑level profiling (Nsight Systems/Compute) and performance methodology.
  • Strong C/C++ skills.
Nice‑to‑have
  • CUTLASS experience and tensor core utilization strategies.
  • Triton kernel experience and/or PyTorch custom op integration.
  • Experience building benchmark harnesses and perf regression tests.
What Success Looks Like (Year One)
  • Measurable improvement on at least one critical end‑to‑end pipeline (throughput and/or latency), validated by repeatable benchmarks.
  • At least one research‑driven technique shipped as a production kernel and maintained over time.
  • Performance regressions are detectable early via benchmarks/guardrails, not discovered late.
What We Offer
  • Unique challenges: Our architectural innovations and efficiency requirements offer unique optimization challenges. High ownership from day one.
  • Compensation: Competitive base salary with equity in a unicorn‑stage company.
  • Health: We pay 100% of medical, dental, and vision premiums for employees and dependents.
  • Financial: 401(k) matching up to 4% of base pay.
  • Time Off: Unlimited PTO plus company‑wide Ref­ill Days throughout the year.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding GPU Kernel Engineer
Founding GPU Kernel Engineer

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Comprehensive benefits package
Founding GPU Kernel Engineer
Founding GPU Kernel Engineer

SF Tensor • San Francisco (CA)

On-site
USD 285,000 - 315,000
Member of Technical Staff - GPU Infrastructure Engineer
Member of Technical Staff - GPU Infrastructure Engineer

Liquid AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
Health benefits
401k matching
+2
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 150,000 - 350,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
Member of Technical Staff, Performance Optimization
Member of Technical Staff, Performance Optimization

Fireworks AI • San Mateo (CA)

On-site
USD 180,000 - 260,000
Senior Software Engineer - GPU Kernel Authoring & Optimization
Senior Software Engineer - GPU Kernel Authoring & Optimization

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
GPU Performance / Kernel Engineer
GPU Performance / Kernel Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan
Annual bonus
+1
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000