Research Engineer, GPU Performance

Harnham

California (MO)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Harnham is seeking a systems engineer to optimize large-scale AI systems for real-time performance. You will work on cutting-edge multimodal models, improving training efficiency and system architecture.

This role requires extensive experience in GPU programming and distributed systems, where you will directly influence AI advancements. Join Harnham to contribute to the future of interactive AI systems.

Qualifications

  • 4+ years of experience in systems engineering, ML infrastructure, or performance optimization.
  • Strong experience with GPU programming (CUDA, Triton, or similar).
  • Experience with distributed systems and large-scale training (NCCL, model parallelism).

Responsibilities

  • Optimize training throughput across large GPU clusters.
  • Implement techniques such as mixed precision, memory-efficient attention, and checkpointing.
  • Profile and optimize inference pipelines for real-time multimodal generation.

Skills

GPU programming
Performance optimization
Distributed systems
Machine Learning frameworks
Low-precision techniques

Job description

We’re partnered with a well-funded AI research company focused on building next-generation multimodal models for media and interactive experiences. Their work spans cutting-edge generative systems and is increasingly moving toward real-time, interactive environments, pushing beyond static outputs into dynamic, AI-driven applications.

This is a deeply technical, high-impact role focused on making large-scale AI systems faster, more efficient, and capable of running in real time. You’ll work across the stack, from low-level GPU kernels to distributed training systems, directly influencing what is computationally possible for next-generation AI models.

What You’ll Do
  • Optimize training throughput across large GPU clusters, improving efficiency and utilization
  • Implement techniques such as mixed precision (FP8, BF16), memory-efficient attention, and checkpointing
  • Design and scale distributed training systems (tensor parallelism, FSDP, multi-node setups)
  • Profile and optimize inference pipelines for real-time multimodal generation
  • Improve latency through CUDA graphing, KV cache optimization, and operator fusion
  • Contribute across the stack, from kernel-level optimization to system-level architecture
Requirements
  • 4+ years of experience in systems engineering, ML infrastructure, or performance optimization
  • Strong experience with GPU programming (CUDA, Triton, or similar)
  • Experience with distributed systems and large-scale training (NCCL, model parallelism)
  • Familiarity with ML framework internals such as PyTorch or JAX
  • Experience with mixed or low-precision techniques (FP8, INT8, BF16)
  • Proven experience building and operating scalable, fault-tolerant training systems
  • Strong interest in pushing the limits of performance for cutting-edge AI systems
Nice to Have
  • Experience with compiler optimizations or model compilation (e.g., PyTorch compile)
  • Background working on large multimodal or generative models
  • Exposure to real-time inference systems

If you're interested in working on the systems that enable next-generation AI models to train faster and run in real time, this is a rare opportunity to operate at the cutting edge of research and infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer
Research Engineer

Harnham • United States

On-site
USD 120,000 - 150,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
Senior AI Performance Engineer
Senior AI Performance Engineer

Brillfy Technology Inc • United States

On-site
USD 150,000 - 210,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Software Engineer, CUDA Deep Learning Systems
Software Engineer, CUDA Deep Learning Systems

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Member of Technical Staff, Performance Optimization
Member of Technical Staff, Performance Optimization

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000
Competitive compensation
Inclusive environment
Ownership & impact
AI Inference Performance Engineer - New College Grad 2026
AI Inference Performance Engineer - New College Grad 2026

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
GPU Performance / Kernel Engineer
GPU Performance / Kernel Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan
Annual bonus
+1
Senior AI Performance and Efficiency Engineer
Senior AI Performance and Efficiency Engineer

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Competitive salaries
Comprehensive benefits package
Equity eligibility