Research Engineer, GPU Performance

Harnham

California (MO)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Harnham is seeking a systems engineer to optimize large-scale AI systems for real-time performance. You will work on cutting-edge multimodal models, improving training efficiency and system architecture.

This role requires extensive experience in GPU programming and distributed systems, where you will directly influence AI advancements. Join Harnham to contribute to the future of interactive AI systems.

Qualifications

  • 4+ years of experience in systems engineering, ML infrastructure, or performance optimization.
  • Strong experience with GPU programming (CUDA, Triton, or similar).
  • Experience with distributed systems and large-scale training (NCCL, model parallelism).

Responsibilities

  • Optimize training throughput across large GPU clusters.
  • Implement techniques such as mixed precision, memory-efficient attention, and checkpointing.
  • Profile and optimize inference pipelines for real-time multimodal generation.

Skills

GPU programming
Performance optimization
Distributed systems
Machine Learning frameworks
Low-precision techniques

Job description

We’re partnered with a well-funded AI research company focused on building next-generation multimodal models for media and interactive experiences. Their work spans cutting-edge generative systems and is increasingly moving toward real-time, interactive environments, pushing beyond static outputs into dynamic, AI-driven applications.

This is a deeply technical, high-impact role focused on making large-scale AI systems faster, more efficient, and capable of running in real time. You’ll work across the stack, from low-level GPU kernels to distributed training systems, directly influencing what is computationally possible for next-generation AI models.

What You’ll Do
  • Optimize training throughput across large GPU clusters, improving efficiency and utilization
  • Implement techniques such as mixed precision (FP8, BF16), memory-efficient attention, and checkpointing
  • Design and scale distributed training systems (tensor parallelism, FSDP, multi-node setups)
  • Profile and optimize inference pipelines for real-time multimodal generation
  • Improve latency through CUDA graphing, KV cache optimization, and operator fusion
  • Contribute across the stack, from kernel-level optimization to system-level architecture
Requirements
  • 4+ years of experience in systems engineering, ML infrastructure, or performance optimization
  • Strong experience with GPU programming (CUDA, Triton, or similar)
  • Experience with distributed systems and large-scale training (NCCL, model parallelism)
  • Familiarity with ML framework internals such as PyTorch or JAX
  • Experience with mixed or low-precision techniques (FP8, INT8, BF16)
  • Proven experience building and operating scalable, fault-tolerant training systems
  • Strong interest in pushing the limits of performance for cutting-edge AI systems
Nice to Have
  • Experience with compiler optimizations or model compilation (e.g., PyTorch compile)
  • Background working on large multimodal or generative models
  • Exposure to real-time inference systems

If you're interested in working on the systems that enable next-generation AI models to train faster and run in real time, this is a rare opportunity to operate at the cutting edge of research and infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer
Research Engineer

Harnham • United States

On-site
USD 120,000 - 150,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
GPU Optimization Engineer
GPU Optimization Engineer

Vincere • San Francisco (CA)

On-site
USD 230,000 - 300,000
Machine Learning Engineer, GPU Performance
Machine Learning Engineer, GPU Performance

Brahma Consulting Group • San Francisco (CA)

On-site
USD 150,000 - 210,000
Software Engineer - GPU Kernel
Software Engineer - GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Performance Engineer
Performance Engineer

Tangier Partners • Boston (MA)

On-site
USD 150,000 - 210,000
Senior AI Performance Engineer
Senior AI Performance Engineer

Brillfy Technology Inc • United States

On-site
USD 150,000 - 210,000
AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • Spain (TX)

Remote
USD 135,000 - 203,000
Remote-first
International team
Cutting-edge AI
+1
Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Equity
Visa sponsorship
Relocation assistance
+1
Senior AI Performance and Efficiency Engineer
Senior AI Performance and Efficiency Engineer

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits