Research Engineer

Harnham

United States

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Harnham is looking for a technical expert to accelerate AI systems performance for next-generation models. The role involves optimising GPU training throughput, implementing advanced techniques, and designing scalable systems. Candidates should have over 4 years of relevant experience in performance optimisation, distributed systems, and strong GPU programming skills. This is a unique opportunity to work with cutting-edge AI technology and contribute significantly to real-time systems.

Qualifications

  • 4+ years of experience in systems engineering, ML infrastructure, or performance optimisation.
  • Strong experience with GPU programming.
  • Proven experience building scalable, fault-tolerant training systems.

Responsibilities

  • Optimize training throughput across large GPU clusters.
  • Implement mixed precision and memory-efficient techniques.
  • Design and scale distributed training systems.
  • Profile and optimise inference pipelines for real-time multimodal generation.

Skills

GPU programming (CUDA, Triton)
Performance optimisation
Distributed systems
ML framework internals (PyTorch, JAX)
Mixed or low‑precision techniques (FP8, INT8, BF16)

Job description

We're partnered with a well‑funded AI research company focused on building next‑generation multimodal models for media and interactive experiences. Their work spans cutting‑edge generative systems and is increasingly moving toward real‑time, interactive environments, pushing beyond static outputs into dynamic, AI‑driven applications.

This is a deeply technical, high‑impact role focused on making large‑scale AI systems faster, more efficient, and capable of running in real time. You'll work across the stack, from low‑level GPU kernels to distributed training systems, directly influencing what is computationally possible for next‑generation AI models.

What You'll Do
  • Optimize training throughput across large GPU clusters, improving efficiency and utilisation
  • Implement techniques such as mixed precision (FP8, BF16), memory‑efficient attention, and checkpointing
  • Design and scale distributed training systems (tensor parallelism, FSDP, multi‑node setups)
  • Profile and optimise inference pipelines for real‑time multimodal generation
  • Improve latency through CUDA graphing, KV cache optimisation, and operator fusion
  • Contribute across the stack, from kernel‑level optimisation to system‑level architecture
Requirements
  • 4+ years of experience in systems engineering, ML infrastructure, or performance optimisation
  • Strong experience with GPU programming (CUDA, Triton, or similar)
  • Experience with distributed systems and large‑scale training (NCCL, model parallelism)
  • Familiarity with ML framework internals such as PyTorch or JAX
  • Experience with mixed or low‑precision techniques (FP8, INT8, BF16)
  • Proven experience building and operating scalable, fault‑tolerant training systems
  • Strong interest in pushing the limits of performance for cutting‑edge AI systems
Nice to Have
  • Experience with compiler optimisations or model compilation (e.g., PyTorch compile)
  • Background working on large multimodal or generative models
  • Exposure to real‑time inference systems

If you're interested in working on the systems that enable next‑generation AI models to train faster and run in real time, this is a rare opportunity to operate at the cutting edge of research and infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Research Engineer, Data
Research Engineer, Data

Harnham • California (MO)

On-site
USD 100,000 - 130,000
Research Engineer
Research Engineer

Mind Robotics • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research Engineer, AI Models
Research Engineer, AI Models

EnCharge AI • Germany (OH)

On-site
USD 132,000 - 176,000
Research Engineer [ Performance Engineering ]
Research Engineer [ Performance Engineering ]

Metamorphic • Palo Alto (CA)

On-site
USD 200,000 - 280,000
Visa sponsorship
Competitive compensation
Mentorship and career development
+1
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Research Engineer, ML Infrastructure
Research Engineer, ML Infrastructure

cognition • San Francisco (CA)

On-site
USD 180,000 - 250,000
Research Engineer [ Distributed Training ]
Research Engineer [ Distributed Training ]

Metamorphic • Palo Alto (CA)

On-site
USD 200,000 - 280,000
Visa sponsorship
Competitive compensation
Equity package
+1
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1