Turn this role into an interview — a resume and cover letter built around what this employer wants.
Get past ATS filters
Job summary
Harnham is looking for a technical expert to accelerate AI systems performance for next-generation models. The role involves optimising GPU training throughput, implementing advanced techniques, and designing scalable systems. Candidates should have over 4 years of relevant experience in performance optimisation, distributed systems, and strong GPU programming skills. This is a unique opportunity to work with cutting-edge AI technology and contribute significantly to real-time systems.
Qualifications
4+ years of experience in systems engineering, ML infrastructure, or performance optimisation.
Strong experience with GPU programming.
Proven experience building scalable, fault-tolerant training systems.
Responsibilities
Optimize training throughput across large GPU clusters.
Implement mixed precision and memory-efficient techniques.
Design and scale distributed training systems.
Profile and optimise inference pipelines for real-time multimodal generation.
Skills
GPU programming (CUDA, Triton)
Performance optimisation
Distributed systems
ML framework internals (PyTorch, JAX)
Mixed or low‑precision techniques (FP8, INT8, BF16)
Job description
Harnham is looking for a technical expert to accelerate AI systems performance for next-generation models. The role involves optimising GPU training throughput, implementing advanced techniques, and designing scalable systems. Candidates should have over 4 years of relevant experience in performance optimisation, distributed systems, and strong GPU programming skills. This is a unique opportunity to work with cutting-edge AI technology and contribute significantly to real-time systems.