Get more replies from employers
Send a job-specific resume in minutes.
Our client is a rapidly growing organization at the forefront of the AI revolution, specializing in providing high-performance computing infrastructure to run heavy LLM models and AI products. They operate a global network of data centers with capacity specifically designed and tailored for extreme-scale computational workloads.
We are seeking an experienced HPC Engineer to join a dedicated high-performance computing optimization team. This team sits at the intersection of R&D, hardware engineering, and distributed systems, focusing on maximizing computational throughput and efficiency rather than traditional system administration.
In this role you will focus on optimizing the performance of large-scale GPU clusters, targeting latency reduction, computational efficiency, and enhanced parallel processing capabilities. Working with InfiniBand networks and high-performance computing infrastructure, you will collaborate with cross-functional teams to deliver scalable HPC solutions for client needs. The role requires balancing operational optimization and troubleshooting (50%) with HPC architecture design and performance tuning projects (50%). You will maintain and optimize distributed computing systems, managing over 30,000 GPUs across 10+ InfiniBand networks, while ensuring the optimal performance of global HPC infrastructure and driving continuous computational improvements.
Salary: up to 160k + 25% bonus (200k OTE)