Comprehensive medical, dental, and vision insurance
Fully paid parental leave
Paid time off
Daily lunch and dinner provided
Job summary
A technology company in the UK is seeking a skilled team member to enhance their Compute Platform. The role focuses on managing a K8s-based platform, ensuring system health, and improving performance. Responsibilities include cluster management, designing effective monitoring strategies, and preparing infrastructure for future GPU deployments. Ideal candidates will have strong systems-level engineering abilities, cloud storage expertise, and deep knowledge of GPU hardware in a Kubernetes environment. This position offers top-tier compensation and a collaborative work environment.
Qualifications
Experience focusing on cluster-wide behavior and maintenance.
Proven skills in systems or GPU infrastructure.
Familiarity with NCCL and multi-GPU environments.
Responsibilities
Build and maintain tools for automatic remediation and capacity planning.
Design cluster management stack for large-scale workloads.
Implement cluster-wide monitoring and performance benchmarking.
Prepare infrastructure for next-generation GPU deployments.
Skills
Systems-level engineering experience
Strong coding ability
Deep GPU hardware knowledge
Alignment with a K8s-first architecture
Cloud storage expertise
Job description
A technology company in the UK is seeking a skilled team member to enhance their Compute Platform. The role focuses on managing a K8s-based platform, ensuring system health, and improving performance. Responsibilities include cluster management, designing effective monitoring strategies, and preparing infrastructure for future GPU deployments. Ideal candidates will have strong systems-level engineering abilities, cloud storage expertise, and deep knowledge of GPU hardware in a Kubernetes environment. This position offers top-tier compensation and a collaborative work environment.