Kraken is seeking an experienced individual to join its AI Compute and Infrastructure team in Romblon, España. The role involves managing GPU and accelerator clusters, optimizing AI workloads, and enhancing the compute infrastructure to support AI initiatives. Candidates should have over 5 years of experience in infrastructure engineering, particularly with GPU systems and ML infrastructure. The position offers the opportunity to impact Kraken's AI strategy significantly by ensuring efficient, reliable, and robust infrastructure for AI operations.
Qualifications
5+ years of infrastructure engineering experience, focusing on GPU compute.
Hands-on experience operating GPU clusters in production environments.
Strong systems engineering fundamentals across Linux, networking, and containers.
Responsibilities
Own and operate GPU and accelerator clusters for AI workloads.
Design infrastructure for local model training on GPUs.
Build observability tools for GPU usage and system performance.
Skills
Infrastructure engineering experience
Hands-on experience with GPU clusters
Strong systems engineering fundamentals
Experience with ML serving frameworks
Proficiency in Python
Performance tradeoffs understanding
Optimize compute costs
Observability systems building
Working in high-stakes environments
Clear communication skills
Tools
Kubernetes
Linux
CUDA
Rust
C++
Go
Job description
Kraken is seeking an experienced individual to join its AI Compute and Infrastructure team in Romblon, España. The role involves managing GPU and accelerator clusters, optimizing AI workloads, and enhancing the compute infrastructure to support AI initiatives. Candidates should have over 5 years of experience in infrastructure engineering, particularly with GPU systems and ML infrastructure. The position offers the opportunity to impact Kraken's AI strategy significantly by ensuring efficient, reliable, and robust infrastructure for AI operations.