A leading AI technology firm in San Francisco is seeking an AI Infra Engineer to enhance their infrastructure. The successful candidate will design and maintain Kubernetes clusters and manage Slurm for distributed training. Important skills include extensive experience in Kubernetes and Slurm, and a strong foundation in Python and C++. Join a dynamic team aiming at advancements in AI and ML infrastructure.
Qualifications
Expert-level Kubernetes administration and YAML configuration management.
Proficiency with Slurm job scheduling and resource management.
Hands-on experience with ML frameworks like PyTorch in distributed training contexts.
Responsibilities
Design, deploy, and maintain scalable Kubernetes clusters for AI workloads.
Manage and optimize Slurm-based HPC environments for distributed training.
Develop robust APIs for training pipelines and inference services.
Skills
Kubernetes administration
Slurm workload management
Python programming
C++ programming
ML frameworks (PyTorch)
Distributed systems architecture
API development
Debugging and monitoring
Tools
Kubernetes
Slurm
Terraform
Ansible
Job description
A leading AI technology firm in San Francisco is seeking an AI Infra Engineer to enhance their infrastructure. The successful candidate will design and maintain Kubernetes clusters and manage Slurm for distributed training. Important skills include extensive experience in Kubernetes and Slurm, and a strong foundation in Python and C++. Join a dynamic team aiming at advancements in AI and ML infrastructure.