Remote GPU Cloud Platform Engineer: Scale AI Compute
Yotta Labs
United States
Remote
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Flexible remote work environment
Innovative team collaboration
Cutting-edge technology challenges
Job summary
A pioneering AI infrastructure company is seeking a GPU Cloud Platform Engineer to design and operate large-scale GPU clusters. This remote position aims to ensure high availability and performance of containerized AI workloads across cloud environments. The ideal candidate should have extensive experience with Kubernetes, cloud-native development, and strong skills in monitoring tools. They will work in an innovative team dedicated to revolutionizing AI infrastructure.
Qualifications
3+ years of experience in system engineering or DevOps.
5+ years of experience in cloud-native development or AI engineering.
Hands-on experience with multi-cluster deployment, upgrade, and scaling.
Responsibilities
Build and operate large-scale, high-performance GPU clusters.
Conduct performance testing of multi-node GPU clusters.
Deploy large models across multi-cluster environments.
Skills
Kubernetes multi-cluster management
Cloud-native development
Docker
Monitoring tools (Prometheus, Grafana)
Communication protocols (IB, RoCE)
Team collaboration
Education
Bachelor's degree in Computer Science, Software Engineering, or related fields
Tools
Helm
kubectl
AWS
GCP
Azure
Job description
A pioneering AI infrastructure company is seeking a GPU Cloud Platform Engineer to design and operate large-scale GPU clusters. This remote position aims to ensure high availability and performance of containerized AI workloads across cloud environments. The ideal candidate should have extensive experience with Kubernetes, cloud-native development, and strong skills in monitoring tools. They will work in an innovative team dedicated to revolutionizing AI infrastructure.