ML Training Platform Engineer | Multi-Cloud & Decentralized
Pluralis Research
San Francisco (CA)
Remote
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A decentralized AI platform company in the United States is seeking an experienced ML Training Platform Engineer to design and build robust infrastructure for ML training. The ideal candidate has over 5 years in infrastructure and platform engineering, with expertise in multi-cloud deployments and distributed systems. Responsibilities include architecting fault-tolerant infrastructure and simulating real-world network conditions. This role is essential for enabling large-scale, collaborative AI development.
Qualifications
5+ years of experience with deep experience in infrastructure and platform engineering.
Production experience managing multi-cloud deployments.
Strong Python engineering skills with a focus on reliability and observability.
Responsibilities
Design resource management systems across AWS, GCP, and Azure.
Architect fault-tolerant infrastructure for distributed ML.
Build systems simulating real-world network conditions for ML training.
Skills
Infrastructure & Platform Engineering
Distributed Systems & ML Infrastructure
Systems Programming & Reliability
Tools
Pulumi
Terraform
Docker
Kubernetes
Prometheus
Grafana
Job description
A decentralized AI platform company in the United States is seeking an experienced ML Training Platform Engineer to design and build robust infrastructure for ML training. The ideal candidate has over 5 years in infrastructure and platform engineering, with expertise in multi-cloud deployments and distributed systems. Responsibilities include architecting fault-tolerant infrastructure and simulating real-world network conditions. This role is essential for enabling large-scale, collaborative AI development.