ML Training Platform Engineer | Multi-Cloud & Decentralized

Pluralis Research

San Francisco (CA)

Remote

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A decentralized AI platform company in the United States is seeking an experienced ML Training Platform Engineer to design and build robust infrastructure for ML training. The ideal candidate has over 5 years in infrastructure and platform engineering, with expertise in multi-cloud deployments and distributed systems. Responsibilities include architecting fault-tolerant infrastructure and simulating real-world network conditions. This role is essential for enabling large-scale, collaborative AI development.

Qualifications

  • 5+ years of experience with deep experience in infrastructure and platform engineering.
  • Production experience managing multi-cloud deployments.
  • Strong Python engineering skills with a focus on reliability and observability.

Responsibilities

  • Design resource management systems across AWS, GCP, and Azure.
  • Architect fault-tolerant infrastructure for distributed ML.
  • Build systems simulating real-world network conditions for ML training.

Skills

Infrastructure & Platform Engineering
Distributed Systems & ML Infrastructure
Systems Programming & Reliability

Tools

Pulumi
Terraform
Docker
Kubernetes
Prometheus
Grafana

Job description

A decentralized AI platform company in the United States is seeking an experienced ML Training Platform Engineer to design and build robust infrastructure for ML training. The ideal candidate has over 5 years in infrastructure and platform engineering, with expertise in multi-cloud deployments and distributed systems. Responsibilities include architecting fault-tolerant infrastructure and simulating real-world network conditions. This role is essential for enabling large-scale, collaborative AI development.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Platform Engineer: Scale AI Infra, Deploy & Optimize
ML Platform Engineer: Scale AI Infra, Deploy & Optimize

United States Digital Space LLC • United States

Remote
USD 120,000 - 180,000
ML Systems Engineer: Cloud‑Scale Training Infra
ML Systems Engineer: Cloud‑Scale Training Infra

Basis Research Institute • New York (NY)

On-site
USD 150,000 - 200,000
Competitive salary
Collaborative work environment
In-person events
Staff ML Engineer - Build Scalable Multimodal AI Platform
Staff ML Engineer - Build Scalable Multimodal AI Platform

kadence • San Francisco (CA)

On-site
USD 180,000 - 250,000
ML Platform Engineer
ML Platform Engineer

Jobtailor • North Carolina

On-site
USD 100,000 - 180,000
ML Platform Engineer: Scalable AI Infra & Reliability
ML Platform Engineer: Scalable AI Infra & Reliability

Bjak • Town of Poland (NY)

On-site
USD 140,000 - 220,000
Lead Platform ML Engineer — Scale AI Infra & Cloud
Lead Platform ML Engineer — Scale AI Infra & Cloud

Gen • New York (NY)

On-site
USD 197,000 - 219,000
ML Training Platform Architect for Large-Scale GPU Clusters
ML Training Platform Architect for Large-Scale GPU Clusters

Scale AI • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 216,000 - 270,000
Comprehensive health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
Senior ML Platform Engineer for Gen AI & MLOps
Senior ML Platform Engineer for Gen AI & MLOps

Bloomberg L.P. • New York (NY)

On-site
USD 160,000 - 240,000
401(k) + match
Paid time off
Medical and dental benefits
+1
ML Platform Engineer: Scalable AI Infrastructure
ML Platform Engineer: Scalable AI Infrastructure

Bjak • Germany (OH)

On-site
USD 120,000 - 180,000
Senior ML Platform Engineer — Training & Deployment
Senior ML Platform Engineer — Training & Deployment

Maxinsights Corporation • Santa Clara (CA)

On-site
USD 180,000 - 240,000