Machine Learning Senior Platform & Infrastructure Engineer

WorkGenius Group

Los Angeles (CA)

On-site

USD 117,000 - 186,000

Full time

18 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

WorkGenius Group is seeking a Machine Learning Senior Platform & Infrastructure Engineer to build and operate scalable GPU clusters and ML training platforms in Los Angeles. You will design game-simulation infrastructure for parallel rollouts, data collection, training, and evaluation.

The role involves developing CI/CD pipelines, IaC, and internal tools across cloud environments, with focus on reliability, performance, and cost efficiency.

Qualifications

  • Bachelor's degree in Computer Science or related field, or equivalent practical experience, with 3+ years of software engineering experience.
  • Production experience operating distributed systems and high-scale infrastructure, including Kubernetes, AWS or GCP, IaC and CI/CD.
  • Experience with GPU infrastructure, including scheduling, multi-node orchestration, and optimization for long ML workloads.
  • Strong Python skills and understanding of networking, microservices, and distributed systems.
  • Familiarity with MLOps practices such as model versioning, pipeline orchestration, experiment tracking, and reproducible ML workflows.

Responsibilities

  • Build and operate Kubernetes, multi-node GPU clusters, networking infrastructure, and distributed ML training platforms.
  • Design and scale game simulation infrastructure for parallel rollouts, data collection, training, and evaluation.
  • Develop CI/CD pipelines, deployment automation, artifact management, infrastructure-as-code, and internal developer tools across cloud environments.
  • Improve reliability, scalability, performance, cost efficiency, observability, reproducibility, auditability, and SLO-based operations.
  • Support MLOps, security governance, incident response, root-cause remediation, and mentoring of platform engineering talent.

Skills

Kubernetes
GPU infrastructure
Distributed systems
MLOps
Python
AWS
GCP
Infrastructure as code
CI/CD
Distributed ML training

Education

Bachelor’s degree in Computer Science or related field

Tools

Terraform

Job description

Title: Machine Learning Senior Platform & Infrastructure Engineer

Industry: Gaming

Location: Los Angeles, CA

Duration: 12 months

Responsibilities
  • Build and operate Kubernetes, multi-node GPU clusters, networking infrastructure, and distributed ML training platforms.
  • Design and scale game simulation infrastructure for parallel rollouts, data collection, training, and evaluation.
  • Develop CI/CD pipelines, deployment automation, artifact management, infrastructure-as-code, and internal developer tools across cloud environments.
  • Improve reliability, scalability, performance, cost efficiency, observability, reproducibility, auditability, and SLO-based operations.
  • Support MLOps, security governance, incident response, root-cause remediation, and the mentoring and development of platform engineering talent.
Requirements
  • Bachelor’s degree in Computer Science or a related field, or equivalent practical experience, with 3+ years of software engineering experience.
  • Production experience operating distributed systems and reliable, high-scale infrastructure, including Kubernetes, AWS or GCP, infrastructure-as-code, and CI/CD.
  • Experience with GPU infrastructure, including scheduling, multi-node orchestration, and resource optimization for long-running ML workloads.
  • Strong Python skills and understanding of networking, microservices, infrastructure services, and distributed systems.
  • Familiarity with MLOps practices such as model versioning, pipeline orchestration, experiment tracking, artifact management, and reproducible ML workflows; experience with distributed training, HPC, inference serving, high-performance networking, or game simulation is a plus.
Skills
  • Kubernetes
  • GPU infrastructure
  • Distributed systems
  • MLOps
  • Python
  • AWS
  • GCP
  • Infrastructure as code
  • CI/CD
  • Distributed ML training

Hourly rate is commensurate with experience and is an estimated range provided by WorkGenius.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Platform Engineer — GPU, Kubernetes Infra
Senior ML Platform Engineer — GPU, Kubernetes Infra

WorkGenius Group • Los Angeles (CA)

On-site
USD 117,000 - 186,000
Senior Platform & Infrastructure Engineer – ML Platforms
Senior Platform & Infrastructure Engineer – ML Platforms

IDR, Inc. • Los Angeles (CA)

On-site
USD 180,000 - 240,000
MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
Senior ML Platform & Infra Engineer (Kubernetes, GPUs)
Senior ML Platform & Infra Engineer (Kubernetes, GPUs)

IDR, Inc. • Los Angeles (CA)

On-site
USD 180,000 - 240,000
AI/ML Infra Engineer - Hosting
AI/ML Infra Engineer - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California)
Senior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California)

Match Group • West Hollywood (CA)

On-site
USD 190,000 - 246,000
MLOps / ML Platform Engineer
MLOps / ML Platform Engineer

sumersports • Germany (OH)

Hybrid
USD 100,000 - 130,000
Competitive Salary and Bonus Plan
Comprehensive health insurance plan
Retirement savings plan (401k) with company match
+2
Software Engineer/Senior Software Engineer, Data & ML Platform
Software Engineer/Senior Software Engineer, Data & ML Platform

PlusAI • Santa Clara (CA)

On-site
USD 135,000 - 200,000
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options