Senior Machine Learning Infrastructure Engineer

Bee Talent Solutions

Los Angeles (CA)

On-site

USD 150,000 - 210,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, vision coverage

Job summary

Bee Talent Solutions seeks a Senior Platform & Infrastructure Engineer to design, build, and operate core infra and ML platforms powering Understanding Agents. You’ll focus on compute/orchestration for large-scale distributed training, simulation environments, and policy evaluation, plus CI/CD, IaC, observability, and developer tooling for production-grade systems.

You will close platform gaps, drive automation and maturity, and influence technical direction beyond your team, operating

Qualifications

  • Bachelor's degree in CS or equivalent practical experience.
  • 3+ years in software engineering with infra/platform/SRE focus.
  • Experience operating distributed systems in production.
  • Strong Kubernetes, AWS or GCP, IaC, CI/CD, and prod tooling.
  • GPU compute infra experience for long-running training workloads.
  • Proficiency in Python and networking fundamentals.
  • Familiar with MLOps workflows: versioning, pipelines, artifacts.

Responsibilities

  • Build and operate Kubernetes clusters for distributed ML bot training.
  • Design infra for large-scale game simulation environments.
  • Build CI/CD, deployment automation, and IaC across clouds.
  • Improve platform reliability, cost efficiency, and observability.
  • Develop internal APIs, control planes, and developer tooling.
  • Support MLOps workflows including training pipelines and experiments.
  • Manage production incidents and mentor engineers.

Skills

Kubernetes
Distributed systems
Python
Networking basics
CI/CD
AWS/GCP
MLOps
GPU compute

Education

Bachelor's degree in Computer Science or related field

Tools

Infrastructure-as-code

Job description

As a Senior Platform & Infrastructure Engineer on the MLBots team, you will design, build, and operate the core infrastructure and ML platforms behind Understanding Agents. Your focus will be on the compute and orchestration platforms that power large-scale distributed training of Game Understanding Agents (e.g. via RL, IL, and other techniques), simulation environments, and policy evaluation, as well as the CI/CD, infrastructure-as-code, observability, and developer tooling that keep these systems production-grade.

You will close critical infrastructure gaps across the team's stack, driving improvements to standards, automation, and operational maturity. You will operate independently on multi-month work efforts and begin to influence technical direction beyond your immediate team.

Responsibilities:
  • Build and operate Kubernetes, multi-node GPU clusters, and networking infrastructure for distributed ML bot training and large-scale policy evaluation.
  • Design infrastructure for running game simulation environments at scale, enabling parallel rollouts, data collection, training, and evaluation.
  • Build CI/CD, deployment automation, artifact management, and infrastructure-as-code across cloud environments.
  • Improve platform reliability, cost efficiency, performance, reproducibility, auditability, and operational maturity.
  • Build observability, monitoring, alerting, health indicators, and SLO-aligned dashboards for infrastructure and ML workloads.
  • Develop internal APIs, control planes, templates, and developer tooling for distributed training and evaluation workflows.
  • Support MLOps workflows including automated training pipelines, model artifact management, experiment tracking, and reproducible ML lifecycle operations.
  • Build security and governance controls, manage production incidents, drive root-cause remediation, mentor engineers, and support recruiting for platform roles.
Qualifications:
  • Bachelor’s degree in Computer Science or a related field, or equivalent practical experience.
  • 3+ years of software engineering experience, with meaningful experience in infrastructure, platform engineering, or SRE roles.
  • Experience operating distributed systems in production and keeping them healthy under real load.
  • Strong experience with Kubernetes, AWS or GCP, infrastructure-as-code, CI/CD, deployment automation, and production tooling.
  • Experience with GPU compute infrastructure, including scheduling, multi-node orchestration, and resource optimization for long-running training workloads.
  • Proficiency in Python and solid understanding of networking, microservices, core infrastructure services, and distributed systems fundamentals.
  • Familiarity with MLOps workflows such as model versioning, pipeline orchestration, experiment tracking, artifact management, and reproducible ML workflows.
  • Bonus: experience with distributed training or HPC frameworks, inference serving, systems languages, high-performance networking, game AI or simulation, Unreal/client-server architecture, AI-assisted development tools, or a passion for games and player experience.
Benefits:

Bee Talent Solutions offers a competitive benefits package for eligible full-time employees, including medical, dental, and vision coverage. Additional details regarding eligibility and plan offerings will be provided during the interview and onboarding process.

It’s our policy to provide equal employment opportunity for all applicants and employees of Bee Talent Solutions. The Company makes reasonable accommodations for handicapped and disabled employees and does not unlawfully discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, handicap, veteran status, marital status, criminal history, or any other category protected by applicable federal and state law.

We consider for employment all qualified applicants, including those with criminal histories, in a manner consistent with applicable federal, state and local law, including, but not limited to, the California Fair Chance Act, the City of Los Angeles Fair Chance Initiative for Hiring Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, the San Francisco Fair Chance Ordinance, and the Washington Fair Chance Act.

Per the Los Angeles County Fair Chance Ordinance, the following core duties may create a basis for disqualifying candidates with relevant criminal histories:

  • Safeguarding confidential and sensitive data while employed by us and while on assignment at a customer of ours
  • Communication with others, including employees and third parties such as vendors, customers (including their employees), and/or players, including minors
  • Accessing our or our customer’s assets, secure digital systems, and networks
  • Ensuring a safe interactive environment for players, employees, and temporary workers

These duties are directly related to essential operations, safety, trust, and compliance obligations within our organization and within the organization of any customer to whom you may be assigned while employed by us. Please note that job duties may evolve based on business needs and additional responsibilities may be assigned as necessary to maintain operational efficiency and security.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Platform & Infrastructure - Riot Technology
Senior Software Engineer, Platform & Infrastructure - Riot Technology

Riot Games • Town of Montana (WI)

On-site
USD 140,000 - 190,000
Open PTO
Medical/Dental/Life Insurance
401k with company match
Senior Software Engineer, Platform & Infrastructure - Riot Technology
Senior Software Engineer, Platform & Infrastructure - Riot Technology

Riot Games • Mercer Island (WA)

On-site
USD 170,000 - 230,000
Open PTO policy
Medical, dental, and life insurance
401k with company match
Senior Software Engineer, Platform & Infrastructure - Riot Technology
Senior Software Engineer, Platform & Infrastructure - Riot Technology

Socket.dev • Los Angeles (CA)

On-site
USD 180,000 - 250,000
Medical insurance
Dental insurance
Life insurance
+3
Senior Software Engineer, Platform & Infrastructure - Riot Technology
Senior Software Engineer, Platform & Infrastructure - Riot Technology

Riot Games • Los Angeles (CA)

On-site
USD 162,000 - 227,000
Medical insurance
Dental insurance
Life insurance
+4
Staff Machine Learning Engineer (Applied Modeling) - League of Legends
Staff Machine Learning Engineer (Applied Modeling) - League of Legends

Riot Games • Los Angeles (CA)

On-site
USD 210,000 - 320,000
Open paid time off policy
Medical, dental, and life insurance
401k with company match
Senior Data Scientist_Hybrid
Senior Data Scientist_Hybrid

Bee Talent Solutions • Los Angeles (CA)

On-site
Principal Machine Learning Engineer - League of Legends
Principal Machine Learning Engineer - League of Legends

Latitude • Los Angeles (CA), Northern (KY)

Hybrid
USD 292,000 - 438,000
Medical insurance
Dental insurance
Life insurance
+2
Principal Machine Learning Engineer - League of Legends
Principal Machine Learning Engineer - League of Legends

Socket.dev • Los Angeles (CA)

On-site
USD 292,000 - 438,000
Medical insurance
Dental insurance
Life insurance
+4
Staff Machine Learning Engineer - Game Tech Group, ML Platform
Staff Machine Learning Engineer - Game Tech Group, ML Platform

Riot Games • Los Angeles (CA)

On-site
USD 229,000 - 320,000
Medical, dental, and life insurance
401(k) with company match
Open PTO
Software Engineer, RL Training Infra
Software Engineer, RL Training Infra

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000