Senior Software Engineer, Platform & Infrastructure - Riot Technology

Socket.dev

Los Angeles (CA)

On-site

USD 180,000 - 250,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
Dental insurance
Life insurance
Parental leave (you, spouse/partner, &
401k with company match
Flexible work schedules

Job summary

Riot Games is seeking a Senior Platform & Infrastructure Engineer to design, build, and operate the core infrastructure and ML platforms that power large-scale distributed training and simulations. You will work on computing and orchestration platforms, CI/CD, infrastructure-as-code, observability, and developer tooling to keep systems production-grade.

You will collaborate with engineers, data scientists, designers, and product teams to close critical gaps, improve reliability, cost efficiency,

Qualifications

  • Bachelor’s degree in Computer Science or a related field, or equivalent practical experience.
  • 3+ years of software engineering experience in infrastructure, platform engineering, or SRE roles.
  • Experience operating distributed systems in production and under real load.
  • Strong experience with Kubernetes, AWS or GCP, infrastructure-as-code, CI/CD, deployment automation, and production tooling.
  • Experience with GPU compute infrastructure, including scheduling, multi-node orchestration, and resource optimization for long-running training workloads.
  • Proficiency in Python and solid understanding of networking, microservices, core infrastructure services, and distributed systems fundamentals.

Responsibilities

  • Build and operate Kubernetes, multi-node GPU clusters, and networking infrastructure for distributed ML bot training and large-scale policy evaluation.
  • Design infrastructure for running simulation environments at scale, enabling parallel rollouts, data collection, training, and evaluation.
  • Build CI/CD, deployment automation, artifact management, and infrastructure-as-code across cloud environments.
  • Improve platform reliability, cost efficiency, performance, reproducibility, auditability, and operational maturity.
  • Build observability, monitoring, alerting, health indicators, and SLO-aligned dashboards for infrastructure and ML workloads.
  • Develop internal APIs, control planes, templates, and developer tooling for distributed training and evaluation workflows.
  • Support MLOps workflows including automated training pipelines, model artifact management, experiment tracking, and reproducible ML lifecycle operations.
  • Build security and governance controls, manage production incidents, drive root-cause remediation, mentor engineers, and support recruiting for platform roles.

Skills

Kubernetes
Cloud (AWS/GCP)
Infrastructure as code
CI/CD
GPU compute infra
Python
Distributed systems
Networking fundamentals

Education

Bachelor’s degree in Computer Science or related field

Tools

Terraform
Docker
Prometheus
CI/CD tooling

Job description

Platform and infrastructure engineers at Riot build the foundational systems that enable teams to develop, deploy, and operate systems at global scale. They partner across disciplines with other engineers, data scientists, designers, and product teams to ensure reliable, scalable, and secure infrastructure underpins every ML capability that reaches players.

As a Senior Platform & Infrastructure Engineer on the Riot Technology team, you will design, build, and operate the core infrastructure and ML platforms. Your focus will be on the computing and orchestration platforms that power large-scale distributed training of agents (e.g. via RL, IL, and other techniques), simulation environments, and policy evaluation, as well as the CI/CD, infrastructure-as-code, observability, and developer tooling that keep these systems production-grade. You will close critical infrastructure gaps across the team's stack, driving improvements to standards, automation, and operational maturity. You will operate independently on multi-month work efforts and begin to influence technical direction beyond your immediate team. You will report to the Manager of Machine Learning.

Responsibilities:
  • Build and operate Kubernetes, multi-node GPU clusters, and networking infrastructure for distributed ML bot training and large-scale policy evaluation.
  • Design infrastructure for running simulation environments at scale, enabling parallel rollouts, data collection, training, and evaluation.
  • Build CI/CD, deployment automation, artifact management, and infrastructure-as-code across cloud environments.
  • Improve platform reliability, cost efficiency, performance, reproducibility, auditability, and operational maturity.
  • Build observability, monitoring, alerting, health indicators, and SLO-aligned dashboards for infrastructure and ML workloads.
  • Develop internal APIs, control planes, templates, and developer tooling for distributed training and evaluation workflows.
  • Support MLOps workflows including automated training pipelines, model artifact management, experiment tracking, and reproducible ML lifecycle operations.
  • Build security and governance controls, manage production incidents, drive root-cause remediation, mentor engineers, and support recruiting for platform roles.
Required Qualifications:
  • Bachelor’s degree in Computer Science or a related field, or equivalent practical experience.
  • 3+ years of software engineering experience, with meaningful experience in infrastructure, platform engineering, or SRE roles.
  • Experience operating distributed systems in production and keeping them healthy under real load.
  • Strong experience with Kubernetes, AWS or GCP, infrastructure-as-code, CI/CD, deployment automation, and production tooling.
  • Experience with GPU compute infrastructure, including scheduling, multi-node orchestration, and resource optimization for long-running training workloads.
  • Proficiency in Python and solid understanding of networking, microservices, core infrastructure services, and distributed systems fundamentals.
Desired Qualifications:
  • Familiarity with MLOps workflows such as model versioning, pipeline orchestration, experiment tracking, artifact management, and reproducible ML workflows.
  • Experience with distributed training or HPC frameworks, inference serving, systems languages, high-performance networking, Unreal/client-server architecture, AI-assisted development tools
  • Passion for games and player experience.

For this role, you'll find success through craft expertise, a collaborative spirit, and decision‑making that prioritizes the delight of players. We will be looking at your past studies, experience, and your personal relationship with games. If you embody player empathy and care about players' experiences, this could be your role!

Our Perks:

Riot focuses on work/life balance, shown by our open paid time off policy and other perks such as flexible work schedules. We offer medical, dental, and life insurance, parental leave for you, your spouse/domestic partner, and children, and a 401k with company match. Check out our benefits pages for more information.

At Riot Games, we put players first. That mission drives every decision in our quest to create games and experiences that make it better to be a player. Whether you’re working directly on a new player‑facing experience or you’re supporting the company as a whole, everyone at Riot is part of our mission. And just like in our games, we’re better when we work together. Our goal is to create collaborative teams where you are empowered to bring your unique perspective everyday. If that sounds like the kind of place you want to work, we’re looking forward to your application.

It’s our policy to provide equal employment opportunity for all applicants and members of Riot Games, Inc. Riot Games makes reasonable accommodations for handicapped and disabled Rioters and does not unlawfully discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, handicap, veteran status, marital status, criminal history, or any other category protected by applicable federal and state law. We consider for employment all qualified applicants, including those with criminal histories, in a manner consistent with applicable federal, state and local law, including the California Fair Chance Act, the City of Los Angeles Fair Chance Initiative for Hiring Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, the San Francisco Fair Chance Ordinance, and the Washington Fair Chance Act.

Per the Los Angeles County Fair Chance Ordinance, the following core duties may create a basis for disqualifying candidates with relevant criminal histories:

  • Safeguarding confidential and sensitive Company data
  • Communication with others, including Rioters and third parties such as vendors, and/or players, including minors
  • Accessing Company assets, secure digital systems, and networks
  • Ensuring a safe interactive environment for players and other Rioters

These duties are directly related to essential operations, safety, trust, and compliance obligations within our organization. Please note that job duties may evolve based on business needs and additional responsibilities may be assigned as necessary to maintain operational efficiency and security.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Platform & Infrastructure - Riot Technology
Senior Software Engineer, Platform & Infrastructure - Riot Technology

Riot Games • Mercer Island (WA)

On-site
USD 170,000 - 230,000
Open PTO policy
Medical, dental, and life insurance
401k with company match
Senior Software Engineer, Platform & Infrastructure - Riot Technology
Senior Software Engineer, Platform & Infrastructure - Riot Technology

Riot Games • Los Angeles (CA)

On-site
USD 162,000 - 227,000
Medical insurance
Dental insurance
Life insurance
+4
Staff Machine Learning Engineer - Game Tech Group, ML Platform
Staff Machine Learning Engineer - Game Tech Group, ML Platform

Riot Games • Los Angeles (CA)

On-site
USD 229,000 - 320,000
Medical, dental, and life insurance
401(k) with company match
Open PTO
Principal Machine Learning Engineer (Platform & Integrations)- League of Legends
Principal Machine Learning Engineer (Platform & Integrations)- League of Legends

Riot Games • Los Angeles (CA)

On-site
USD 292,000 - 438,000
Medical insurance
Dental insurance
Life insurance
+3
Principal Software Engineer - Developer Connections, Game Release
Principal Software Engineer - Developer Connections, Game Release

Riot Games • Los Angeles (CA)

On-site
USD 180,000 - 240,000
Health insurance
Dental insurance
Life insurance
+3
Staff Machine Learning Engineer (Applied Modeling) - League of Legends
Staff Machine Learning Engineer (Applied Modeling) - League of Legends

Riot Games • Los Angeles (CA)

On-site
USD 210,000 - 320,000
Open paid time off policy
Medical, dental, and life insurance
401k with company match
Staff Software Engineer - Esports Platforms
Staff Software Engineer - Esports Platforms

Riot Games • Los Angeles (CA)

On-site
USD 180,000 - 280,000
Open PTO
Medical insurance
Dental insurance
+3
Staff Data and AI Engineer - Enterprise
Staff Data and AI Engineer - Enterprise

Riot Games • Los Angeles (CA)

On-site
USD 180,000 - 240,000
Medical insurance
401k with company match
Flexible work schedules
Principal Data Engineer, Personalization - Central Product Insights
Principal Data Engineer, Personalization - Central Product Insights

Riot Games • Los Angeles (CA)

On-site
USD 180,000 - 240,000
Open PTO
Medical insurance
Dental insurance
+3
Staff Software Engineer - Data Foundations
Staff Software Engineer - Data Foundations

Riot Games • Los Angeles (CA)

On-site
USD 180,000 - 280,000
Work-life balance
Medical & dental
Parental leave
+2