Senior Site Reliability Engineer

31st Union

Austin (TX)

Hybrid

USD 140,000 - 210,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

2K is seeking a Senior Site Reliability Engineer to lead production infrastructure across multi-cloud and hybrid environments in Austin. You will own Kubernetes platforms end-to-end, drive reliability through architecture decisions, and collaborate with network engineers, game studios, and platform teams.

You will build observability, implement SLI/SLOs, lead incident response and post-mortems, and push for security and scalable operations across 2K's live-service titles.

Qualifications

  • 5+ years in SRE, platform engineering, or equivalent infrastructure at production scale.
  • Deep experience with cloud environments (EKS/GKE), including networking and multi-cluster setups.
  • Infrastructure as Code with Terraform and/or Pulumi; experience with Helm, Terragrunt, and GitOps.
  • Experience with AWS, GCP, VMware, and bare metal environments.
  • Configuration management with Ansible, Puppet, and AWS Systems Manager.
  • Observability with Datadog, Prometheus, Grafana, and OpenTelemetry.
  • Production-grade coding in Go, Python, or TypeScript for tools and automation.
  • Solid Linux internals, TCP/IP, DNS, TLS, for system-level debugging.
  • Incident management and post-mortem leadership with systemic follow-through.

Responsibilities

  • Design, build, and operate scalable multi-cloud and hybrid infrastructure.
  • Own Kubernetes platforms end-to-end lifecycle, multi-tenancy, networking, and autoscaling.
  • Build observability stack and define SLI/SLO policies; implement alerting.
  • Lead chaos engineering exercises to surface failure modes before players encounter them.
  • Drive incident response and post-mortems with a focus on systemic fixes.
  • Embed security at the platform layer through secrets management and policy-as-code.
  • Promote SRE practices across 2K studios through reliability reviews and RFCs.

Skills

Kubernetes
Terraform
Pulumi
GitOps
Go
Python
TypeScript
Networking
Datadog
Incident response
CI/CD

Tools

ArgoCD
GitHub Actions
Helm
Terragrunt
OpenTelemetry
Datadog

Job description

At 2K, we create some of the most iconic and culture-shaping video games in entertainment, including NBA® 2K, one of the top-selling franchises in the world, and legendary titles like BioShock®, Borderlands®, Mafia, Sid Meier’s Civilization®, and XCOM®, as well as fan favorites WWE® 2K, TopSpin®, and PGA TOUR® 2K. We build unforgettable experiences by pushing the boundaries of creativity, authenticity and innovation across every genre.

Our portfolio is brought to life by some of the most influential game development studios in the world. Visual Concepts, Firaxis Games, Hangar 13, Cat Daddy Games, 31st Union, Cloud Chamber, Gearbox, HB Studios, and 2K SportsLab create world-class experiences across platforms. But what truly powers 2K is our people. We believe the best ideas come from teams that feel empowered, supported, and inspired. As an equal opportunity employer, we are committed to fostering a diverse, inclusive workplace where people are encouraged to come as they are and do their best work.

The Team

The 2K SRE team owns the infrastructure behind every player connection—All 2K game services, account platforms, CI/CD pipelines, and developer tooling spanning AWS, GCP, and on-premises data centers across multiple global regions. Global launch windows and live-service events push systems to their limits, and this team is expected to hold the line.

Post-mortems here focus on systems, not people. Automation is the default answer to repetitive work. The infrastructure keeps millions of players connected, and the team takes that seriously!

The Role

The Senior SRE at 2K is a hands-on technical leader—shaping production infrastructure across multiple clouds and regions while partnering with network engineers, systems architects, and game studio developers. This is an ownership role: driving technical direction, influencing reliability from architecture review through production operation, and closing the gap between what engineering ships and what players experience.

What You'll Do
Platform & Infrastructure

Design, build, and operate scalable multi-cloud and hybrid infrastructure using Terraform, Pulumi, and GitOps workflows (ArgoCD, Flux).

Own Kubernetes platforms (EKS, GKE) end-to-end cluster lifecycle, multi-tenancy, networking (Istio, Cilium), and autoscaling.

Observability & Reliability

Build and run the full observability stack: Prometheus + Grafana + Datadog.

Define SLI/SLO/error budget policies and build alerting that cuts through the noise.

Lead chaos engineering exercises to surface failure modes before players encounter them.

Drive incident response and post-mortems with a focus on systemic fixes and real follow-through.

Automation, Security & Developer Experience

Eliminate toil through self-service provisioning, automated remediation, and intelligent scaling.

Embed security at the platform layer through secrets management (PasswordState, 1Password, and AWS Secrets Manager) and policy-as-code (OPA/Gatekeeper).

Leadership

Promote SRE practices across 2K studios through reliability reviews, runbooks, and embedded collaboration.

Shape architectural decisions and author engineering RFCs that move the platform forward.

Required Qualifications

Experience: 5+ years in SRE, Platform Engineering, or equivalent infrastructure work at production scale.

Kubernetes: Deep experience in cloud environments (EKS or GKE preferred), including networking, storage, and multi-cluster patterns.

Infrastructure as Code (IaC): Strong proficiency with Terraform and/or Pulumi; hands-on with Helm, Terragrunt, and GitOps tooling (ArgoCD or GitHub Actions).

Environments: Experience with modern and legacy tech, including AWS, GCP, VMware, and Bare metal servers.

Configuration Management: Server configuration using Ansible, Puppet, and AWS Systems Manager.

Observability: Experience with Datadog, Prometheus + Grafana, and OpenTelemetry; fluency in operationalizing SLI/SLO/error budgets inside engineering teams.

Software Engineering: Production-quality code in Go, Python, or TypeScript for tools, automation, and internal libraries.

Systems & Networking: Solid understanding of Linux internals, TCP/IP networking, DNS, and TLS proven enough to debug at the system level.

Incident Management: Incident response and post-mortem leadership with a track record of systemic follow-through.

Preferred Qualifications

Live-service game or large-scale consumer internet experience dealing with millions of concurrent users.

Deep knowledge of Service mesh (Istio, Cilium) and advanced Kubernetes networking.

Experience with FinOps and managing resources efficiently at cloud scale.

Experience with AI and Agentic Development.

Cloud certifications (AWS Solutions Architect, GCP Professional Cloud Architect, CKA/CKS, or equivalent).

Experience mentoring SREs or leading reliability working groups.

As an equal opportunity employer, we are committed to ensuring that qualified individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform their essential job functions, and to receive other benefits and privileges of employment. Please contact us if you need reasonable accommodation.

Please note that 2K Games and its studios never uses instant messaging apps or personal email accounts to contact prospective employees or conduct interviews and when emailing, only use 2K.com accounts.

2K develops and publishes interactive video games for consoles, PCs, and mobile devices, featuring renowned franchises like NBA 2K, BioShock, and Borderlands.

Computer Games Software Development Technology, Information and Internet Technology, Information and Media

Apply For This Role

Company 2K

Location Hybrid - Austin, Texas, United States

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

2K • Austin (TX)

On-site
USD 120,000 - 160,000
Staff Platform Engineer
Staff Platform Engineer

2K • Austin (TX)

On-site
USD 180,000 - 240,000
Staff Platform Engineer
Staff Platform Engineer

2K • Los Angeles (CA)

On-site
USD 160,000 - 180,000
Lead Security Architect
Lead Security Architect

Linuxconfig • Austin (TX)

On-site
USD 170,000 - 260,000
Lead Security Architect
Lead Security Architect

2K • Austin (TX)

On-site
USD 140,000 - 210,000
Technical Director, Product Engineering
Technical Director, Product Engineering

Socket.dev • Austin (TX)

On-site
USD 150,000 - 230,000
Senior Technical Program Manager Austin, Texas, United States
Senior Technical Program Manager Austin, Texas, United States

2K Games, Inc. • Austin (TX)

On-site
USD 120,000 - 190,000
Manager, Engineering
Manager, Engineering

2K • Austin (TX)

On-site
USD 130,000 - 160,000
Senior Systems Engineer
Senior Systems Engineer

31st Union • Novato (CA)

On-site
USD 109,000 - 161,000
Medical, dental, vision, and basicLife
14 paid holidays
Paid vacation (15–25 days)
+6
Site Reliability Engineer II
Site Reliability Engineer II

Sony Playstation • Aliso Viejo (CA)

On-site
USD 120,000 - 150,000