Site Reliability Engineer, Infrastructure Platforms — UK (Intermediate to Senior Staff)

GitLab

United States

Remote

USD 140,000 - 190,000

Full time

13 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health benefits
Flexible Paid Time Off
Team Member Resource Groups
Equity Compensation & Employee Stock P
Growth and Development Fund
Parental Leave

Job summary

GitLab is seeking Site Reliability Engineers for Infrastructure Platforms. This remote-first posting evaluates candidates holistically to match you to the best opportunity across Intermediate to Senior Staff roles.

You’ll work on reliability, automation, and IaC, operating Kubernetes at scale, and contribute to observability and incident response in a global, async team. UK location is emphasized in the posting.

Qualifications

  • Experience keeping production systems reliable and scalable.
  • Experience building automation and tooling for infra, e.g., Terraform modules, Kubernetes operators.
  • Ability to read and reason about Go/Ruby code and discuss behavior, performance, and failure modes.
  • Deep knowledge of Kubernetes and cloud providers (GCP or AWS).
  • Familiarity with observability, metrics, logs, alerts, and SLOs/SLI, using data for operations.
  • Comfort in on-call and incident response, with structured troubleshooting.
  • Strong written communication and ability to work async in distributed environments.
  • Track record of using automation and AI to reduce toil.
  • Alignment with GitLab values and working accordingly.

Responsibilities

  • Keep user-facing services reliable, scalable, and efficient.
  • Build automation and tooling to reduce toil using IaC-driven workflows.
  • Operate and troubleshoot production systems on Kubernetes, including deployments and scaling.
  • Write and maintain infrastructure as code, shipping changes via CI/CD and GitOps.
  • Participate in on-call, triage alerts, and improve runbooks.
  • Contribute to observability using metrics, logs, and SLOs to detect signals early.
  • Participate in incident response and post-incident reviews to drive improvements.
  • Document runbooks and architecture decisions to make findings repeatable.

Skills

Production reliability
Automation tooling
Go language
Ruby language
Kubernetes experience
Observability
On-call experience
Written communication
Automation
Cloud providers

Tools

Terraform
Kubernetes
CI/CD tooling
GitOps

Job description

An overview of this role

Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve our production infrastructure.

This is a single application for Site Reliability Engineering opportunities across Infrastructure Platforms. Rather than asking you to choose the right team or level upfront, we evaluate your skills holistically and match you to the opportunity that best aligns with your experience and our hiring needs. We hire Site Reliability Engineers from Intermediate through Senior Staff across multiple Infrastructure Platforms teams.

We don't expect every candidate to have experience with every technology in our environment. We're looking for engineers with strong technical fundamentals, a growth mindset, and the ability to learn quickly. We'll support you in becoming successful with GitLab's tools, systems, and ways of working.

Please note: This position is open to candidates based in the United Kingdom only. Candidates based in the United States or Canada can apply to this posting:

Site Reliability Engineer, Infrastructure Platforms - AMER (Intermediate to Senior Staff)

How our SRE hiring works

Because this is a single application for SRE roles across Infrastructure Platforms, our process is built to evaluate you once and match you well, rather than interviewing separately for every team.

  • Recruiter Screen: A conversation about your background, what you're looking for, and the level and teams that fit, so we can point your process in the right direction.
  • Core Technical: The shared assessment every SRE candidate takes, regardless of eventual team. A low-stress, collaborative discussion covering system architecture and incident review.
  • Hiring Manager Interview: A conversation about ownership, judgment, execution, collaboration, and growth, the non-technical signals that make an SRE effective at GitLab.
  • Peer Technical: Team-specific depth, run by SREs from the team you're most likely to join, focused on the problems that team actually works on.
  • Skip-Level Interview: A conversation with a senior leader on values alignment, and how you'll work across teams.

After your interviews, we consider your performance alongside our current hiring needs to confirm the level and team where you'll do your best work. Interview results are a major factor, and final placement also reflects our active hiring priorities at the time.

We’ll calibrate your level throughout the interview process based on the scope and impact of your experience.
  • Intermediate: You independently deliver meaningful reliability improvements within a defined area.
  • Senior: You own complex reliability work end to end and raise the effectiveness of your team.
  • Staff: You shape reliability across multiple teams, solving systemic problems and creating approaches others can reuse.
  • Senior Staff: You set technical direction across a broader Infrastructure area and influence reliability strategy at organizational scale.
What you'll do
  • Keep user-facing services and production systems reliable, scalable, and efficient
  • Build automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflows
  • Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling
  • Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps
  • Participate in on-call, triage alerts, follow and improve runbooks, and elevate appropriately
  • Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outages
  • Take part in incident response and post-incident reviews, turning learnings into changes in automation and process
  • Document runbooks, architecture decisions, and reviews so your findings become repeatable practices
What you'll bring
  • Experience keeping production systems reliable, combining an operations mindset with real software engineering practice
  • Experience building net-new infrastructure tooling and automation, not just configuring existing tools. For example, Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratch
  • The ability to read, debug, and reason about code. Most of our teams work in Go; some work in Ruby. You can discuss a piece of code's behavior, performance, and failure modes
  • Experience with infrastructure as code, and with Kubernetes and its ecosystem, at a depth appropriate to your level
  • Hands-on experience with at least one major cloud provider (GCP or AWS)
  • Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs, and using data to inform operational decisions
  • Comfort participating in on-call and incident response, with a structured approach to troubleshooting under pressure
  • Strong written communication and the ability to operate as a manager-of-one in an async, distributed environment
  • A track record of using automation, and increasingly AI, to reduce toil and improve how you and your team work
  • Alignment with GitLab's values and a commitment to working in accordance with them
About the team

Infrastructure Platforms is responsible for the availability, reliability, performance, and scalability of GitLab’s user-facing services, most notably GitLab.com. The organization spans teams across Production Engineering, GitLab Dedicated, GitLab Delivery, and Developer Experience, covering everything from the production fleet and networking platform to observability, incident response, deployment infrastructure, tenant scale, and our single-tenant Dedicated offering. We are a globally distributed, remote-first organization that works asynchronously, favors automation over toil, and uses monitoring, metrics, and clear ownership to continuously improve the reliability of GitLab at scale. For more on how we work, see the Infrastructure Handbook Page.

How GitLab Supports Full-Time Employees
  • Benefits to support your health, finances, and well-being
  • Flexible Paid Time Off
  • Team Member Resource Groups
  • Equity Compensation & Employee Stock Purchase Plan
  • Growth and Development Fund
  • Parental Leave
Country Hiring Guidelines

GitLab hires new team members in countries around the world. All of our roles are remote, however some roles may carry specific location-based eligibility requirements. Our Talent Acquisition team can help answer any questions about location after starting the recruiting process.

Privacy Policy

Please review our Recruitment Privacy Policy.

GitLab is proud to be an equal opportunity workplace and is an affirmative action employer. GitLab’s policies and practices relating to recruitment, employment, career development and advancement, promotion, and retirement are based solely on merit, regardless of race, color, religion, ancestry, sex (including pregnancy, lactation, sexual orientation, gender identity, or gender expression), national origin, age, citizenship, marital status, mental or physical disability, genetic information (including family medical history), discharge status from the military, protected veteran status (which includes disabled veterans, recently separated veterans, active duty wartime or campaign badge veterans, and Armed Forces service medal veterans), or any other basis protected by law. GitLab will not tolerate discrimination or harassment based on any of these characteristics. See also GitLab’s EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know during the recruiting process.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer, Platform Engineering: Dedicated New Remote, Canada; Remote, United Kingdom; Remote, United States
Principal Site Reliability Engineer, Platform Engineering: Dedicated New Remote, Canada; Remote, United Kingdom; Remote, United States

GitLab Inc. • Northern (KY)

Hybrid
USD 180,000 - 240,000
Director, Engineering, Platform Operations & Productivity
Director, Engineering, Platform Operations & Productivity

GitLab • United States

Remote
USD 260,000 - 360,000
Flexible Paid Time Off
Equity Compensation & Employee Stock-P
Growth and Development Fund
+1
Distinguished Engineer, Core DevOps
Distinguished Engineer, Core DevOps

Arbeitnow • United States

Remote
INR 23,992,000 - 33,493,000
Flexible Paid Time Off
Equity Compensation & Employee Stock-P
Growth and Development Fund
+2
Senior Engineering Manager - Continuous Deployment
Senior Engineering Manager - Continuous Deployment

GitLab • United States

On-site
USD 170,000 - 230,000
Flexible Paid Time Off
Team Member Resource Groups
Equity Compensation & Employee Stock P
+2
Distinguished Engineer, Core DevOps
Distinguished Engineer, Core DevOps

GitLab • United States

Remote
USD 250,000 - 349,000
Health benefits
Flexible PTO
Employee Resource Groups
+3
Senior Backend Engineer, Platform Enablement
Senior Backend Engineer, Platform Enablement

GitLab • United States

Remote
USD 157,000 - 235,000
Flexible Paid Time Off
Equity Compensation & Employee Stock-
Growth and Development Fund
+1
Senior Assigned Support Engineer (EMEA)
Senior Assigned Support Engineer (EMEA)

GitLab • United States

Remote
USD 110,000 - 160,000
Health benefits
Flexible PTO
Employee resource groups
+3
Senior Backend Engineer, Platform Enablement
Senior Backend Engineer, Platform Enablement

GitLab Inc. • Northern (KY)

Remote
USD 140,000 - 210,000
Remote Site Reliability Engineer - Infrastructure Platforms
Remote Site Reliability Engineer - Infrastructure Platforms

GitLab • United States

Remote
USD 140,000 - 190,000
Health benefits
Flexible Paid Time Off
Team Member Resource Groups
+3
Distinguished Engineer, Core DevOps New Remote, Canada; Remote, United Kingdom; Remote, United States
Distinguished Engineer, Core DevOps New Remote, Canada; Remote, United Kingdom; Remote, United States

GitLab Inc. • Northern (KY)

Hybrid
USD 210,000 - 270,000