Senior Site Reliability Engineer

theaccessgroup

United States

Hybrid

USD 140,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
401(k) match
PTO 22 days
Holidays 11 days

Job summary

The Access Group is seeking a Senior Site Reliability Engineer in the United States to own cross-system incidents, drive architectural remediation, and lead platform-level reviews.

You will design and operate cloud infrastructure in Azure, run Kubernetes in production, and manage IaC with Terraform while improving observability with Datadog and PagerDuty in collaboration with Engineering, Security, and Operations teams.

Qualifications

  • 8+ years in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering.
  • Senior escalation point for cross-team incidents requiring architectural decision-making.
  • Expert-level Azure cloud design and operation with AWS knowledge.

Responsibilities

  • Serve as senior escalation point for complex production incidents (P1/P2).
  • Lead platform-level architecture reviews ensuring reliability and security before implementation.
  • Identify systemic failure patterns and translate into architectural changes and tooling improvements.
  • Own availability, reliability, and performance of production systems daily.

Skills

SRE/Platform Engineering
Azure cloud
Kubernetes
Terraform IaC
Bash scripting
Networking basics
Observability (Datadog, PagerDuty)
CI/CD pipelines

Tools

Terraform
Datadog
PagerDuty

Job description

We're looking for people to join the Access family, who share our passion for believing in better, and who will help us continue to grow.

Love Work. Love Life. Be You. - is central to our success and how we give our customers the freedom to do more of what's important to them.

What does Access offer you?

We offer a blended approach to office working, encouraging you to collaborate and connect in one of our thriving offices. We deliver on what we say, taking the development of our people seriously. We'll work with you to progress your success plan and provide opportunities to accelerate your career.

On top of a competitive salary, you'll receive 22 days paid time off, plus 11 company paid holidays. Also, medical, dental & vision insurance, 5% 401(k) company match, plus a range of other benefits that you can choose from.

About You:

You think in systems, not just technologies. You are the engineer your peers escape to when the problem is hard, the blast radius is unclear, and the path forward requires both depth and judgment. You are equally at home designing cloud architecture, leading a complex P1 incident response, and writing the Terraform module that ensures it never happens again.

You bring discipline to post-incident reviews, urgency to production issues, and genuine satisfaction to the quieter work - a clean runbook, a well-structured SLO, a piece of toil that no longer exists. You don't separate strategy from operations. You understand that the best Site Reliability Engineers do both, every day, at a senior level.

You are a clear communicator who can lead a technical architecture review with engineers and then explain the same decision to an executive stakeholder without losing either audience. You lead through credibility, influence without authority, and take ownership of outcomes - not just tasks. If this sounds like the way you already work, we'd love to talk.

Day-to-day, you will:
  • Serve as the senior escalation point for complex production incidents (P1/P2), owning cross-system triage and leading permanent architectural remediation - not just tactical fixes.
  • Lead platform-level architecture reviews, ensuring cloud infrastructure designs meet reliability, scalability, security, and operational standards before implementation.
  • Identify systemic failure patterns across incidents and translate them into architectural changes, design standards, and lasting platform improvements.
  • Own the availability, reliability, performance, and scalability of production systems on a daily basis.
  • Define, track, and improve Service Level Objectives (SLOs), Service Level Indicators (SLIs), and operational KPIs across critical services.
  • Develop and maintain Infrastructure-as-Code (IaC) solutions using Terraform, including module design, state management, and governance standards.
  • Identify and systematically eliminate operational toil through automation, self-service capabilities, and platform-level tooling.
  • Build and maintain automation frameworks and operational tooling using Bash, PowerShell, and related scripting technologies.
  • Administer and architect solutions within Microsoft Azure as the primary cloud platform, with working knowledge of AWS.
  • Operate Kubernetes in production, including cluster management, workload operations, and platform-level maintenance.
  • Manage hybrid-cloud environments including virtual machines, networking, and distributed infrastructure.
  • Maintain and improve Datadog observability and PagerDuty alerting configurations, championing observability standards across metrics, logging, tracing, and alerting.
  • Design infrastructure controls that meet PCI-DSS, SOC 1/2, and ISO 27001 compliance requirements, and support evidence collection during audits.
  • Partner with Engineering, Product, Security, and Operations teams to improve CI/CD pipelines, release processes, and overall DevOps maturity.
  • Mentor peers and junior engineers, and influence organizational engineering standards as the internal technical authority on infrastructure design.
Your skills and experiences might also include:
Required
  • 8+ years in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering, with direct ownership of complex production platforms at scale.
  • Proven track record as a senior technical escalation point for cross-team incidents requiring architectural-level decision-making and permanent remediation.
  • Expert-level experience designing and operating cloud infrastructure in Microsoft Azure, with working knowledge of AWS.
  • Deep expertise running Kubernetes in production - cluster design, workload operations, and platform-level maintenance.
  • Advanced Terraform and Infrastructure-as-Code (IaC) skills, including module design, state management, and governance.
  • Strong Bash scripting and automation development for operational tooling and self-service platform capabilities.
  • Solid networking fundamentals: firewalls, DNS, routing, VPN, and network troubleshooting, including Cloudflar
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

On-site
USD 140,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Level • Bellevue (WA)

On-site
USD 150,000 - 195,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Platform Engineer
Platform Engineer

Synergy • Chicago (IL)

On-site
USD 100,000 - 150,000
Senior Site Reliability Engineer - Cloud Platform | Energy Trading & Infrastructure Firm
Senior Site Reliability Engineer - Cloud Platform | Energy Trading & Infrastructure Firm

Techfellow Limited • New York (NY)

Hybrid
USD 270,000 - 330,000
Senior IT Reliability & Automation Lead
Senior IT Reliability & Automation Lead

First Horizon Bank • Memphis (TN)

On-site
USD 120,000 - 180,000
Senior CloudOps & Site Reliability Engineer
Senior CloudOps & Site Reliability Engineer

MangoApps • Seattle (WA)

On-site
USD 110,000 - 150,000
Senior Site Reliability Engineer, Colorado Springs
Senior Site Reliability Engineer, Colorado Springs

Onebrief • Colorado Springs (CO)

On-site
USD 140,000 - 190,000
Relocation assistance
On-site in Colorado Springs