Senior Site Reliability Engineer

Jobtailor

Arlington (VA)

On-site

USD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor in Arlington, VA seeks a Senior Platform/SRE to own the reliability, scalability, and security of our production app and platform. You will design, implement, and manage our observability stack and automate infrastructure with Terraform/Ansible.

You will lead incident response, define SLIs/SLOs, and drive blameless post-mortems to prevent recurrence, partnering with platform engineers to secure Kubernetes clusters and cloud/on-prem environments.

Qualifications

  • Active Top Secret clearance
  • 5+ years in Platform, DevOps, or SRE with infra focus
  • Proven cross-functional collaboration
  • Strong incident response and root-cause analysis experience
  • Terraform/Ansible IaC expertise
  • Kubernetes design, deployment and operations
  • CI/CD pipelines experience
  • Scripting in Python/Go/Bash
  • AWS or GovCloud experience
  • Observability toolchain (Grafana/ELK/Datadog)
  • Networking fundamentals and secure configurations

Responsibilities

  • Own the monitoring, logging, and alerting stack for production platforms.
  • Define SLIs/SLOs and establish reliability metrics.
  • Lead incident response and post-mortems to drive long-term fixes.
  • Build secure, scalable Kubernetes clusters with IaC; embed RMF/STIG controls.
  • Reduce toil by automating operations and sharing best practices.

Skills

Site Reliability Engineering
Platform/DevOps collaboration
Incident response
Infrastructure as Code
Kubernetes design/ops
CI/CD pipelines
Scripting (Python/Go/Bash)
AWS / GovCloud familiarity
Observability tooling
Networking fundamentals

Education

Active Top Secret clearance

Tools

Terraform
Ansible
Kubernetes
Prometheus
Grafana
ELK
Datadog
GitLab CI/CD
Jenkins
GitHub Actions
AWS
CloudFormation

Job description

Responsibilities

You will own the reliability, scalability, and security of the production application and/or platform. You will do this by:

  • Implementing a World-Class Observability Platform: Design, implement, and manage our monitoring, logging, and alerting stack (e.g., Prometheus, Loki, Alloy, and Grafana). You won't just track metrics; you'll create the actionable insights and automated alerting that allow teams to identify and resolve issues before they impact users.
  • Defining and Upholding Reliability: Define, measure, and own alerting that feeds into our Service Level Indicators (SLIs) and Service Level Objectives (SLOs), increasing trust internally and externally. You will be the organization's expert on what it means for our systems to be reliable and how to measure it.
  • Leading Incident Response: Act as the incident responder and potentially incident commander during critical incidents who will lead blameless post-mortems / After Action Reviews (AARs) that identify true root causes and drive automated, long-term solutions to prevent recurrence.
  • Automating for Scale and Security: Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible). You will embed security and compliance controls (RMF, STIGs) directly into this automation.
  • Eliminating Toil and Scaling the Team: Proactively identify and eliminate operational toil by building automation. You will partner with other teams to share best practices for air-gapped environments and support their readiness for production.
Requirements
  • An active Top Secret clearance
  • 5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.
  • Proven partner to DevOps/Platform and application teams; collaborates well across functions and shares context openly.
  • A deep understanding of incident response processes, with experience conducting thorough root cause analyses and driving continuous improvement.
  • Technical expertise
  • Infrastructure as Code: Terraform (or CloudFormation), Ansible.
  • Containers and orchestration: Kubernetes design, deployment, and operations.
  • CI/CD: experience building and maintaining pipelines (GitLab CI/CD, Jenkins, GitHub Actions).
  • Scripting: proficiency with at least one of Python, Go, or Bash.
  • Cloud: Familiarity with AWS or AWS GovCloud.
  • Observability: Grafana stack, ELK stack, or Datadog.
  • Networking fundamentals: core protocols and secure configurations.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • California (MO)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

TechDigital Group • Houston (TX), Juno Beach (FL)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Shrive Technologies LLC • Schaumburg (IL)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Shrive Technologies • Schaumburg (IL)

On-site
USD 120,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Buffalo (NY)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer – Secret Clearance Required
Senior Site Reliability Engineer – Secret Clearance Required

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 190,000