Senior Site Reliability Engineer (Arlington, VA) - Relocation Provided

Onebrief

United States

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Relocation assistance

Job summary

Onebrief is hiring a Site Reliability Engineer to join our Infrastructure & Security team. You will own the reliability, scalability, and security of production deployments in DoD environments and AWS/cloud/on-prem. You’ll work with fellow SREs, security, and customer success, and contribute to automated guardrails and scalable architectures.

You must hold an active Top Secret clearance and be prepared for on-site work at customer locations in Arlington, VA with relocation support if needed.

Qualifications

  • Active Top Secret clearance required.
  • 5+ years in Platform/DevOps/SRE with an infrastructure and operations focus.
  • Proven partner to DevOps/Platform and application teams; collaborates well across functions and shares context openly.

Responsibilities

  • Own the reliability, scalability, and security of the production application and/or platform.
  • Implement a World-Class Observability Platform: design and manage monitoring, logging, and alerting stack (Prometheus, Loki, Alloy, Grafana).
  • Define and uphold reliability: measure and own SLIs/SLOs to increase trust.
  • Lead incident response: act as incident responder or commander during critical incidents and conduct blameless postmortems/AARs.
  • Automate for scale and security: build secure, resilient Kubernetes clusters and IaC (Terraform, Ansible).
  • Eliminate toil and scale the team: automate to improve reliability and efficiency.

Skills

Top Secret clearance
SRE / Platform focus
Cross-functional collaboration
Kubernetes
Terraform
Ansible
Security compliance (RMF/STIGs)
Observability
Prometheus
Grafana

Tools

Prometheus
Loki
Alloy
Grafana
Terraform
Ansible

Job description

Consequential Work. Dedicated People.

About Onebrief

Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination.

Military planning is complex by nature, requiring teams to coordinate information, people, and decisions across systems and locations. Onebrief brings planning, collaboration, simulation, and AI into one connected environment, helping teams test strategies, adapt to changing conditions, and make decisions with greater clarity when the stakes are real.

We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world.

Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth.

Security Clearance, Location, and Onsite Notice:

This role requires regularly working on-site at customer locations in Arlington, VA.

If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance).

Active Secret Clearance required; SCI eligibility is a plus.

About The Role

We are hiring a Site Reliability Engineer to join our Infrastructure & Security team. You’ll work closely with fellow SREs, security, and customer success.

You will be the first line of support for our mission critical deployments, and responsible for ensuring best-in-class service quality and issue resolution. You will work in both on-premise DoD environments and AWS cloud environments. Your lessons from the field will shape how our team works, from policy to implementation.

In addition to working at the customer, you will contribute directly to solutions that increase stability, performance, and security of our deployments, and improve the overall experience of deploying and managing Onebrief on premise.

About You

You care deeply about reliability and treat it as a core feature of any application or platform, with a bias toward “reliability over novelty.” You think about infrastructure and operability as products to be automated, well-documented, and continuously improved, and you aim to leave systems easier to operate than you found them.

You are equally comfortable leading a post-incident review, or diving into a kubectl shell to triage a complex production issue. You don't just fix problems; you translate constraints and failure modes into clear, automated guardrails and scalable, resilient architecture. For you, robust monitoring, actionable alerting, and insightful runbooks are core parts of the engineering process, not afterthoughts.

You mentor others, fostering a culture of blameless postmortems and proactive reliability. You collaborate naturally with application and platform teams, helping them move quickly but safely by building the tools, processes, and observability that make "fast recovery" a reality.

What You'll Do

You'll own the reliability, scalability, and security of the production application and/or platform. You will do this by:

  • Implementing a World-Class Observability Platform: Design, implement, and manage our monitoring, logging, and alerting stack (e.g., Prometheus, Loki, Alloy, and Grafana). You won't just track metrics; you'll create the actionable insights and automated alerting that allow teams to identify and resolve issues before they impact users.

  • Defining and Upholding Reliability: Define, measure, and own alerting that feeds into our Service Level Indicators (SLIs) and Service Level Objectives (SLOs), increasing trust internally and externally. You will be the organization's expert on what it means for our systems to be reliable and how to measure it.

  • Leading Incident Response: Act as the incident responder and potentially incident commander during critical incidents who will lead blameless post-mortems / After Action Reviews (AARs) that identify true root causes and drive automated, long‑term solutions to prevent recurrence.

  • Automating for Scale and Security: Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible). You will embed security and compliance controls (RMF, STIGs) directly into this automation.

  • Eliminating Toil and Scaling the Team: Proactively identify and eliminate operational toil by building automation. You will partner with other teams to share best practices for air-gapped environments and support their readiness for production.

What We Look For
  • An active Top Secret clearance

  • 5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.

  • Proven partner to DevOps/Platform and application teams; collaborates well across functions and shares context openly.

  • <
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (Arlington, VA) - Relocation Provided
Senior Site Reliability Engineer (Arlington, VA) - Relocation Provided

Onebrief • Arlington (VA), Northern (KY)

Hybrid
USD 140,000 - 170,000
Senior Site Reliability Engineer, Colorado Springs
Senior Site Reliability Engineer, Colorado Springs

Onebrief • Colorado Springs (CO)

On-site
USD 140,000 - 190,000
Relocation assistance
On-site in Colorado Springs
Senior Site Reliability Engineer, Application Reliability (Arlington, VA) - Secret Clearance Required - Relocation Provided
Senior Site Reliability Engineer, Application Reliability (Arlington, VA) - Secret Clearance Required - Relocation Provided

InvestedintheMission • United States

On-site
USD 120,000 - 180,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Onebrief • Arlington (VA)

On-site
USD 190,000 - 240,000
Relocation assistance
Senior Site Reliability Engineer (Arlington, VA) - Secret Clearance Required - Relocation Provided
Senior Site Reliability Engineer (Arlington, VA) - Secret Clearance Required - Relocation Provided

Onebrief • Arlington (VA)

On-site
USD 140,000 - 200,000
Relocation assistance
Senior SRE - Secret Clearance, On-Prem & Cloud
Senior SRE - Secret Clearance, On-Prem & Cloud

Onebrief • Arlington (VA), Northern (KY)

Hybrid
USD 140,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

United States Digital Space LLC • United States

On-site
USD 145,000 - 200,000
Remote-first
Senior SRE: On-Prem & Cloud Reliability for DoD Ops
Senior SRE: On-Prem & Cloud Reliability for DoD Ops

Onebrief • United States

Remote
USD 140,000 - 190,000
Relocation assistance
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

On-site
USD 210,000 - 230,000