Senior Site Reliability Engineer, Application Reliability (Arlington, VA) - Secret Clearance Required - Relocation Provided

InvestedintheMission

United States

Hybrid

USD 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Onebrief is seeking a Site Reliability Engineer to join our Infrastructure & Security team. You will work with product engineers, fellow SREs, security, and customer success to ship reliable software that supports mission-critical deployments in DoD and AWS environments.

You’ll write code (TypeScript primarily), improve observability, define and measure SLIs/SLOs, and lead incident response with blameless postmortems, blending development and operations to keep systems secure and scalable.

Qualifications

  • Five plus years in software engineering, SRE, or a related role.
  • Strong TypeScript or comparable modern language experience.
  • Deep understanding of full SDLC and how reliability fits into each stage.
  • Experience with incident response and root-cause analysis.
  • Collaborative across product, platform, and DevOps teams.

Responsibilities

  • Fix reliability and performance in the codebase (TypeScript).
  • Design and run monitoring, logging, and alerting to tie to real application behavior.
  • Define and measure SLIs and SLOs; wire up corresponding alerting.
  • Lead incident response and post-mortems to identify root causes and fixes.
  • Automate repetitive operational work to reduce toil across environments.

Skills

TypeScript
SRE
CI/CD
Incident response
Blameless postmortems
Observability
Team collaboration
Security

Tools

Prometheus
Grafana
Loki
GitHub Actions
GitLab CI/CD
Jenkins
Terraform
Ansible
Kubernetes
AWS
Python
Go

Job description

Consequential Work. Dedicated People.
About Onebrief

Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination.

Military planning is complex by nature, requiring teams to coordinate information, people, and decisions across systems and locations. Onebrief brings planning, collaboration, simulation, and AI into one connected environment, helping teams test strategies, adapt to changing conditions, and make decisions with greater clarity when the stakes are real.

We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world.

Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth.

Security Clearance, Location, and Onsite Notice:

This is a hybrid role, requiring regular work on-site at customer locations in Arlington, VA - about 50/50 on-site vs remote.

If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance).

Active Secret Clearance required.

About The Role

We're hiring a Site Reliability Engineer to join our Infrastructure & Security team. You'll work closely with product engineers, fellow SREs, security, and customer success.

This is an SRE role for someone who's comfortable in application code. Much of the reliability and performance work happens in the codebase (primarily TypeScript), so you'll fix problems at the source rather than working around them in the infrastructure. You'll be a first line of support for our mission-critical deployments across on-prem DoD and AWS environments, and what you learn in the field will feed directly back into the product.

You'll ship code that makes Onebrief more stable, faster, and easier to deploy and operate. The work sits at the seam between engineering and operations, and it's weighted toward engineering.

About You

You treat reliability as a feature, not an afterthought, and you'd rather fix a problem in the code than route around it. You understand the full software development lifecycle (design, review, testing, release) and you know where reliability fits into each step.

You're comfortable reading and writing application code, and you're just as happy dropping into a kubectl shell to triage a production issue. You turn failure modes into guardrails, and you think monitoring, alerting, and clear runbooks are part of building software, not extra credit.

You mentor others and push a culture of blameless postmortems. You work naturally with product and platform teams, helping them move fast without breaking things by giving them the tools, tests, and observability that make quick recovery real.

What You'll Do

You'll help make our production application reliable, scalable, and secure by improving the software itself, not just the systems it runs on. Day to day that looks like:

  • Improving the application: Work directly in the codebase (primarily TypeScript) to fix reliability and performance problems at the source. You'll partner with product engineers on design decisions, review code with reliability and security in mind, and treat "make the app better" as a first-class part of the job rather than something you hand off.

  • Building observability that developers actually use: Design and run our monitoring, logging, and alerting (Prometheus, Loki, Alloy, Grafana). The goal is alerts and dashboards tied to real application behavior, so teams catch issues before users do.

  • Owning reliability targets: Define and measure SLIs and SLOs, wire up alerting that feeds them, and be the person who can say what "reliable" means for our systems and prove it with data.

  • Leading incident response: Act as incident responder, and incident commander when needed. Run blameless post-mortems (AARs) that find the actual root cause and turn it into a code or process fix so it doesn't happen again.

  • Automating away toil: Spot the repetitive operational work and write software to kill it. Share what works with other teams, including those running in air-gapped environments, and help them get production-ready.

What We Look For
  • An active Secret clearance

  • 5+ years in software engineering, SRE, or a related role, with real time spent writing and shipping application code

  • Strong TypeScript (or comparable modern language experience with willingness to work primarily in TypeScript)

  • Solid grasp of the full SDLC: design, code review, testing, release, and how reliability fits into each stage

  • Experience with incident response, root cause analysis, and turning findings into lasting fixes

  • A collaborator who works well across product, platform, and DevOps teams and shares context openly

Technical expertise
  • Application development in TypeScript (Node and/or a modern front-end framework)

  • CI/CD: building and maintaining pipelines (GitHub Actions, GitLab CI/CD, Jenkins)

  • Testing and quality practices as part of the delivery process

  • Comfort with at least one of Python, Go, or Bash for tooling and automation

  • Working knowledge of containers and Kubernetes (enough to debug and deploy, not necessarily to stand up clusters from scratch)

  • Networking fundamentals and secure configuration basics

Bonus points (nice to have)
  • Observability: Grafana stack, ELK, or Datadog

  • Infrastructure as Code (Terraform, Ansible) and cloud experience (AWS or AWS GovCloud)

  • Kubernetes cluster design and operations

  • Designing meaningful SLIs/SLOs with error budgets for distributed systems

  • GitOps practices and toolchains

  • DoD environments and compliance frameworks (RMF, STIGs, ICD 503)

  • Service mesh (Istio, Linkerd)

  • On-prem virtualization (VMware, Proxmox, Nutanix, Hyper-V)

  • Relevant certs (AWS DevOps Engineer, CKA/CKAD)

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (Arlington, VA) - Secret Clearance Required - Relocation Provided
Senior Site Reliability Engineer (Arlington, VA) - Secret Clearance Required - Relocation Provided

Onebrief • Arlington (VA)

On-site
USD 140,000 - 200,000
Relocation assistance
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Onebrief • Arlington (VA)

On-site
USD 190,000 - 240,000
Relocation assistance
Senior Site Reliability Engineer, Colorado Springs (Top Secret Clearance Required, Relocation Provided)
Senior Site Reliability Engineer, Colorado Springs (Top Secret Clearance Required, Relocation Provided)

Socket.dev • Colorado Springs (CO)

On-site
USD 140,000 - 200,000
Senior Site Reliability Engineer (Arlington, VA) - Relocation Provided
Senior Site Reliability Engineer (Arlington, VA) - Relocation Provided

Onebrief • Arlington (VA), Northern (KY)

Hybrid
USD 140,000 - 170,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Onebrief • Colorado Springs (CO)

On-site
USD 205,000 - 255,000
Relocation assistance
On-site customer deployments
Senior Site Reliability Engineer, Colorado Springs
Senior Site Reliability Engineer, Colorado Springs

Onebrief • Colorado Springs (CO)

On-site
USD 140,000 - 190,000
Relocation assistance
On-site in Colorado Springs
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

Hybrid
USD 210,000 - 230,000