Forward Deployed Site Reliability Engineer (TS/SCI Required)

Twenty Technologies

Chantilly

Sur place

EUR 124 000 - 169 000

Plein temps

14 jours+
Générateur de candidature

Une candidature sur mesure pour ce poste — un CV personnalisé et une lettre de motivation qui correspondent directement à l’offre.

Passez les filtres ATS

Avantages offerts par ce poste

Health, dental, and vision plans
Life/AD&D and disability coverage
Parental leave

Résumé du poste

Twenty Technologies seeks an on-site Reliability Engineer to ensure the platform operates reliably within a restricted, air-gapped AWS environment at a government customer site.

You will define how we measure reliability, lead incident response with autonomy, and act as the primary technical link between the site and engineering in Arlington.

Qualifications

  • 5+ years in site reliability engineering or production operations.
  • Proven experience defining and tracking SLIs, SLOs, and error budgets in production.
  • Hands-on Docker, Docker Compose, and AWS in production deployments.
  • Solid Linux/Unix systems administration in constrained environments.
  • Experience with Terraform within guardrails.
  • Experience with the LGTM observability stack or Grafana/Loki/Prometheus/Mimir.
  • Strong incident response experience with post-mortems and runbooks.
  • Scripting with Python or Bash; Go is a plus.

Responsabilités

  • Define, track, and report on SLIs and SLOs for platform services in the customer environment.
  • Use error budgets to guide reliability discussions with Arlington engineering.
  • Identify and eliminate toil; automate repetitive tasks within a secure environment.
  • Conduct post-incident reviews, root cause analysis, and durable fixes with engineering.
  • Own the on-site observability posture — dashboards, alerts, log pipelines.
  • Lead on-site incident response: triage, containment, and customer communication.
  • Maintain and improve runbooks and emergency response procedures.
  • Serve as the on-site technical interface between the government customer and Twenty's team.

Connaissances

Incident response
Automation
Linux/Unix
Telemetry & observability

Formation

TS/SCI security clearance

Outils

Docker
Docker Compose
AWS
Grafana
Loki
Tempo
Mimir
Terraform
Python
Bash

Description du poste

About the Company

America is under sustained cyber attack. Our adversaries infiltrate our networks, steal our IP, and degrade the digital infrastructure that modern life runs on. They’ve learned—correctly—that those attacks rarely produce consequences.

Twenty was founded to change that, by making our adversaries think twice before they attack us. Our vision is American and allied primacy in cyberspace—a future where they cannot contest us, deterrence is assured, and the free world remains secure.

Founded in 2024, Twenty Technologies (www.twenty.io) industrializes offensive cyber operations for the U.S. and its allies. Headquartered in Arlington, Virginia, Twenty has raised $168M from Khosla Ventures, Accel, Caffeinated Capital, Friends & Family Capital, Point72 Ventures, General Catalyst, and In-Q-Tel.

Mission | On Site | Full Time | Active TS/SCI with Full Scope Polygraph

Role Summary

You'll be our eyes, ears, and hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical platform running in a restricted, air-gapped AWS environment. This role sits at the intersection of deep technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident response in a constrained environment, and serve as the primary technical link between what's happening on-site and the engineering team back in Arlington. You'll work closely with the DevSecOps engineer to ensure the platform operates within government security and compliance requirements, and with product engineers to translate operational reality into actionable feedback. You'll report directly to the VP of Engineering. If you thrive operating with autonomy in high-stakes environments and find satisfaction in making complex systems provably reliable, this role is for you.

What You'll Do
Reliability Engineering
  • Define, track, and report on SLIs and SLOs for platform services running in the customer environment.

  • Use error budgets to drive reliability conversations with the Arlington engineering team, translating operational data into prioritized engineering work.

  • Identify and eliminate toil: build automation for repetitive operational tasks within the constraints of the secure environment.

  • Conduct post-incident reviews, own root cause analysis, and drive durable fixes in partnership with the engineering team.

Observability & Incident Response
  • Own the observability posture for the on-site deployment — dashboards, alerting thresholds, and log pipelines using the LGTM stack (Grafana, Loki, Tempo, Mimir).

  • Lead incident response on-site: triage, containment, coordination with Arlington, and customer communication.

  • Maintain and continuously improve runbooks for operational procedures and emergency response protocols.

  • Serve as the on-call anchor for the customer environment, with clear escalation paths to the engineering team.

Deployment & Infrastructure Operations
  • Work with the customer deployment team to get Twenty's platform stood up and updated within the restricted environment.

  • Manage containerized services (Docker, Docker Compose) across deployment lifecycle — configuration, updates, rollbacks.

  • Apply and validate Terraform-based infrastructure changes within the enclave, in coordination with the DSO engineer who owns IaC policy and guardrails.

  • Perform capacity planning and flag scaling requirements to the Arlington team before they become incidents.

Customer Liaison & Engineering Feedback
  • Serve as the primary technical interface between the government customer and Twenty's engineering team — translating operational requirements, constraints, and issues in both directions.

  • Represent the operational environment accurately in engineering discussions: what the team in Arlington can't see, you make visible.

  • Partner with the DevSecOps engineer on compliance, logging, and audit requirements specific to the customer environment.

  • Provide technical guidance and support to customer stakeholders on system behavior and troubleshooting procedures.

Who You Are
  • You own reliability outcomes, not just uptime dashboards — you define what "healthy" means and hold the system to it.

  • You're as comfortable writing a runbook as you are deep in a production incident with limited tooling and no safety net.

  • You operate well with minimal remote support — ambiguity doesn't paralyze you, and you know when to elevate versus when to solve it yourself.

  • You build trust naturally with external stakeholders, including government customers, and can translate complex technical situations into plain language under pressure.

  • You treat toil as a bug: if you're doing something manually more than twice, you automate it.

  • You communicate with precision — your incident reports and runbooks are read by people who weren't in the room, and they need to be right.

  • You understand that in a restricted environment, you are the feedback loop — and you take that responsibility seriously.

Must Have
  • 5+ years of professional experience in site reliability engineering, production operations, or a closely related infrastructure role.

  • Proven experience defining and tracking SLIs, SLOs, and error budgets in a production environment.

  • Hands-on experience with Docker, Docker Compose, and AWS (EC2, ECS, RDS, VPCs, security groups) in production deployments.

  • Solid Linux/Unix systems administration skills; productive in constrained environments where GUI tooling may be limited or unavailable.

  • Experience with Terraform for infrastructure provisioning and configuration, working within DSO-provided policy guardrails.

  • Experience with the LGTM observability stack or equivalent (Grafana, Loki, Prometheus/Mimir, distributed tracing).

  • Strong incident response experience: you've led responses, written post-mortems and runbooks, and shipped the preventive fix.

  • Scripting proficiency in Python or Bash for operational automation, with familiarity in Go a plus; experience with PagerDuty or equivalent on-call tooling.

  • Experience working in or directly supporting government or defense environments, including air-gapped or enclave deployments.

Nice to Have
  • Experience with NATS or similar pub/sub messaging systems in production.

  • Background in cyber operations, intelligence systems, or signals environments.

  • AWS certifications (Solutions Architect, SysOps, or DevOps Engineer).

Security Requirements
  • Must possess and be able to maintain a TS/SCI security clearance with appropriate polygraph

  • U.S. citizenship required

  • Willingness to travel occasionally for customer engagements and operational support

Benefits
What's on the table:

  • Health. Medical, dental, and vision plan options. Life / AD&D, disability coverage options.

  • Family. Paid parental leave for eligible full-time employees. 12 weeks for birthing parents, 4 for non-birthing parents, 6 weeks for adoptive, foster, or intended parents through surrogacy.

  • Vacation. Paid holidays and flexible PTO. Take what you need.

  • Retirement. 401(k) with pre-tax and Roth options. HSA/FSA options, dependent care FSA.

Benefits vary by location, role, and eligibility. Full plan details provided during the interview and offer process.

Due to U.S. government contract and security requirements, this role is limited to U.S. citizens. Some positions may also require eligibility to obtain and maintain a U.S. Government security clearance. Any active clearance requirement will be listed in the role description.

Twenty is an equal opportunity employer. We consider all qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, veteran status, disability, or any other protected status, consistent with applicable law.

If you need a reasonable accommodation during the hiring process, let us know and we will work with you.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Staff Data Engineer - TS/SCI Cleared
Staff Data Engineer - TS/SCI Cleared

Twenty Technologies • Chantilly

Sur place
EUR 106 000 - 142 000
Health insurance
Paid parental leave
Flexible PTO
+1
Senior/Staff Forward Deployed Software Engineer - TS/SCI Cleared
Senior/Staff Forward Deployed Software Engineer - TS/SCI Cleared

Twenty Technologies • Chantilly

Sur place
EUR 130 000 - 165 000
Health benefits
Parental leave
Flexible PTO
+1
Senior Forward Deployed Analyst
Senior Forward Deployed Analyst

Twenty Technologies • Chantilly

Sur place
EUR 124 000 - 169 000
Health plan
Dental & vision
Life/AD&D
+3
Mission Deployment Lead
Mission Deployment Lead

Twenty Technologies • Chantilly

Sur place
EUR 130 000 - 165 000
Health plan
Parental leave
Paid holidays
+1
Senior On-Site Software Engineer for Cyber Operations
Senior On-Site Software Engineer for Cyber Operations

Twenty Technologies • Chantilly

Sur place
EUR 130 000 - 165 000
Health benefits
Parental leave
Flexible PTO
+1
Mission Deployment Lead — IC Cyber Ops Enablement
Mission Deployment Lead — IC Cyber Ops Enablement

Twenty Technologies • Chantilly

Sur place
EUR 130 000 - 165 000
Health plan
Parental leave
Paid holidays
+1
Software Systems Engineer w/Secret Clearance
Software Systems Engineer w/Secret Clearance

TekSynap • Les Ulis

À distance
EUR 79 000 - 123 000
Network Intrusion Specialist
Network Intrusion Specialist

LV8D Solutions, LLC • Chantilly

Sur place
EUR 105 000 - 167 000
Performance bonuses
Retention incentives
Referral bonuses
+2
Incident Operations Lead
Incident Operations Lead

Jobgether • France

Sur place
EUR 120 000 - 190 000
Stock options
Health benefits
One-time USD 500 home-office setup
+1
Infrastructure DevOps Lead (TS/SCI with Poly Required)
Infrastructure DevOps Lead (TS/SCI with Poly Required)

GCI • Chantilly

Sur place
EUR 164 000 - 279 000