Twenty is looking for a highly skilled Site Reliability Engineer to join their team in Arlington, VA. You will ensure the reliability of mission-critical platforms in a secure AWS environment, leading incident responses and collaborating with internal teams. The ideal candidate has over 5 years of experience in reliability engineering, strong knowledge of AWS and Docker, and exceptional incident response skills. Full benefits, including flexible PTO and health coverage, are provided.
Qualifications
5+ years of professional experience in site reliability engineering or closely related role.
Proven experience defining and tracking SLIs, SLOs, and error budgets.
Hands-on experience with Docker and AWS in production deployments.
Solid Linux/Unix systems administration skills.
Experience with Terraform within policy guardrails.
Strong incident response experience including writing post-mortems and runbooks.
Responsibilities
Define, track, and report on SLIs and SLOs for platform services.
Lead incident response on-site, handling triage and coordination.
Own the observability posture, including dashboards and alerting.
Manage containerized services across the deployment lifecycle.
Serve as primary technical interface between customer and engineering team.
Skills
Site reliability engineering
Docker
AWS
Incident response
Scripting in Python or Bash
Linux/Unix administration
Tools
Terraform
LGTM stack (Grafana, Loki, Mimir)
Job description
Twenty is looking for a highly skilled Site Reliability Engineer to join their team in Arlington, VA. You will ensure the reliability of mission-critical platforms in a secure AWS environment, leading incident responses and collaborating with internal teams. The ideal candidate has over 5 years of experience in reliability engineering, strong knowledge of AWS and Docker, and exceptional incident response skills. Full benefits, including flexible PTO and health coverage, are provided.