Site Reliability Engineer

Vannevar Labs

San Diego (CA)

Hybrid

USD 140,000 - 200,000

Full time

44 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Unlimited PTO
Remote-friendly culture
WeWork access
Lifestyle stipend
Parental leave
Child care reimbursement
Pet care reimbursement

Job summary

Vannevar Labs is seeking a Site Reliability Engineer to own platform reliability, health, and deployment automation. You will monitor dashboards, handle incidents end-to-end, and build observability tooling to mature SRE practices.

The role emphasizes careful judgment, clear communication, and collaboration across teams as the system scales for high-stakes operations. The ideal candidate has 5+ years in SRE/DevOps, strong scripting, IaC experience (Terraform/Pulumi), and is comfortable working

Qualifications

  • 5+ years of experience in SRE, DevOps, or software engineering.
  • Hands-on experience monitoring production systems and responding to incidents.
  • Excellent communication skills, especially during incidents and day-to-day work.
  • Experience with the PLG stack, Datadog, or other enterprise monitoring/observability tools.
  • Experience with AWS cloud technologies.
  • Familiarity with infrastructure-as-code technologies such as Terraform and Pulumi.
  • Experience with Python, Bash, or other scripting languages.
  • Experience working in an agile scrum environment, with the ability to work independently.
  • Able to quickly learn new and existing technologies.
  • Willingness and ability to work on-site in San Diego, CA.
  • U.S. Citizenship status is required, as this position requires access to U.S.-only data systems and export-controlled data.
  • TS/SCI Clearance required.
  • Experience defining and tracking SLOs/SLIs and error budgets.
  • Experience crafting CI/CD processes and automation.
  • Proficient with containerization technologies like Docker.
  • Experience working in AWS GovCloud.
  • Experience with modern web services architectures.
  • Experience with relational database systems, including SQL and relational design.
  • Experience working with Elasticsearch/OpenSearch.
  • Strong collaboration and negotiation skills for cross-functional projects.

Responsibilities

  • Monitor dashboards and system telemetry to detect health issues, performance degradation, and reliability risks.
  • Own the debugging and incident response process end to end, exercising judgment on deeper investigation or escalation.
  • Build logging, monitoring, and observability tooling to visualize the state of the platform and mature SRE practices.
  • Develop, maintain, and be responsible for overall platform health, scaling, and capacity planning.
  • Understand and help improve the deployment process, and automate build & deployment pipelines.
  • Identify bottlenecks in engineering workflows and drive improvements for speed and reliability.
  • Develop self-service tools and automation to improve engineering efficiency.
  • Play a critical part in implementing a secure, robust, high-availability delivery pipeline.
  • Communicate system status, trade-offs, and post-incident learnings clearly with teammates and stakeholders.

Skills

SRE experience
Incident response
AWS cloud
Terraform
Pulumi
Docker
CI/CD automation
Python scripting
On-call rotations
SLOs/SLIs
SQL
Relational DB
Elasticsearch
OpenSearch
Communication

Tools

Datadog
Elasticsearch/OpenSearch
AWS GovCloud
Docker
Terraform
Pulumi
CI/CD tools

Job description

Vannevar builds AI systems for the Department of War's most consequential missions. We have 125 deployments across every branch and combatant command spanning our core platform and five distinct products. An an example of how we work, when major combat operations started with Iran, we fielded a new product supporting 24/7 operations and 9,000+ users in three months.

How we build is our advantage: we deploy forward with the people who own the mission. Our engineers, product team, and CTO deploy forward to the point of friction, including visiting units in Ukraine. We focus on embedding directly with operational units supporting great power competition with China, combat operations with Iran, and counter-narcotics missions.

We are looking for an Site Reliability Engineer to own the reliability, health, and deployment automation of the platform at Vannevar Labs. In this role you'll be the person watching the system's pulse — monitoring dashboards, catching health issues before they become incidents, and owning the debugging process from first alert to resolution. Your decisions today will have a large impact on the company's future.

We believe that simple systems are easier to understand, maintain, and scale. You will be making trade-offs as you work to ensure that our systems are prepared to operate reliably in high-side environments at scale. A strong sense of judgment matters here: knowing when to dig deeper into a problem yourself and when to pull in the right people to pull in the right people to pull in the right people. Clear, calm communication — during an incident and in day-to-day work — is a must.

  • Monitor dashboards and system telemetry to detect health issues, performance degradation, and reliability risks — often before anyone else notices them.
  • Own the debugging and incident response process end to end, exercising good judgment about when to investigate more deeply and when to elevate.
  • Build logging, monitoring, and observability tooling to visualize the state of the platform and continuously mature our SRE practices.
  • Develop, maintain, and be responsible for overall platform health, scaling, and capacity planning.
  • Understand and help improve the deployment process, and automate build & deployment pipelines.
  • Identify bottlenecks in engineering workflows and drive improvements that make the whole team faster and more reliable.
  • Develop self-service tools and automation to improve engineering efficiency.
  • Play a critical part in implementing a secure, robust, high-availability delivery pipeline.
  • Communicate system status, trade-offs, and post-incident learnings clearly with teammates and stakeholders.
  • 5+ years of experience in SRE, DevOps, or software engineering.
  • Hands-on experience monitoring production systems and responding to incidents — comfortable owning a debugging process and making the call on when to dig in versus pull in the right people.
  • Excellent communication skills, especially the ability to stay clear and organized while troubleshooting live issues.
  • Experience with the PLG stack, Datadog, or other enterprise monitoring/observability tools.
  • Experience participating in an on-call rotation and running or contributing to post-mortems.
  • Knowledge of AWS cloud technologies.
  • Familiarity with infrastructure-as-code technologies such as Terraform and Pulumi.
  • Experience with Python, Bash, or other scripting languages.
  • Experience working in an agile scrum environment, with the ability to work independently.
  • Able to quickly learn new and existing technologies.
  • Strong attention to detail and analytical capabilities.
  • Willingness and ability to work on-site in San Diego, CA.
  • U.S. Citizenship status is required, as this position requires the ability to access U.S.-only data systems and export-controlled data.
  • TS/SCI Clearance required.
  • Experience defining and tracking SLOs/SLIs and error budgets.
  • Experience crafting CI/CD processes and automation.
  • Proficient with containerization technologies like Docker.
  • Experience working in AWS GovCloud.
  • Experience with modern web services architectures.
  • Experience with relational database systems, including SQL and relational design.
  • Experience working with Elasticsearch/OpenSearch.
  • Strong collaboration and negotiation skills, with the ability to work on cross-functional projects with internal partner engineering teams.
  • Health, dental, and vision insurance
  • 100% remote first culture. You can work from anywhere in the US and all full time employees have WeWork access
  • Unlimited PTO including competitive vacation and holiday schedules
  • Lifestyle stipends - Monthly mental health, wellness & fitness stipend, in-home office setup stipend and family planning assistance
  • Salary top-up during military reserve duty
  • Fully paid parental leave
  • Child and pet care reimbursement during travel

We encourage candidates from all backgrounds to apply, even if you don't feel like you're a perfect fit. If you're passionate about contributing to our mission, we'd love to hear from you!

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Vannevarlabs • San Diego (CA)

On-site
USD 140,000 - 190,000
Health insurance
Remote-first culture
Unlimited PTO
+4
Senior Full-Stack Engineer - Product (TS/SCI Required)
Senior Full-Stack Engineer - Product (TS/SCI Required)

Vannevar • New York (NY)

On-site
USD 150,000 - 210,000
Health, dental, and vision insurance
Remote friendly with WeWork access
Unlimited PTO and holidays off
+5
Senior Full-Stack Engineer - Product (TS/SCI Required) Software
Senior Full-Stack Engineer - Product (TS/SCI Required) Software

Front Door Defense • New York (NY), Northern (KY)

On-site
USD 150,000 - 190,000
Health, dental, and vision insurance
Remote friendly with WeWork access
401(k) match
+2
Forward Deployed Engineer (TS/SCI Clearance Required)
Forward Deployed Engineer (TS/SCI Clearance Required)

Vannevar • United States

On-site
USD 135,000 - 205,000
Health, dental, and vision insurance
Unlimited PTO
401(k) match
+1
Mission Lead - Tampa (Clearance Required)
Mission Lead - Tampa (Clearance Required)

Vannevar Labs • Tampa (FL)

Hybrid
USD 130,000 - 170,000
Health insurance
Remote-friendly
401(k) match
+2
Senior Mission Manager (Clearance Required)
Senior Mission Manager (Clearance Required)

InvestedintheMission • Charlotte (NC)

On-site
USD 180,000 - 220,000
Health, dental and vision insurance
Remote-friendly with WeWork access
Unlimited PTO
+3
Senior Full-Stack Engineer - Product (TS/SCI Required)
Senior Full-Stack Engineer - Product (TS/SCI Required)

Vannevarlabs • New York (NY)

On-site
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+7
Forward Deployed Engineer - Oahu (TS/SCI Clearance Required)
Forward Deployed Engineer - Oahu (TS/SCI Clearance Required)

Vannevar • Honolulu (HI)

On-site
USD 135,000 - 205,000
Health, dental, and vision insurance
Unlimited PTO
401(k) match
+1
Senior Full-Stack Engineer - Product (TS/SCI Required)
Senior Full-Stack Engineer - Product (TS/SCI Required)

InvestedintheMission • New York (NY)

On-site
USD 170,000 - 230,000
Health Insurance
Dental Insurance
Vision Insurance
+7
Senior Mission Manager - International (Clearance Required)
Senior Mission Manager - International (Clearance Required)

Vannevar Labs • Washington

Hybrid
USD 180,000 - 220,000
Health, dental, and vision insurance
Remote friendly with WeWork access
Unlimited PTO
+2