Site Reliability Engineer — Remote-First, High Impact

Vannevarlabs

San Diego (CA)

On-site

USD 140,000 - 190,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
Remote-first culture
Unlimited PTO
Lifestyle stipend
Military reserve top-up
Parental leave
Travel care reimbursement

Job summary

Vannevar Labs seeks an experienced Site Reliability Engineer to own the reliability, health, and deployment automation of our platform. You will monitor dashboards, detect issues, and drive incident response from first alert to resolution, making prudent trade-offs to keep systems reliable at scale.

You will build observability tooling, manage capacity planning, improve deployment pipelines, and ensure secure, high-availability delivery.

Qualifications

  • 5+ years of experience in SRE, DevOps, or software engineering.
  • Hands-on experience monitoring production systems and responding to incidents.
  • Excellent communication skills, especially when troubleshooting live issues.
  • Experience with the PLG stack, Datadog, or other enterprise monitoring/observability tools.
  • Experience participating in an on-call rotation and post-mortems.
  • Knowledge of AWS cloud technologies.
  • Familiarity with Terraform and Pulumi.
  • Experience with Python, Bash, or other scripting languages.
  • Experience in an agile scrum environment.
  • Able to learn new technologies quickly.
  • Strong attention to detail and analytical capabilities.

Responsibilities

  • Monitor dashboards and system telemetry to detect health issues and reliability risks.
  • Own the debugging and incident response process end to end.
  • Build logging, monitoring, and observability tooling.
  • Develop and maintain overall platform health, scaling, and capacity planning.
  • Understand and improve deployment processes and automate pipelines.
  • Identify bottlenecks in engineering workflows and drive improvements.
  • Develop self-service tools to improve engineering efficiency.
  • Contribute to a secure, high-availability delivery pipeline.
  • Communicate system status and learnings clearly with teammates.

Skills

SRE/DevOps experience
Incident response
Communication
On-call experience
Python
Bash
Agile/Scrum
Learning agility
Analytical thinking

Tools

Datadog
Terraform
Pulumi
AWS
Docker
SQL
Elasticsearch/OpenSearch

Job description

Vannevar Labs seeks an experienced Site Reliability Engineer to own the reliability, health, and deployment automation of our platform. You will monitor dashboards, detect issues, and drive incident response from first alert to resolution, making prudent trade-offs to keep systems reliable at scale.

You will build observability tooling, manage capacity planning, improve deployment pipelines, and ensure secure, high-availability delivery.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote SRE for High-Impact AI Platform
Remote SRE for High-Impact AI Platform

Vannevar Labs • San Diego (CA)

Hybrid
USD 140,000 - 200,000
Health insurance
Dental insurance
Vision insurance
+7
Site Reliability Engineer
Site Reliability Engineer

Vannevarlabs • San Diego (CA)

On-site
USD 140,000 - 190,000
Health insurance
Remote-first culture
Unlimited PTO
+4
Site Reliability Engineer
Site Reliability Engineer

Vannevar Labs • San Diego (CA)

Hybrid
USD 140,000 - 200,000
Health insurance
Dental insurance
Vision insurance
+7
Senior Site Reliability Engineer — Remote, AWS & Observability
Senior Site Reliability Engineer — Remote, AWS & Observability

Prove • United States

Hybrid
USD 140,000 - 190,000
Wellbeing reimbursement
401k Match
Parental Leave Policy
+5
Remote DevOps/SRE Engineer – Site Reliability
Remote DevOps/SRE Engineer – Site Reliability

indexventures • United States

Remote
USD 100,000 - 150,000
Senior Site Reliability Engineer (Remote) – Scale & Reliability
Senior Site Reliability Engineer (Remote) – Scale & Reliability

WellSaid • United States

Remote
USD 140,000 - 190,000
Stock options
Medical, dental, and vision insurance
401(k) plan matching
+4
Senior Site Reliability Engineer: Scalable Infra & Observability
Senior Site Reliability Engineer: Scalable Infra & Observability

Early Warning • Chicago (IL)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Matching
Paid Time Off
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Senior Site Reliability Engineer — Remote • Flexible PTO
Senior Site Reliability Engineer — Remote • Flexible PTO

ACI Infotech • Seattle (WA), Northern (KY)

Hybrid
USD 120,000 - 170,000
Health insurance
Dental insurance
Vision insurance
+3
Site Reliability Engineer — Platform Observability & Autonomy
Site Reliability Engineer — Platform Observability & Autonomy

MaintainX • San Francisco (CA), Northern (KY)

Hybrid
USD 130,000 - 180,000