Remote SRE for High-Impact AI Platform

Vannevar Labs

San Diego (CA)

Hybrid

USD 140,000 - 200,000

Full time

28 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Unlimited PTO
Remote-friendly culture
WeWork access
Lifestyle stipend
Parental leave
Child care reimbursement
Pet care reimbursement

Job summary

Vannevar Labs is seeking a Site Reliability Engineer to own platform reliability, health, and deployment automation. You will monitor dashboards, handle incidents end-to-end, and build observability tooling to mature SRE practices.

The role emphasizes careful judgment, clear communication, and collaboration across teams as the system scales for high-stakes operations. The ideal candidate has 5+ years in SRE/DevOps, strong scripting, IaC experience (Terraform/Pulumi), and is comfortable working

Qualifications

  • 5+ years of experience in SRE, DevOps, or software engineering.
  • Hands-on experience monitoring production systems and responding to incidents.
  • Excellent communication skills, especially during incidents and day-to-day work.
  • Experience with the PLG stack, Datadog, or other enterprise monitoring/observability tools.
  • Experience with AWS cloud technologies.
  • Familiarity with infrastructure-as-code technologies such as Terraform and Pulumi.
  • Experience with Python, Bash, or other scripting languages.
  • Experience working in an agile scrum environment, with the ability to work independently.
  • Able to quickly learn new and existing technologies.
  • Willingness and ability to work on-site in San Diego, CA.
  • U.S. Citizenship status is required, as this position requires access to U.S.-only data systems and export-controlled data.
  • TS/SCI Clearance required.
  • Experience defining and tracking SLOs/SLIs and error budgets.
  • Experience crafting CI/CD processes and automation.
  • Proficient with containerization technologies like Docker.
  • Experience working in AWS GovCloud.
  • Experience with modern web services architectures.
  • Experience with relational database systems, including SQL and relational design.
  • Experience working with Elasticsearch/OpenSearch.
  • Strong collaboration and negotiation skills for cross-functional projects.

Responsibilities

  • Monitor dashboards and system telemetry to detect health issues, performance degradation, and reliability risks.
  • Own the debugging and incident response process end to end, exercising judgment on deeper investigation or escalation.
  • Build logging, monitoring, and observability tooling to visualize the state of the platform and mature SRE practices.
  • Develop, maintain, and be responsible for overall platform health, scaling, and capacity planning.
  • Understand and help improve the deployment process, and automate build & deployment pipelines.
  • Identify bottlenecks in engineering workflows and drive improvements for speed and reliability.
  • Develop self-service tools and automation to improve engineering efficiency.
  • Play a critical part in implementing a secure, robust, high-availability delivery pipeline.
  • Communicate system status, trade-offs, and post-incident learnings clearly with teammates and stakeholders.

Skills

SRE experience
Incident response
AWS cloud
Terraform
Pulumi
Docker
CI/CD automation
Python scripting
On-call rotations
SLOs/SLIs
SQL
Relational DB
Elasticsearch
OpenSearch
Communication

Tools

Datadog
Elasticsearch/OpenSearch
AWS GovCloud
Docker
Terraform
Pulumi
CI/CD tools

Job description

Vannevar Labs is seeking a Site Reliability Engineer to own platform reliability, health, and deployment automation. You will monitor dashboards, handle incidents end-to-end, and build observability tooling to mature SRE practices.

The role emphasizes careful judgment, clear communication, and collaboration across teams as the system scales for high-stakes operations. The ideal candidate has 5+ years in SRE/DevOps, strong scripting, IaC experience (Terraform/Pulumi), and is comfortable working

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — Remote-First, High Impact
Site Reliability Engineer — Remote-First, High Impact

Vannevarlabs • San Diego (CA)

On-site
USD 140,000 - 190,000
Health insurance
Remote-first culture
Unlimited PTO
+4
Remote SRE: AI Platform Reliability & Automation
Remote SRE: AI Platform Reliability & Automation

Runpod • United States

On-site
USD 150,000 - 200,000
Remote work first
Competitive base salary
Stock options equity
+2
Senior SRE: Scale Resilient AI Platforms & Automation
Senior SRE: Scale Resilient AI Platforms & Automation

Relx Plc • Philadelphia

Hybrid
USD 95,000 - 159,000
Site Reliability Engineer
Site Reliability Engineer

Vannevarlabs • San Diego (CA)

On-site
USD 140,000 - 190,000
Health insurance
Remote-first culture
Unlimited PTO
+4
Senior SRE: Platform Reliability & AI-Driven Ops
Senior SRE: Platform Reliability & AI-Driven Ops

Block • New York (NY)

On-site
USD 170,100 - 283,600
Healthcare coverage
Retirement plans
Employee Stock Purchase Program
+1
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

IntraEdge • Austin (TX)

On-site
USD 120,000 - 180,000
Senior SRE: Reliability, Automation & AI Platforms
Senior SRE: Reliability, Automation & AI Platforms

RX Brasil • Philadelphia

On-site
USD 95,000 - 159,000
Annual incentive bonus