Systems & Reliability Engineer (SRE / Platform / DevOps) Full-Time, Bay Area, Hybrid See Job Details

Getpeer

San Francisco, Northern (CA, KY)

Hybrid

USD 190,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Meaningful equity grant
Remote-first with quarterly San Francs
Health coverage
12 weeks parental leave
Home-office stipend
Unlimited PTO

Job summary

Peer AI is seeking a Senior / Staff Systems & Reliability Engineer to own the release path from merge to production and to establish a lightweight on-call and incident-response framework. The role focuses on improving observability, CI/CD, and AI-assisted incident analysis within a hybrid Bay Area environment.

The team builds an AI-native platform for pharma workflows, with on-sites in San Francisco and a remote-friendly culture. Equity and parental leave are offered.

Qualifications

  • Strong software engineering fundamentals; you should be comfortable solving problems in application code as well as infrastructure.
  • Experience operating distributed production systems and debugging failures across multiple services.
  • Hands-on experience with AWS, containers, CI/CD, infrastructure automation, and modern observability tooling.
  • Experience with production on-call, incident response, deployment strategy, and safe rollback or recovery patterns.
  • A bias toward removing toil through software, with strong judgment about where process adds safety and where it simply adds bureaucracy.

Responsibilities

  • Own and evolve the path from merge to production, including release coordination, deployment safety, rollback, and release-health signals.
  • Establish and run a lightweight, effective engineering on-call and incident-response process, with clear ownership and useful escalation paths.
  • Improve observability across services, asynchronous jobs, and complete customer workflows, and define reliability signals that reflect what users actually experience.
  • Build and improve CI/CD, automated deployment checks, feature rollout controls, environment reliability, and developer tooling.
  • Use AI to accelerate incident triage, log and trace analysis, root-cause investigation, runbook execution, and post-release analysis where it is genuinely useful.
  • Lead incident reviews focused on systemic fixes, automate recurring remediation, and partner with Quality and Security as we move toward GxP with reproducible change control, release evidence, auditability, and production monitoring.

Skills

Strong software engineering
Distributed systems
On-call experience
Safety-conscious coding
Toil reduction via software

Tools

AWS
Containers
CI/CD
Infrastructure automation
Observability tooling

Job description

Systems & Reliability Engineer (SRE / Platform / DevOps)

Full-Time, Bay Area, Hybrid


Location: Bay Area / Hybrid


Type: Full-Time


Level: Senior / Staff


Why this role is exciting

We have built a culture that can ship very quickly. Your job is to make that speed sustainable: releases that are observable and reversible, an on-call process engineers trust, and production systems that surface customer impact early. This role spans release engineering, SRE, platform, production operations, and developer experience. The goal is not to create more process; it is to turn recurring operational work into software and make the safe path the easy path.


About Peer AI

We're a Bay Area-founded, remote-friendly company backed by investors with deep roots in both technology and life sciences. Our platform is live with top-tier pharma customers today, and we're growing fast in a $23B addressable market spanning document authoring, submission workflow management, and agency interactions with the FDA and EMA.


Our Vision

At Peer AI, we are working to clear the path for important scientific and medical discoveries, so that treatments reach the patients who need them as quickly as possible.


Our Values


  • Drive Impact: We focus on delivering real results for our users. Making their work easier, better, and more impactful.


  • Be the Expert: We lead with deep expertise, curiosity, and honesty to guide others toward the best outcomes.


  • Go for Great: We take pride in pushing past “good enough” to deliver standout work in every detail.


  • Win as a Team: We succeed together—building trust, sharing ownership, and helping each other grow every day.



How We Work

Peer AI is an AI-native company, and we want our engineering organization to be AI-native too. We use AI throughout how we build, test, analyze, debug, and operate software. We look for people who ask what can be automated, what should become a system instead of a recurring task, and where human judgment creates the most value. These roles are intentionally broader than their traditional equivalents because we want the people who join us to help redefine the function itself.


KeyResponsibilities


  • Own and evolve the path from merge to production, including release coordination, deployment safety, rollback, and release-health signals.


  • Establish and run a lightweight, effective engineering on-call and incident-response process, with clear ownership and useful escalation paths.


  • Improve observability across services, asynchronous jobs, and complete customer workflows, and define reliability signals that reflect what users actually experience.


  • Build and improve CI/CD, automated deployment checks, feature rollout controls, environment reliability, and developer tooling.


  • Use AI to accelerate incident triage, log and trace analysis, root-cause investigation, runbook execution, and post-release analysis where it is genuinely useful.


  • Lead incident reviews focused on systemic fixes, automate recurring remediation, and partner with Quality and Security as we move toward GxP with reproducible change control, release evidence, auditability, and production monitoring.



Must‑HaveQualifications


  • Strong software engineering fundamentals; you should be comfortable solving problems in application code as well as infrastructure.


  • Experience operating distributed production systems and debugging failures across multiple services.


  • Hands‑on experience with AWS, containers, CI/CD, infrastructure automation, and modern observability tooling.


  • Experience with production on‑call, incident response, deployment strategy, and safe rollback or recovery patterns.


  • A bias toward removing toil through software, with strong judgment about where process adds safety and where it simply adds bureaucracy.



Nice‑to‑Have


  • Experience with AWS ECS/Fargate, SQS, PostgreSQL, CloudWatch, or similar cloud‑native stacks.


  • Experience using LLMs or agents for developer productivity, incident response, or production operations.


  • Experience in a regulated, security‑sensitive, or otherwise high‑reliability software environment.



WhatWeOffer


  • Meaningful equity grant in a high‑growth, venture‑backed company.


  • Remote‑first setup with quarterly on‑sites in San Francisco.


  • Comprehensive health coverage, 12 weeks paid parental leave, and monthly home‑office stipends.


  • Unlimited PTO and a mission that directly accelerates life‑saving clinical research.



Equal Opportunity

Peer AI is an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other characteristic protected by applicable law. If you need a reasonable accommodation during the hiring process, please let us know.


Hiring Manager

Head of Engineering

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Product Data Scientist (Product Analytics / ML)
Product Data Scientist (Product Analytics / ML)

Getpeer • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 210,000
Equity grant
Remote-first with SF on-sites
Health coverage
+3
Product Quality Engineer (SDET / Test Automation) Full-Time, Bay Area, Hybrid See Job Details
Product Quality Engineer (SDET / Test Automation) Full-Time, Bay Area, Hybrid See Job Details

Getpeer • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Equity grant
Remote-first with quarterly on-sites
Health coverage
+3
Full Stack Software Engineer
Full Stack Software Engineer

Getpeer • San Francisco (CA), Northern (KY)

On-site
USD 180,000 - 240,000
Equity grant
Remote-first with San Francisco on-sit
Health coverage
+3
Technical Pre-Sales Engineer Full-Time, Remote See Job Details
Technical Pre-Sales Engineer Full-Time, Remote See Job Details

Peer TechBio Inc. • San Francisco (CA)

Hybrid
USD 120,000 - 180,000
Equity grant
Remote-first
On-sites in SF
+4
Software Engineer, Applications (App Foundations) (High Seniority)
Software Engineer, Applications (App Foundations) (High Seniority)

F-Prime Capital • San Francisco (CA)

On-site
USD 150,000 - 190,000
Hybrid work model
On-site 3 days per week
Staff Engineer
Staff Engineer

Mira Mace • San Francisco (CA)

On-site
USD 180,000 - 240,000
DevOps Engineer
DevOps Engineer

Prudentia Sciences • San Francisco (CA)

On-site
USD 180,000 - 200,000
DevOps Engineer
DevOps Engineer

Prudentia Sciences • Cambridge (MA)

On-site
USD 180,000 - 200,000
Competitive salary and equity
Hybrid onsite 2 days/week (Boston, NYC
San Francisco hub
Platform Engineer
Platform Engineer

Harper • San Francisco (CA)

On-site
USD 140,000 - 280,000
Uber commuter benefits
Meals provided (breakfast, lunch, and/
Snacks, drinks and coffee daily
+2
Senior Software Engineer (Full Stack)
Senior Software Engineer (Full Stack)

Prudentia Sciences • United States

On-site
USD 153,000 - 235,000
Impact on pharma tech
Ownership and leadership
Growth opportunities
+2