Systems & Reliability Engineer (SRE / Platform / DevOps)

Getpeer

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity grant
Remote-first with SF on-sites
Health coverage
12 weeks parental leave
Unlimited PTO
Home-office stipend

Job summary

Peer AI in the Bay Area is seeking a Systems & Reliability Engineer to own the path from merge to production, improve release safety, and automate recurring work. The role spans release engineering, SRE, platform, and developer experience with an AI-native approach.

The candidate will drive CI/CD improvements, resilient deployment strategies, and post-release analyses while aligning with security and regulatory expectations. Hybrid role with quarterly on-sites in San Francisco.

Qualifications

  • Strong software engineering fundamentals applied to both app and infra.
  • Experience operating distributed production systems across services.
  • Hands-on with AWS, containers, CI/CD, infrastructure automation, observability.

Responsibilities

  • Own and evolve path from merge to production, including release coordination and rollback.
  • Establish lightweight on-call and incident-response with clear ownership.
  • Improve observability across services and customer workflows.
  • Build and improve CI/CD, deployment checks, and environment reliability.
  • Use AI to accelerate incident triage, log analysis, and runbook execution.
  • Lead incident reviews and drive systemic fixes with auditability and monitoring.

Skills

Software engineering fundamentals
Distributed systems
AWS / containers / CI/CD
On-call / incident response
Toil reduction via software

Tools

Observability tooling

Job description

Systems & Reliability Engineer (SRE / Platform / DevOps)

Full-Time, Bay Area, Hybrid

Location: Bay Area / Hybrid

Type: Full-Time

Level: Senior / Staff

Why this role is exciting

We have built a culture that can ship very quickly. Your job is to make that speed sustainable: releases that are observable and reversible, an on-call process engineers trust, and production systems that surface customer impact early. This role spans release engineering, SRE, platform, production operations, and developer experience. The goal is not to create more process; it is to turn recurring operational work into software and make the safe path the easy path.

About Peer AI

We're a Bay Area-founded, remote-friendly company backed by investors with deep roots in both technology and life sciences. Our platform is live with top-tier pharma customers today, and we're growing fast in a $23B addressable market spanning document authoring, submission workflow management, and agency interactions with the FDA and EMA.

Our Vision

At Peer AI, we are working to clear the path for important scientific and medical discoveries, so that treatments reach the patients who need them as quickly as possible.

Our Values
  • Drive Impact: We focus on delivering real results for our users. Making their work easier, better, and more impactful.
  • Be the Expert: We lead with deep expertise, curiosity, and honesty to guide others toward the best outcomes.
  • Go for Great: We take pride in pushing past "good enough" to deliver standout work in every detail.
  • Win as a Team: We succeed together—building trust, sharing ownership, and helping each other grow every day.
How We Work

Peer AI is an AI-native company, and we want our engineering organization to be AI-native too. We use AI throughout how we build, test, analyze, debug, and operate software. We look for people who ask what can be automated, what should become a system instead of a recurring task, and where human judgment creates the most value. These roles are intentionally broader than their traditional equivalents because we want the people who join us to help redefine the function itself.

KeyResponsibilities
  • Own and evolve the path from merge to production, including release coordination, deployment safety, rollback, and release-health signals.
  • Establish and run a lightweight, effective engineering on-call and incident-response process, with clear ownership and useful escalation paths.
  • Improve observability across services, asynchronous jobs, and complete customer workflows, and define reliability signals that reflect what users actually experience.
  • Build and improve CI/CD, automated deployment checks, feature rollout controls, environment reliability, and developer tooling.
  • Use AI to accelerate incident triage, log and trace analysis, root-cause investigation, runbook execution, and post-release analysis where it is genuinely useful.
  • Lead incident reviews focused on systemic fixes, automate recurring remediation, and partner with Quality and Security as we move toward GxP with reproducible change control, release evidence, auditability, and production monitoring.
Must‑HaveQualifications
  • Strong software engineering fundamentals; you should be comfortable solving problems in application code as well as infrastructure.
  • Experience operating distributed production systems and debugging failures across multiple services.
  • Hands-on experience with AWS, containers, CI/CD, infrastructure automation, and modern observability tooling.
  • Experience with production on-call, incident response, deployment strategy, and safe rollback or recovery patterns.
  • A bias toward removing toil through software, with strong judgment about where process adds safety and where it simply adds bureaucracy.
Nice‑to‑Have
  • Experience with AWS ECS/Fargate, SQS, PostgreSQL, CloudWatch, or similar cloud-native stacks.
  • Experience using LLMs or agents for developer productivity, incident response, or production operations.
  • Experience in a regulated, security-sensitive, or otherwise high-reliability software environment.
WhatWeOffer
  • Meaningful equity grant in a high-growth, venture-backed company.
  • Remote-first setup with quarterly on-sites in San Francisco.
  • Comprehensive health coverage, 12 weeks paid parental leave, and monthly home-office stipends.
  • Unlimited PTO and a mission that directly accelerates life-saving clinical research.
Equal Opportunity

Peer AI is an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other characteristic protected by applicable law. If you need a reasonable accommodation during the hiring process, please let us know.

Hiring Manager

Head of Engineering

Resources

Blog

Case studies

Events

Press

Trust center

Webinars

White papers

Company

About us

Careers

Contact us

Disclaimer

Privacy policy

Terms of use

© 2026 Peer AI. All rights reserved.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Full Stack Software Engineer
Full Stack Software Engineer

Getpeer • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity grant
Remote-first with San Francisco on-sit
Health coverage
+3
Product Quality Engineer (SDET / Test Automation)
Product Quality Engineer (SDET / Test Automation)

Getpeer • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Meaningful equity grant
Remote-first with quarterly on-sites (
Health coverage, parental leave, home‑
Product Data Scientist (Product Analytics / ML)
Product Data Scientist (Product Analytics / ML)

Getpeer • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Equity grant
Remote-first with SF on-sites
Health coverage
+3
Senior SRE / Platform Engineer — Remote, Unlimited PTO
Senior SRE / Platform Engineer — Remote, Unlimited PTO

Getpeer • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity grant
Remote-first with SF on-sites
Health coverage
+3
Senior DevOps Engineer
Senior DevOps Engineer

Qualified Health • Palo Alto (CA)

Hybrid
USD 170,000 - 220,000
Equity
Medical insurance
Dental insurance
+3
Senior Director, Quality Engineering
Senior Director, Quality Engineering

Springhealth66 • San Francisco (CA)

Hybrid
USD 230,000 - 250,000
Health benefits
401(k) matching
Noom program
+5
Sr. Software Engineer
Sr. Software Engineer

Intertech, Inc • San Francisco (CA)

On-site
USD 170,000 - 230,000
Founding Senior AI Engineer
Founding Senior AI Engineer

Peerbound • New York (NY)

On-site
USD 180,000 - 220,000
Meaningful equity
Medical vision dental
401K
+1
Senior DevOps Engineer
Senior DevOps Engineer

Transformcap • Palo Alto (CA)

Hybrid
USD 170,000 - 220,000
Equity
Medical insurance
Flexible hours
+1
Senior Distributed Systems SWE
Senior Distributed Systems SWE

Acceler8 Talent • San Francisco (CA)

Hybrid
USD 130,000 - 160,000