SRE & Platform Engineer — Automate Reliability & Releases

Getpeer

San Francisco, Northern (CA, KY)

Hybrid

USD 190,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Meaningful equity grant
Remote-first with quarterly San Francs
Health coverage
12 weeks parental leave
Home-office stipend
Unlimited PTO

Job summary

Peer AI is seeking a Senior / Staff Systems & Reliability Engineer to own the release path from merge to production and to establish a lightweight on-call and incident-response framework. The role focuses on improving observability, CI/CD, and AI-assisted incident analysis within a hybrid Bay Area environment.

The team builds an AI-native platform for pharma workflows, with on-sites in San Francisco and a remote-friendly culture. Equity and parental leave are offered.

Qualifications

  • Strong software engineering fundamentals; you should be comfortable solving problems in application code as well as infrastructure.
  • Experience operating distributed production systems and debugging failures across multiple services.
  • Hands-on experience with AWS, containers, CI/CD, infrastructure automation, and modern observability tooling.
  • Experience with production on-call, incident response, deployment strategy, and safe rollback or recovery patterns.
  • A bias toward removing toil through software, with strong judgment about where process adds safety and where it simply adds bureaucracy.

Responsibilities

  • Own and evolve the path from merge to production, including release coordination, deployment safety, rollback, and release-health signals.
  • Establish and run a lightweight, effective engineering on-call and incident-response process, with clear ownership and useful escalation paths.
  • Improve observability across services, asynchronous jobs, and complete customer workflows, and define reliability signals that reflect what users actually experience.
  • Build and improve CI/CD, automated deployment checks, feature rollout controls, environment reliability, and developer tooling.
  • Use AI to accelerate incident triage, log and trace analysis, root-cause investigation, runbook execution, and post-release analysis where it is genuinely useful.
  • Lead incident reviews focused on systemic fixes, automate recurring remediation, and partner with Quality and Security as we move toward GxP with reproducible change control, release evidence, auditability, and production monitoring.

Skills

Strong software engineering
Distributed systems
On-call experience
Safety-conscious coding
Toil reduction via software

Tools

AWS
Containers
CI/CD
Infrastructure automation
Observability tooling

Job description

Peer AI is seeking a Senior / Staff Systems & Reliability Engineer to own the release path from merge to production and to establish a lightweight on-call and incident-response framework. The role focuses on improving observability, CI/CD, and AI-assisted incident analysis within a hybrid Bay Area environment.

The team builds an AI-native platform for pharma workflows, with on-sites in San Francisco and a remote-friendly culture. Equity and parental leave are offered.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Platform Engineer - AI Control Plane
SRE Platform Engineer - AI Control Plane

Speakeasy Events, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Platform Engineer (SRE) - AI Control Plane
Senior Platform Engineer (SRE) - AI Control Plane

Speakeasy • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior SRE — AI-Driven Reliability & Oncall Leadership
Senior SRE — AI-Driven Reliability & Oncall Leadership

Block • San Francisco (CA)

On-site
USD 160,700 - 283,600
Healthcare coverage
Health Savings Account
Retirement Plans
+5
Systems & Reliability Engineer (SRE / Platform / DevOps) Full-Time, Bay Area, Hybrid See Job Details
Systems & Reliability Engineer (SRE / Platform / DevOps) Full-Time, Bay Area, Hybrid See Job Details

Getpeer • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 240,000
Meaningful equity grant
Remote-first with quarterly San Francs
Health coverage
+3
Senior SRE: Reliability, Automation & AI Platforms
Senior SRE: Reliability, Automation & AI Platforms

RX Brasil • Philadelphia

On-site
USD 95,000 - 159,000
Annual incentive bonus
Senior SRE: AI-Driven Reliability & Incident Leadership
Senior SRE: AI-Driven Reliability & Incident Leadership

Salesforce • San Francisco (CA)

On-site
USD 149,000 - 246,000
Senior SRE: Platform Reliability & AI-Driven Ops
Senior SRE: Platform Reliability & AI-Driven Ops

Block • New York (NY)

On-site
USD 170,100 - 283,600
Healthcare coverage
Retirement plans
Employee Stock Purchase Program
+1
AI Platform DevOps & SRE Lead
AI Platform DevOps & SRE Lead

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Competitive salary and early equity
Visa sponsorship
Generous health, dental, and vision coverage
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform

Instrumental Inc. • Palo Alto (CA)

On-site
USD 175,000 - 229,000
Health benefits
Commuter plans
Parental leave
Senior SRE: AI-Driven Kubernetes Reliability at Scale
Senior SRE: AI-Driven Kubernetes Reliability at Scale

fal - Features & Labels • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+1