SRE Engineer — Scale AI Platforms (Hybrid Work)

Plaud

San Francisco (CA)

Hybrid

USD 180,000 - 230,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

ESOP
Hybrid work model
Health benefits

Job summary

Plaud Inc. in San Francisco, CA, is seeking an experienced Site Reliability/Platform Engineer to ensure reliability and performance of Plaud.ai at scale.

You will design and operate cloud-native systems for AI workloads, own production reliability, incident response, and drive observability improvements. The role emphasizes on-call rotation, defining SLOs/SLIs, collaborating with product and engineering teams, and advancing multi-region, high-availability infrastructure in a fast-growing,

Qualifications

  • 5+ years in SRE, infra, or platform engineering.
  • Experience with cloud platforms (AWS/GCP/Azure/OCI).
  • Kubernetes & distributed systems experience.
  • On-call rotation and incident management experience.
  • Programming in Go, Python, or Java.

Responsibilities

  • Ensure reliability and performance of Plaud.ai’s AI products at scale.
  • Design and operate highly available, scalable cloud-native systems for AI workloads.
  • Own production reliability, incident response, and on-call practices.
  • Build observability (metrics, logs, tracing) and reliability automation.
  • Define and manage SLOs, SLIs, and error budgets with engineering teams.
  • Drive postmortems and reliability improvements across the platform.
  • Lead incident response and continuous reliability improvement.
  • Partner with product and engineering teams on reliability design.
  • Improve observability and operational maturity.

Skills

SRE / Platform engineering
Cloud platforms
Kubernetes
On-call / incident management
Go / Python / Java

Job description

Plaud Inc. in San Francisco, CA, is seeking an experienced Site Reliability/Platform Engineer to ensure reliability and performance of Plaud.ai at scale.

You will design and operate cloud-native systems for AI workloads, own production reliability, incident response, and drive observability improvements. The role emphasizes on-call rotation, defining SLOs/SLIs, collaborating with product and engineering teams, and advancing multi-region, high-availability infrastructure in a fast-growing,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: Scale Reliability Leader (Hybrid)
Senior SRE: Scale Reliability Leader (Hybrid)

Plenful • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Healthcare Coverage
401(k) with Company Match
Equity
+5
Senior SRE: AI-Driven, Cloud-Native Reliability (Hybrid)
Senior SRE: AI-Driven, Cloud-Native Reliability (Hybrid)

OutSystems • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Hybrid work model
Senior SRE: Scale Resilient AI Platforms & Automation
Senior SRE: Scale Resilient AI Platforms & Automation

Relx Plc • Philadelphia

Hybrid
USD 95,000 - 159,000
Sr. SRE Leader — Platform Reliability & AI-Driven Infra
Sr. SRE Leader — Platform Reliability & AI-Driven Infra

Invoca • Denver (CO)

Remote
USD 190,000 - 250,000
Senior SRE: AI-Driven Platform Reliability & Scale
Senior SRE: AI-Driven Platform Reliability & Scale

Medallia • McLean (VA)

On-site
USD 129,000 - 190,000
Health benefits
401(k) matching
Paid parental leave
+1
Remote SRE: AI Platform Reliability & Automation
Remote SRE: AI Platform Reliability & Automation

Runpod • United States

On-site
USD 150,000 - 200,000
Remote work first
Competitive base salary
Stock options equity
+2
Senior SRE Leader: AI-Driven Infra & Platform Reliability
Senior SRE Leader: AI-Driven Infra & Platform Reliability

Invoca • Los Angeles (CA)

Remote
USD 190,000 - 250,000
Health benefits
Mental wellbeing
Wellness subsidy
+6
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity

FLUIX • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Attractive compensation package including equity options
Comprehensive health, dental, and vision insurance
Opportunities for professional growth
Senior SRE for AI-First Platform & Scale
Senior SRE for AI-First Platform & Scale

Replicant • United States

On-site
USD 140,000 - 210,000
Offsites
Tech & learning stipend
Remote by design
+3
AI Platform SRE: Reliability, Observability & Scale
AI Platform SRE: Reliability, Observability & Scale

Schonfeld • New York (NY)

On-site
USD 175,000 - 225,000