Senior SRE - AI Infra, Cloud & Observability

Plaud

Seattle (WA)

Hybrid

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

ESOP
Health & Retirement Benefits
Unlimited PTO
Hybrid work model
Office snacks

Job summary

Plaud is seeking a Senior SRE to ensure reliability of Plaud.ai at scale, designing and operating cloud-native systems for AI workloads. You will own production reliability, incident response, and on-call practices while building observability and automation with the engineering teams.

You should bring 5+ years in SRE/Infra, strong cloud experience, and Kubernetes know-how. This role offers a hybrid work model with generous benefits and career growth opportunities.

Qualifications

  • 5+ years in SRE, Infra, or Platform Engineering.
  • Experience with cloud platforms (AWS/GCP/Azure/OCI).
  • Kubernetes and distributed systems experience.
  • On-call rotation experience and incident response.
  • Proficient in at least one programming language (Go, Python, Java).

Responsibilities

  • Ensure reliability and performance of Plaud.ai’s AI products at scale.
  • Design and operate highly available, scalable cloud-native systems for AI workloads.
  • Own production reliability, incident response, and on-call practices.
  • Build observability (metrics, logs, tracing) and reliability automation.
  • Define and manage SLOs/SLIs and error budgets with engineering teams.
  • Drive postmortems and reliability improvements across the platform.
  • Lead incident response and continuous reliability improvement.
  • Partner with product and engineering teams on reliability design.
  • Improve observability and operational maturity.

Skills

SRE / Platform Eng
Cloud platforms experience
On-call & incident management
Programming: Go/Python/Java

Tools

Kubernetes

Job description

Plaud is seeking a Senior SRE to ensure reliability of Plaud.ai at scale, designing and operating cloud-native systems for AI workloads. You will own production reliability, incident response, and on-call practices while building observability and automation with the engineering teams.

You should bring 5+ years in SRE/Infra, strong cloud experience, and Kubernetes know-how. This role offers a hybrid work model with generous benefits and career growth opportunities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE for Scalable AI Platforms & Reliability
Senior SRE for Scalable AI Platforms & Reliability

Plaud • Seattle (WA)

Hybrid
USD 180,000 - 240,000
ESOP
High-impact environment
Health & retirement benefits
+5
SRE Engineer — Scale AI Platforms (Hybrid Work)
SRE Engineer — Scale AI Platforms (Hybrid Work)

Plaud • San Francisco (CA)

Hybrid
USD 180,000 - 230,000
ESOP
Hybrid work model
Health benefits
Senior SRE Engineer: Scale AI Platforms & Reliability
Senior SRE Engineer: Scale AI Platforms & Reliability

Plaud • Seattle (WA)

On-site
USD 140,000 - 190,000
SRE Engineer - Seattle
SRE Engineer - Seattle

Plaud • Seattle (WA)

Hybrid
USD 140,000 - 190,000
ESOP ownership
Hybrid work model
Health & retirement benefits
+5
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior SRE: AI Cloud Platform & Kubernetes Expert
Senior SRE: AI Cloud Platform & Kubernetes Expert

Lambda • Bellevue (WA)

On-site
USD 180,000 - 260,000
Health insurance
Dental insurance
Vision insurance
+3
Senior SRE: AI Cloud Platform & Kubernetes
Senior SRE: AI Cloud Platform & Kubernetes

Lambda Inc. • San Francisco (CA)

Hybrid
USD 190,000 - 270,000
Health insurance
401k with company match
Flexible PTO
+2
SRE Engineer - San Francisco
SRE Engineer - San Francisco

Plaud • San Francisco (CA)

Hybrid
USD 180,000 - 230,000
ESOP
Hybrid work model
Health benefits
Senior SRE: Scale Reliability for Health AI Platform
Senior SRE: Scale Reliability for Health AI Platform

RXinsider LTD. • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Healthcare Coverage
401(k) Match
Equity
+5
Senior Cloud SRE: AI-Driven Multi-Cloud Infra
Senior Cloud SRE: AI-Driven Multi-Cloud Infra

Palo Alto Networks • California (MO)

On-site
USD 160,000 - 210,000