SRE Engineer — Scale AI Platforms (Hybrid Work)

Plaud

San Francisco (CA)

Hybrid

USD 180,000 - 230,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

ESOP
Hybrid work model
Health benefits

Job summary

Plaud Inc. in San Francisco, CA, is seeking an experienced Site Reliability/Platform Engineer to ensure reliability and performance of Plaud.ai at scale.

You will design and operate cloud-native systems for AI workloads, own production reliability, incident response, and drive observability improvements. The role emphasizes on-call rotation, defining SLOs/SLIs, collaborating with product and engineering teams, and advancing multi-region, high-availability infrastructure in a fast-growing,

Qualifications

  • 5+ years in SRE, infra, or platform engineering.
  • Experience with cloud platforms (AWS/GCP/Azure/OCI).
  • Kubernetes & distributed systems experience.
  • On-call rotation and incident management experience.
  • Programming in Go, Python, or Java.

Responsibilities

  • Ensure reliability and performance of Plaud.ai’s AI products at scale.
  • Design and operate highly available, scalable cloud-native systems for AI workloads.
  • Own production reliability, incident response, and on-call practices.
  • Build observability (metrics, logs, tracing) and reliability automation.
  • Define and manage SLOs, SLIs, and error budgets with engineering teams.
  • Drive postmortems and reliability improvements across the platform.
  • Lead incident response and continuous reliability improvement.
  • Partner with product and engineering teams on reliability design.
  • Improve observability and operational maturity.

Skills

SRE / Platform engineering
Cloud platforms
Kubernetes
On-call / incident management
Go / Python / Java

Job description

Plaud Inc. in San Francisco, CA, is seeking an experienced Site Reliability/Platform Engineer to ensure reliability and performance of Plaud.ai at scale.

You will design and operate cloud-native systems for AI workloads, own production reliability, incident response, and drive observability improvements. The role emphasizes on-call rotation, defining SLOs/SLIs, collaborating with product and engineering teams, and advancing multi-region, high-availability infrastructure in a fast-growing,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE for Scalable AI Platforms & Reliability
Senior SRE for Scalable AI Platforms & Reliability

Plaud • Seattle (WA)

Hybrid
USD 180,000 - 240,000
ESOP
High-impact environment
Health & retirement benefits
+5
Senior SRE - AI Infra, Cloud & Observability
Senior SRE - AI Infra, Cloud & Observability

Plaud • Seattle (WA)

Hybrid
USD 150,000 - 210,000
ESOP
Health & Retirement Benefits
Unlimited PTO
+2
Senior SRE: Scale Reliability for Health AI Platform
Senior SRE: Scale Reliability for Health AI Platform

RXinsider LTD. • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Healthcare Coverage
401(k) Match
Equity
+5
Senior SRE: Scale Reliability Leader (Hybrid)
Senior SRE: Scale Reliability Leader (Hybrid)

Plenful • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Healthcare Coverage
401(k) with Company Match
Equity
+5
Hybrid SRE Engineer for Scalable Reliability & Equity
Hybrid SRE Engineer for Scalable Reliability & Equity

EarnIn • Mountain View (CA)

Hybrid
USD 139,000 - 232,000
SRE, Cloud Platform – AI-Driven Reliability & Observability
SRE, Cloud Platform – AI-Driven Reliability & Observability

Visa • Austin (TX)

On-site
USD 88,000 - 137,000
Senior SRE: AI-Driven Platform Reliability & Scale
Senior SRE: AI-Driven Platform Reliability & Scale

Medallia • McLean (VA)

On-site
USD 129,000 - 190,000
Health benefits
401(k) matching
Paid parental leave
+1
Remote SRE: AI Platform Reliability & Automation
Remote SRE: AI Platform Reliability & Automation

Runpod • United States

On-site
USD 150,000 - 200,000
Remote work first
Competitive base salary
Stock options equity
+2
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity

FLUIX • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Attractive compensation package including equity options
Comprehensive health, dental, and vision insurance
Opportunities for professional growth
Staff SRE & Platform Reliability Architect
Staff SRE & Platform Reliability Architect

Grailbio • Edison (CA)

On-site
USD 169,000 - 224,000
Flexible time-off
401(k) with employer match
Medical, dental, vision coverage