Senior SRE: AI/ML Ops & Datastore Reliability

Adobe

San Jose (CA)

On-site

USD 208,000 - 302,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Adobe seeks a Senior SRE for RTCDP Datastores & AI/ML Ops in California. You will own production reliability, manage on-call incidents, and drive post-incident reviews for large-scale datastores and services across Aerospike, FoundationDB, Postgres, and CosmosDB/DynamoDB.

You’ll partner with product and platform teams to launch reliable deployments, optimize SLOs, and advance MLOps practices while improving monitoring and automation across regions.

Qualifications

  • 6–10 years in SRE, infrastructure, or platform engineering.
  • Proven track record operating large-scale distributed systems in production.
  • Strong foundation in datastores, reliability engineering, and automation.
  • Hands-on experience with Kubernetes and containerized environments, a major cloud (AWS, Azure, or GCP), and modern observability tooling (Prometheus, Grafana, OpenTelemetry, or equivalents).
  • Real experience in incident response and driving operational improvements out of it.
  • Working knowledge of — or genuine interest in — AI/ML systems or MLOps (expertise not required).
  • Comfortable with scale, ambiguity, and high ownership.

Responsibilities

  • Own production reliability and day-to-day reliability for RTCDP services against SLOs.
  • Participate in on-call rotations and incident response, driving mitigation and recovery.
  • Lead post-incident reviews and follow-up work.
  • Strengthen operational readiness, playbooks, and on-call health.
  • Operate and evolve core datastores (Aerospike, FoundationDB, Postgres, CosmosDB/DynamoDB).
  • Drive reliability, scaling, and observability across platforms.

Skills

SRE experience
Distributed systems
Datastores
Kubernetes
Cloud (AWS/Azure/GCP)
Observability tools
Incident response
Interest in AI/ML systems

Tools

Prometheus
Grafana
OpenTelemetry

Job description

Adobe seeks a Senior SRE for RTCDP Datastores & AI/ML Ops in California. You will own production reliability, manage on-call incidents, and drive post-incident reviews for large-scale datastores and services across Aerospike, FoundationDB, Postgres, and CosmosDB/DynamoDB.

You’ll partner with product and platform teams to launch reliable deployments, optimize SLOs, and advance MLOps practices while improving monitoring and automation across regions.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE - AI/ML Ops & Datastore Reliability
Senior SRE - AI/ML Ops & Datastore Reliability

Adobe • San Jose (CA)

On-site
USD 180,000 - 240,000
Job Posting Title AI/ DevOps Engineer
Job Posting Title AI/ DevOps Engineer

Adobe • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior SRE - Cloud-Native ML Infra & AI Pipelines
Senior SRE - Cloud-Native ML Infra & AI Pipelines

Adobe • New York (NY)

On-site
USD 178,000 - 258,000
Senior AI-Driven Real-Time Data Platform Engineer
Senior AI-Driven Real-Time Data Platform Engineer

Adobe • San Jose (CA)

Hybrid
USD 178,000 - 258,000
Competitive compensation
Equity awards
Health insurance
+3
Senior SRE Ops Lead — AI-Driven Reliability & AWS
Senior SRE Ops Lead — AI-Driven Reliability & AWS

Envision Technology Solutions • Charlotte (NC)

On-site
USD 140,000 - 190,000
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior Site Reliability Engineer – Multi-Cloud SaaS
Senior Site Reliability Engineer – Multi-Cloud SaaS

Adobe Inc. • Lehi (UT)

On-site
USD 139,000 - 258,000
Senior SRE: Scale Resilient AI Platforms & Automation
Senior SRE: Scale Resilient AI Platforms & Automation

Relx Plc • Philadelphia

Hybrid
USD 95,000 - 159,000
Senior SRE: AI-Driven Ops & Incident Leader
Senior SRE: AI-Driven Ops & Incident Leader

Salesforce.com, inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 149,000 - 224,000
Senior SRE: AI Infra, Hybrid Cloud & Automation
Senior SRE: AI Infra, Hybrid Cloud & Automation

d-Matrix • Santa Clara (CA)

On-site
USD 140,000 - 210,000