Job Posting Title AI/ DevOps Engineer

Adobe

San Jose (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Adobe's Real-Time Customer Data Platform (RTCDP) team seeks a Senior SRE to ensure reliability, scalability, and operational excellence across global deployments.

In this hands-on role, you will own day-to-day reliability, operate core datastores, and drive automation, observability, and AI/ML Ops initiatives for RTCDP.

Collaborate with product, platform, engineering, and security teams to deliver resilient services, cost efficiency, and architectural influence at scale.

Qualifications

  • 6–10 years in SRE, infrastructure, or platform engineering.
  • Proven ability to operate large-scale distributed systems in production.
  • Strong foundation in datastores, reliability engineering, and automation.

Responsibilities

  • Own day-to-day reliability for RTCDP services across global deployments.
  • Drive upgrades, capacity management, backup/restore, and DR testing for core datastores.
  • Build automation to provision, scale, and manage lifecycles of services.
  • Improve monitoring, alerting, and observability aligned to customer impact.
  • Mentor engineers and collaborate with product, platform, and security teams.

Skills

SRE mindset
Incident response
Automation
Leadership

Tools

Kubernetes
Prometheus
Grafana
OpenTelemetry
CI/CD tooling

Job description

Senior SRE — RTCDP Datastores & AI/ML Ops

Adobe’s Real‑Time Customer Data Platform (RTCDP) powers personalized experiences for some of the world’s largest brands. As a Senior SRE on this team, you’ll be central to keeping RTCDP reliable, scalable, and operationally excellent at global scale. This is a hands‑on, high‑ownership role at the intersection of production operations (Day’2 ownership) and core datastore engineering, with a growing surface area in operationalizing AI/ML services and workflows.

What You’ll Do
Own production reliability
  • Own day‑to‑day reliability for RTCDP services — availability, performance, and durability against SLOs
  • Participate in on‑call rotations and incident response, driving mitigation and recovery through SEV3–SEV1 events
  • Lead post‑incident reviews and follow‑up work
  • Strengthen operational readiness, playbooks, and on‑call health
  • Partner with product and platform teams on production‑ready launches and regional expansions
Operate and evolve core datastores
  • Drive reliability, scaling, and operational excellence across Aerospike, FoundationDB, Postgres, and CosmosDB/DynamoDB
  • Own upgrades, capacity management, backup/restore, and DR testing
  • Build automation for provisioning, scaling, and lifecycle management
  • Identify and ship cost optimizations (rightsizing, storage/compute efficiency)
Drive automation and observability
  • Build automation‑first solutions that reduce toil and improve system safety
  • Improve monitoring, alerting, and observability — anchored to real customer impact
  • Establish standardized operational patterns across services and regions
  • Support the rollout of SLO‑driven reliability practices
Contribute to AI/ML Ops (emerging area)
  • Support infrastructure and operational needs for AI/ML‑powered services in RTCDP
  • Shape operational practices for model serving and data pipelines — reliability, scaling, monitoring
  • Help land core MLOps patterns where relevant: model deployment workflows, inference observability (latency, errors), data quality and pipeline reliability signals
  • Partner with ML and data teams to get AI‑driven features production‑ready
  • Leverage AI‑assisted tools (e.g., Copilot, Claude Code, Codex, internal tooling) to accelerate debugging, incident response, and operational workflows
Technical leadership and collaboration
  • Operate as a strong IC and technical lead on cross‑cutting projects
  • Mentor junior engineers and raise team practices
  • Partner closely with engineering, infrastructure, and security
  • Live the core SRE/DevOps principles: ownership, automation, error budgets, continuous improvement
Why this role is interesting
  • Deep involvement in production systems at global scale
  • Hands‑on ownership spanning operations and datastore platforms
  • Exposure to next‑generation work: AI/ML systems and AI‑assisted engineering
  • Direct impact on customer reliability, platform scalability, and cost
  • A clear path toward architect‑level influence over time
What We’re Looking For
  • 6–10 years in SRE, infrastructure, or platform engineering
  • Proven track record operating large‑scale distributed systems in production
  • Strong foundation in datastores, reliability engineering, and automation
  • Hands‑on experience with Kubernetes and containerized environments, a major cloud (AWS, Azure, or GCP), and modern observability tooling (Prometheus, Grafana, OpenTelemetry, or equivalents)
  • Real experience in incident response and driving operational improvements out of it
  • Working knowledge of — or genuine interest in — AI/ML systems or MLOps (expertise not required)
  • Comfortable with scale, ambiguity, and high ownership
  • Strong problem‑solving instincts and a bias for action

Adobe is proud to be an Equal Employment Opportunity employer. We do not discriminate based on gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other protected characteristic. Learn more.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI/ML Ops & Datastore Reliability
Senior SRE: AI/ML Ops & Datastore Reliability

Adobe Inc. • San Jose (CA)

On-site
USD 228,000 - 332,000
Job Posting Title AI/ DevOps Engineer
Job Posting Title AI/ DevOps Engineer

Adobe Inc. • San Jose (CA)

On-site
USD 228,000 - 332,000
Head of SRE
Head of SRE

Wand AI • Palo Alto (CA)

On-site
USD 130,000 - 180,000
Software Engineer - Adobe Experience Platform
Software Engineer - Adobe Experience Platform

Adobe Inc. • San Jose (CA)

Hybrid
USD 140,000 - 230,000
Equity awards
Health coverage
Learning budgets
+3
Digital - Principal SRE
Digital - Principal SRE

Huntington National Bank • Columbus (OH)

On-site
USD 100,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

hardrockdigital • United States

Hybrid
USD 120,000 - 160,000
Competitive pay and benefits
Flexible vacation allowance
Startup culture with global brand support
+1
Site Reliability Engineer
Site Reliability Engineer

Optomi • New York (NY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Optomi • Seattle (WA)

On-site
USD 140,000 - 180,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

R&D • United States

On-site
USD 90,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000