Job Posting Title AI/ DevOps Engineer

Adobe

San Jose (CA)

Vor Ort

USD 180.000 - 240.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Adobe's Real-Time Customer Data Platform (RTCDP) team seeks a Senior SRE to ensure reliability, scalability, and operational excellence across global deployments.

In this hands-on role, you will own day-to-day reliability, operate core datastores, and drive automation, observability, and AI/ML Ops initiatives for RTCDP.

Collaborate with product, platform, engineering, and security teams to deliver resilient services, cost efficiency, and architectural influence at scale.

Qualifikationen

  • 6–10 years in SRE, infrastructure, or platform engineering.
  • Proven ability to operate large-scale distributed systems in production.
  • Strong foundation in datastores, reliability engineering, and automation.

Aufgaben

  • Own day-to-day reliability for RTCDP services across global deployments.
  • Drive upgrades, capacity management, backup/restore, and DR testing for core datastores.
  • Build automation to provision, scale, and manage lifecycles of services.
  • Improve monitoring, alerting, and observability aligned to customer impact.
  • Mentor engineers and collaborate with product, platform, and security teams.

Kenntnisse

SRE mindset
Incident response
Automation
Leadership

Tools

Kubernetes
Prometheus
Grafana
OpenTelemetry
CI/CD tooling

Jobbeschreibung

Senior SRE — RTCDP Datastores & AI/ML Ops

Adobe’s Real‑Time Customer Data Platform (RTCDP) powers personalized experiences for some of the world’s largest brands. As a Senior SRE on this team, you’ll be central to keeping RTCDP reliable, scalable, and operationally excellent at global scale. This is a hands‑on, high‑ownership role at the intersection of production operations (Day’2 ownership) and core datastore engineering, with a growing surface area in operationalizing AI/ML services and workflows.

What You’ll Do
Own production reliability
  • Own day‑to‑day reliability for RTCDP services — availability, performance, and durability against SLOs
  • Participate in on‑call rotations and incident response, driving mitigation and recovery through SEV3–SEV1 events
  • Lead post‑incident reviews and follow‑up work
  • Strengthen operational readiness, playbooks, and on‑call health
  • Partner with product and platform teams on production‑ready launches and regional expansions
Operate and evolve core datastores
  • Drive reliability, scaling, and operational excellence across Aerospike, FoundationDB, Postgres, and CosmosDB/DynamoDB
  • Own upgrades, capacity management, backup/restore, and DR testing
  • Build automation for provisioning, scaling, and lifecycle management
  • Identify and ship cost optimizations (rightsizing, storage/compute efficiency)
Drive automation and observability
  • Build automation‑first solutions that reduce toil and improve system safety
  • Improve monitoring, alerting, and observability — anchored to real customer impact
  • Establish standardized operational patterns across services and regions
  • Support the rollout of SLO‑driven reliability practices
Contribute to AI/ML Ops (emerging area)
  • Support infrastructure and operational needs for AI/ML‑powered services in RTCDP
  • Shape operational practices for model serving and data pipelines — reliability, scaling, monitoring
  • Help land core MLOps patterns where relevant: model deployment workflows, inference observability (latency, errors), data quality and pipeline reliability signals
  • Partner with ML and data teams to get AI‑driven features production‑ready
  • Leverage AI‑assisted tools (e.g., Copilot, Claude Code, Codex, internal tooling) to accelerate debugging, incident response, and operational workflows
Technical leadership and collaboration
  • Operate as a strong IC and technical lead on cross‑cutting projects
  • Mentor junior engineers and raise team practices
  • Partner closely with engineering, infrastructure, and security
  • Live the core SRE/DevOps principles: ownership, automation, error budgets, continuous improvement
Why this role is interesting
  • Deep involvement in production systems at global scale
  • Hands‑on ownership spanning operations and datastore platforms
  • Exposure to next‑generation work: AI/ML systems and AI‑assisted engineering
  • Direct impact on customer reliability, platform scalability, and cost
  • A clear path toward architect‑level influence over time
What We’re Looking For
  • 6–10 years in SRE, infrastructure, or platform engineering
  • Proven track record operating large‑scale distributed systems in production
  • Strong foundation in datastores, reliability engineering, and automation
  • Hands‑on experience with Kubernetes and containerized environments, a major cloud (AWS, Azure, or GCP), and modern observability tooling (Prometheus, Grafana, OpenTelemetry, or equivalents)
  • Real experience in incident response and driving operational improvements out of it
  • Working knowledge of — or genuine interest in — AI/ML systems or MLOps (expertise not required)
  • Comfortable with scale, ambiguity, and high ownership
  • Strong problem‑solving instincts and a bias for action

Adobe is proud to be an Equal Employment Opportunity employer. We do not discriminate based on gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other protected characteristic. Learn more.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior SRE: AI/ML Ops & Datastore Reliability
Senior SRE: AI/ML Ops & Datastore Reliability

Adobe • San Jose (CA)

Vor Ort
USD 208.000 - 302.000
Senior SRE - AI/ML Ops & Datastore Reliability
Senior SRE - AI/ML Ops & Datastore Reliability

Adobe • San Jose (CA)

Vor Ort
USD 180.000 - 240.000
Senior SRE
Senior SRE

Accelerant • USA

Remote
USD 140.000 - 210.000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Wand AI • Palo Alto (CA)

Vor Ort
USD 180.000 - 250.000
Software Engineer - Adobe Experience Platform
Software Engineer - Adobe Experience Platform

Adobe Inc. • San Jose (CA)

Vor Ort
USD 140.000 - 230.000
Equity awards
Health coverage
Learning budgets
+3
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Stratitech Services LLC • San Francisco (CA)

Vor Ort
USD 180.000 - 260.000
Site Reliability Engineer
Site Reliability Engineer

Amtex Enterprises Inc. • Irvine (CA)

Vor Ort
USD 120.000 - 180.000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

Vor Ort
USD 120.000 - 150.000
Senior AI/ DevOps Engineer
Senior AI/ DevOps Engineer

Adobe • San Jose (CA)

Vor Ort
USD 208.000 - 302.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

Vor Ort
USD 120.000 - 160.000