Senior SRE - Cloud-Native ML Infra & AI Pipelines

Adobe

New York (NY)

On-site

USD 178,000 - 258,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Adobe is seeking a Senior Site Reliability Engineer to design, build, and operate large-scale, distributed, fault-tolerant systems across cloud platforms. You will own patch workflows, security guardrails, and ML inference infrastructure, partnering with software and ML teams to bake reliability in from the start.

You’ll work in a remote-capable environment, with a focus on scalable, secure operations and incident response.

Qualifications

  • BSc in Computer Science or equivalent experience in a practical setting.
  • Proficiency in Python; familiarity with PHP, Node.js, or Ruby is a plus.
  • Experience deploying/operating ML inference pipelines in production (SageMaker, OpenAI, Bedrock, etc.).
  • Experience operating cloud-native compute over large scale infrastructures (AWS EC2 Auto Scaling Groups; Kubernetes).
  • Experience with vulnerability/patch management and AMI/golden-image lifecycle automation.

Responsibilities

  • Design, build, and operate large-scale, distributed, fault-tolerant systems and the automation tooling behind them.
  • Own patch, vulnerability, and golden-image lifecycle management across the fleet; triage and automate remediation.
  • Contribute to multi-quarter initiatives like compute rationalization and cloud re-platforming.
  • Build and operate ML inference infrastructure including model serving and GPU workloads.
  • Set infrastructure standards and security guardrails for agentic AI and tooling.
  • Partner with software, ML, and platform teams to bake reliability in from the start.
  • Share on-call responsibilities and resolve issues in web services, databases, and pipelines.

Skills

Python
SRE
Distributed systems
On-call rotation
CI/CD
Security guardrails
Linux

Education

BSc in Computer Science or equivalent

Tools

AWS
Azure
GCP
Kubernetes
Terraform
Terragrunt
Atlantis
Chef
Ansible
SSM
Docker

Job description

Adobe is seeking a Senior Site Reliability Engineer to design, build, and operate large-scale, distributed, fault-tolerant systems across cloud platforms. You will own patch workflows, security guardrails, and ML inference infrastructure, partnering with software and ML teams to bake reliability in from the start.

You’ll work in a remote-capable environment, with a focus on scalable, secure operations and incident response.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer, Cloud & ML Infra
Senior DevOps Engineer, Cloud & ML Infra

Adobe • San Jose (CA)

On-site
USD 178,000 - 258,000
Senior SRE: AI-Driven, Cloud-Native Reliability (Hybrid)
Senior SRE: AI-Driven, Cloud-Native Reliability (Hybrid)

OutSystems • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Hybrid work model
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior SRE: ML Infra at Scale, Multi-Cloud K8s
Senior SRE: ML Infra at Scale, Multi-Cloud K8s

The Consensus • New York (NY)

On-site
USD 110,000 - 140,000
Competitive compensation
100% insurance coverage
Flexible PTO policy
+3
Senior DevOps Engineer for Scalable AI Platforms
Senior DevOps Engineer for Scalable AI Platforms

Adobe Inc. • San Jose (CA)

On-site
USD 178,000 - 258,000
Senior SRE - AI/ML Ops & Datastore Reliability
Senior SRE - AI/ML Ops & Datastore Reliability

Adobe • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior Cloud SRE – ML Vision Infra & Kubernetes
Senior Cloud SRE – ML Vision Infra & Kubernetes

Apple Inc. • San Diego (CA)

On-site
USD 142,300 - 263,300
Stock plans
Relocation assistance
Education reimbursement
Senior SRE: AI-Driven Reliability & Automation (Hybrid)
Senior SRE: AI-Driven Reliability & Automation (Hybrid)

Namely • United States

Hybrid
USD 120,000 - 150,000
Senior SRE: AI-Driven Platform Reliability & Scale
Senior SRE: AI-Driven Platform Reliability & Scale

Medallia • McLean (VA)

On-site
USD 129,000 - 190,000
Health benefits
401(k) matching
Paid parental leave
+1