Site Reliability Engineer III: Observability & Automation

Jobtailor

Irvine (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Taco Bell is seeking a Site Reliability Engineer to elevate observability and automate core platform workloads. You will own monitoring, SLOs, and incident response while partnering with product and engineering teams to improve reliability and performance.

You will operate in a fast-paced, on-call driven environment and apply modern cloud practices using DataDog, CloudWatch, OpenTelemetry, Lambda, API Gateway, and DynamoDB to keep systems healthy and scalable.

Qualifications

  • Bachelor's degree in CS/engineering or equivalent work experience.
  • 2+ years in the SRE space with a focus on observability and automation.
  • Familiarity with SRE core principles (SLO, SLA, SLI, Error Budget).
  • Hands-on experience creating monitors, dashboards, SLOs, and observability capabilities.
  • Experience with logging solutions like DataDog, CloudWatch log insights, etc.
  • Understanding of incident management practices, bridge calls, RCAs and postmortems.
  • Familiarity with distributed tracing, APM, OpenTelemetry, and RUM.
  • Excellent communication and collaboration skills in a fast-paced team.

Responsibilities

  • Create, update, or automate internal processes or tools to reduce toil and boost productivity.
  • Understand and monitor the Taco Bell Digital ecosystem for performance and data accuracy.
  • Communicate with technical and non-technical stakeholders about system health and changes.
  • Perform final validation tests on mobile and web applications and report on improvements.
  • Develop expertise in serverless infrastructure and modern SRE practices.
  • Participate in Agile practices and 24/7 on-call rotation.
  • Collaborate with cross-functional partners on high-impact issues affecting revenue and brand.

Skills

Observability
SRE
SQL
AWS
DataDog
CloudWatch
OpenTelemetry
JavaScript
Python
Go
TypeScript
SLA/SLO/SLI
Incident management
On-call
Agile
Troubleshooting
Communication

Education

Bachelor's degree in computer science, engineering, or related field

Tools

DataDog
CloudWatch
Lumigo
OpenTelemetry
Lambda
API Gateway
Fargate
S3
DynamoDB
EventBridge

Job description

Taco Bell is seeking a Site Reliability Engineer to elevate observability and automate core platform workloads. You will own monitoring, SLOs, and incident response while partnering with product and engineering teams to improve reliability and performance.

You will operate in a fast-paced, on-call driven environment and apply modern cloud practices using DataDog, CloudWatch, OpenTelemetry, Lambda, API Gateway, and DynamoDB to keep systems healthy and scalable.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer – III
Site Reliability Engineer – III

Jobtailor • Irvine (CA)

On-site
USD 120,000 - 180,000
Site Reliability Engineer — Observability & Automation
Site Reliability Engineer — Observability & Automation

Randstad Digital Americas • Plano (TX)

On-site
USD 115,000 - 125,000
Site Reliability Engineer: Cloud, Automation & Observability
Site Reliability Engineer: Cloud, Automation & Observability

TechDigital Group • Houston (TX), Juno Beach (FL)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

O.C. Tanner • Salt Lake City (UT)

On-site
USD 130,000 - 180,000
Site Reliability & Observability Engineer
Site Reliability & Observability Engineer

Stability Technology • United States

On-site
USD 120,000 - 170,000
Site Reliability Engineer III — Scale, Observability & Automation
Site Reliability Engineer III — Scale, Observability & Automation

onXmaps, Inc. • Bozeman (MT)

Hybrid
USD 130,000 - 153,000
Health benefits including no monthly–$
401(k) matching
Parental leave
+2
Site Reliability Engineer
Site Reliability Engineer

Stability Technology • United States

On-site
USD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Infosys • Richardson (TX)

On-site
USD 80,000 - 120,000
Senior Site Reliability Engineer – Remote, Impact & Automation
Senior Site Reliability Engineer – Remote, Impact & Automation

Midwest Startups • United States

On-site
USD 175,000 - 185,000
Market-leading medical, dental, and視on
Stock options
Premium-Tier Origin Financial Wellness
+6
Senior SRE: Cloud-Native Reliability & Observability
Senior SRE: Cloud-Native Reliability & Observability

O.C. Tanner • Salt Lake City (UT)

On-site
USD 130,000 - 180,000