Site Reliability Engineer – III

Jobtailor

Irvine (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Taco Bell is seeking a Site Reliability Engineer to elevate observability and automate core platform workloads. You will own monitoring, SLOs, and incident response while partnering with product and engineering teams to improve reliability and performance.

You will operate in a fast-paced, on-call driven environment and apply modern cloud practices using DataDog, CloudWatch, OpenTelemetry, Lambda, API Gateway, and DynamoDB to keep systems healthy and scalable.

Qualifications

  • Bachelor's degree in CS/engineering or equivalent work experience.
  • 2+ years in the SRE space with a focus on observability and automation.
  • Familiarity with SRE core principles (SLO, SLA, SLI, Error Budget).
  • Hands-on experience creating monitors, dashboards, SLOs, and observability capabilities.
  • Experience with logging solutions like DataDog, CloudWatch log insights, etc.
  • Understanding of incident management practices, bridge calls, RCAs and postmortems.
  • Familiarity with distributed tracing, APM, OpenTelemetry, and RUM.
  • Excellent communication and collaboration skills in a fast-paced team.

Responsibilities

  • Create, update, or automate internal processes or tools to reduce toil and boost productivity.
  • Understand and monitor the Taco Bell Digital ecosystem for performance and data accuracy.
  • Communicate with technical and non-technical stakeholders about system health and changes.
  • Perform final validation tests on mobile and web applications and report on improvements.
  • Develop expertise in serverless infrastructure and modern SRE practices.
  • Participate in Agile practices and 24/7 on-call rotation.
  • Collaborate with cross-functional partners on high-impact issues affecting revenue and brand.

Skills

Observability
SRE
SQL
AWS
DataDog
CloudWatch
OpenTelemetry
JavaScript
Python
Go
TypeScript
SLA/SLO/SLI
Incident management
On-call
Agile
Troubleshooting
Communication

Education

Bachelor's degree in computer science, engineering, or related field

Tools

DataDog
CloudWatch
Lumigo
OpenTelemetry
Lambda
API Gateway
Fargate
S3
DynamoDB
EventBridge

Job description

Responsibilities
  • Create, update, or automate internal business processes or tools to reduce toil and improve team productivity.
  • Understand and monitor the Taco Bell Digital ecosystem for performance, availability, and accuracy of transactional data.
  • Communicate and collaborate with both technical and non‑technical stakeholders on issues, upcoming changes, and updates to system health.
  • Perform final validation tests on various mobile and web‑based applications, reporting on, and offering feedback on areas for improvement.
  • Build expertise in serverless infrastructure and initiatives while also learning aspects of modern SRE practices and terms, such as SLIs, SLOs, Observability, toil, and incident response with blameless post‑mortems.
  • Work with and adopt Agile practices while participating in a 24/7 on‑call rotation.
  • Collaborate within the team and with cross‑functional partners on high‑impact business issues that affect revenue and brand reputation.
Requirements
  • Bachelor’s degree in computer science, engineering, OR a related field, OR equivalent work experience.
  • At least 2+ years of experience in the SRE space, with a focus on observability and automation.
  • Familiarity with SRE core principles (e.g. SLO, SLA, SLI, Error Budget, etc.).
  • Hands‑on experience creating monitors, dashboards, SLOs, and other observability capabilities.
  • Experience with logging solutions or platforms such as DataDog, CloudWatch log insights, etc.
  • Understanding of incident management practices, including leading bridge calls, conducting RCAs, and facilitating postmortems.
  • Familiarity with modern observability practices and tools, such as distributed tracing, APM, OpenTelemetry, and RUM.
  • Excellent communication and collaboration skills, with the ability to work effectively in a fast‑paced environment as a member of a team.
  • A fundamentally complete understanding of Observability principles (not just monitoring) + experience using tools like DataDog, Lumigo, CloudWatch, or similar.
  • General level understanding of Agile methods such as Kanban, Scrum, etc.
  • Advanced troubleshooting skills.
  • A curious mindset and the desire to always keep learning.
  • Proactive self‑starter capable of operating autonomously.
  • Ability to participate in an on‑call rotation.
  • Proficiency in SQL.
  • Fundamental understanding of JavaScript, Python, Go, or TypeScript.
  • Working knowledge of AWS services commonly used in serverless and cloud‑native environments, including Lambda, API Gateway, Fargate, S3, DynamoDB, and EventBridge.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer III: Observability & Automation
Site Reliability Engineer III: Observability & Automation

Jobtailor • Irvine (CA)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Hampshire

On-site
USD 110,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Veriipro • Washington

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Madison-Davis, LLC • United States

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TechDigital Group • Englewood Cliffs (NJ)

On-site
USD 110,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000