Senior Site Reliability Engineer (Onsite Role)

O'Reilly Technology Services, Inc.

United States

On-site

USD 120,000 - 190,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive wages
401k with employer contributions
Medical, dental, and vision insurance
Wellbeing programs
Tuition assistance
Career growth opportunities

Job summary

O’Reilly Auto Parts seeks a Site Reliability Engineer to ensure the availability and performance of our customer‑facing microservices powering eCommerce channels. This on‑site role in Springfield, MO applies SRE principles to balance feature velocity with system reliability, using SLOs, SLIs, and error budgets.

You will drive observability, incident response, and automation, collaborating with global teams.

Qualifications

  • 4+ years in SRE, software engineering, or prod ops for large-scale platforms.
  • Hands-on with distributed systems and cloud deployments.
  • Experience defining and measuring reliability metrics (SLOs, SLIs).

Responsibilities

  • Define, own, and improve SLIs, SLOs, and error budgets for key microservices.
  • Lead incident response and blameless postmortems to drive improvements.
  • Design and maintain observability platforms with metrics, logs, traces, and telemetry.
  • Automate toil away from operations and push work into Jira with clear ownership.
  • Collaborate with product and engineering teams to balance reliability and velocity.
  • Promote shift-left reliability and prepare production readiness reviews.

Skills

SRE
Java
Cloud (AWS/Azure/GCP)
Observability
Incident response
Automation
Jira
Team leadership

Tools

Jira
Kubernetes
Docker
CI/CD
Terraform

Job description

Site Reliability Engineers are responsible for ensuring the availability, reliability, scalability, and performance of the firm’s most critical customer-facing microservices that power all eCommerce channels. This role applies Google-inspired SRE principles to balance feature velocity and system reliability using Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets. The role combines software engineering, cloud engineering, automation, and production operations, with a strong emphasis on building systems that are observable, resilient, and operable by default. This is an on-site position located in Springfield, MO. Remote work is not an option for this position.

Primary Responsibilities
  • Define, implement, and own SLIs, SLOs, and error budgets for critical microservices in collaboration with product and engineering teams.
  • Use error budgets to influence release decisions, prioritize reliability initiatives, and manage operational risk.
  • Design and maintain observability platforms, including metrics, logs, traces, and real-time telemetry.
  • Track, manage, and reduce operational toil by converting repetitive operational tasks into Jira stories and epics with clear ownership and measurable outcomes.
  • Design, implement, and validate resiliency mechanisms such as graceful degradation, redundancy, automated failover, and disaster recovery.
  • Lead incident response efforts, act as an escalation point for high-severity incidents, and drive blameless postmortems.
  • Capture incident action items and reliability improvements in Jira, ensuring accountability, closure, and continuous improvement.
  • Partner with Scrum teams to improve reliability through release readiness reviews, production change validation, and testing strategies.
  • Perform deep root cause analysis, debugging, and performance tuning across distributed systems.
  • Promote shift-left reliability practices by embedding operability, monitoring, and failure testing early in the SDLC.
  • Drive continuous improvement through automation, self-healing systems, chaos engineering, and capacity planning.
  • Maintain runbooks, playbooks, and knowledge repositories, linking documentation to Jira tasks to reduce MTTR.
  • Provide technical leadership and mentoring to junior SREs and engineers.
  • Collaborate with global, distributed teams, leveraging Jira for transparent planning, dependency tracking, and execution.
  • Conduct production readiness reviews and ensure services meet operational excellence standards before deployment.
  • Track and improve operational KPIs such as availability, MTTR, MTTD, deployment success rate, and incident recurrence.
  • Collaborate with security and platform teams to ensure reliability, compliance, and operational security best practices are embedded into systems and deployment pipelines.
  • Explore opportunities to leverage AI-driven observability, anomaly detection, and operational automation to improve system reliability and reduce manual effort.
Core Competencies & Qualifications
  • 4+ years of experience in SRE, software engineering, or production operations supporting large-scale eCommerce platforms.
  • Hands‑on experience with Java/J2EE‑based distributed systems; React experience is a plus.
  • Proven ability to design and operate systems using SLO‑driven reliability models.
  • Experience defining and measuring SLIs, including availability, latency, error rates, throughput, and saturation.
  • Good understanding of NoSQL technologies and RDBMS concepts, with the ability to write and troubleshoot database queries.
  • Experience deploying and operating services on cloud platforms such as AWS, Azure, or Google Cloud Platform (GCP).
  • Expertise with observability, APM, and caching tools such as Dynatrace, Splunk, ELK, Akamai, Quantum Metric, and Tealeaf.
  • Strong experience with Jira for backlog management, incident tracking, toil reduction initiatives, and cross‑team coordination.
  • Ability to independently own services and drive reliability initiatives end‑to‑end.
  • Strong communication skills with the ability to influence engineering and product teams.
  • Experience participating in on‑call rotations and handling critical/high‑severity incidents.
Desired Skills
  • Experience building and operating microservices architectures using Spring Boot, Groovy, React, or similar technologies.
  • Strong understanding of CI/CD pipelines, release automation, and progressive delivery practices.
  • Experience working within eCommerce domains such as Catalog, Customer Data, and Order Management.
  • Familiarity with search platforms including Endeca, Solr, Lucene, and Elasticsearch.
  • Proficiency in scripting and automation using Python, Bash, Ruby, Perl, or PowerShell.
  • Experience with ITSM tools integrated with Jira workflows.
  • Exposure to capacity planning, load testing, and chaos engineering practices.
  • Experience with containerization and orchestration technologies such as Docker and Kubernetes (EKS, AKS, or GKE).
  • Familiarity with Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, or Ansible.
  • Understanding of operational KPIs including availability, MTTR, MTTD, deployment success rate, and incident recurrence metrics.
  • Experience conducting production readiness reviews and implementing operational governance processes.
  • Ability to collaborate with security and platform engineering teams to ensure reliability, compliance, and operational security best practices.
  • Exposure to AI‑assisted operations, anomaly detection, intelligent alerting, and automated remediation solutions.
  • Experience designing scalable, self‑healing platforms and automation frameworks for cloud‑native environments.
Total Compensation Package
  • Competitive Wages & Paid Time Off
  • Stock Purchase Plan & 401k with Employer Contributions
  • Starting Day One
  • Medical, Dental, & Vision Insurance with Optional Flexible Spending Account (FSA)
  • Team Member Health/Wellbeing Programs
  • Tuition Educational Assistance Programs
  • Opportunities for Career Growth

O’Reilly Auto Parts is an equal opportunity employer.

The Company does not discriminate on the basis of race, religion, color, national origin or ancestry (including immigration status or citizenship), sex, sexual orientation, gender identity, pregnancy (including childbirth, lactation, and related medical conditions,) age (40 and over), veteran status, uniformed service member status, physical or mental disability, genetic information (including testing or characteristics) or another protected status as defined by local, state, or federal law, as applicable.

Qualified individuals with a disability may be entitled to reasonable accommodation under the Americans with Disabilities Act.

If you require a reasonable accommodation during the application or employment process, please send an email to: rar@oreillyauto.com or call (800) 471-7431 option , and provide your requested accommodation, and position details.

Your first job at O’Reilly Auto Parts is just the beginning! From a comprehensive benefits and compensation package to a rewarding and positive work environment, your leaders will support you and foster your development so you can grow with the company. Our promote‑from‑within philosophy means our top leaders worked their way up – and you can, too. Learn more about our culture, benefits, and history at oreillyauto.com/careers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (Onsite Role)
Senior Site Reliability Engineer (Onsite Role)

O'Reilly Auto Parts • Springfield (MO)

On-site
USD 120,000 - 180,000
Competitive wages
401k with employer contributions
Medical, Dental, Vision insurance
+3
Principal Platform Engineer (Google Cloud) ( Onsite Role)
Principal Platform Engineer (Google Cloud) ( Onsite Role)

O'Reilly Auto Parts • Springfield (MO)

On-site
USD 140,000 - 180,000
Stock Purchase Plan & 401k
Medical, Dental, & Vision Insurance
Tuition Educational Assistance
+1
Sr Platform Engineer (Google Cloud) (Onsite Role)
Sr Platform Engineer (Google Cloud) (Onsite Role)

O'Reilly Auto Parts • Springfield (MO)

On-site
USD 120,000 - 170,000
Competitive wages
401k with employer contributions
Medical, dental, vision insurance
Sr Platform Engineer (Google Cloud) (Onsite Role)
Sr Platform Engineer (Google Cloud) (Onsite Role)

O'Reilly Automotive Stores • Springfield (MO)

On-site
USD 120,000 - 160,000
Stock Purchase Plan
401k with Employer Contributions
Medical, Dental, & Vision Insurance
+2
Platform Engineer II (Virtualization) (Onsite Role)
Platform Engineer II (Virtualization) (Onsite Role)

O'Reilly Auto Parts • Springfield (MO)

On-site
USD 100,000 - 150,000
Competitive wages
Paid time off
401k with match
+6
Senior System Engineer - Kafka/Messaging
Senior System Engineer - Kafka/Messaging

O'Reilly Auto Parts • Springfield (MO)

On-site
USD 120,000 - 190,000
Competitive wages
Stock Purchase Plan
401k with employer contributions
+4
IT Support Analyst I (Onsite Position)
IT Support Analyst I (Onsite Position)

O'Reilly Technology Services, Inc. • United States

On-site
USD 42,000 - 62,000
Competitive wages
401k with employer contributions
Starting Day One medical, dental, and视
+3
IT Support Manager (Onsite Position)
IT Support Manager (Onsite Position)

O'Reilly Technology Services, Inc. • United States

On-site
USD 90,000 - 120,000
Manager AI Enablement
Manager AI Enablement

O'Reilly Technology Services, Inc. • United States

On-site
USD 120,000 - 190,000
Stock purchase plan
401k with employer contributions
Medical, dental, & vision insurance
+3
Material Handler II - Weekend Forklift
Material Handler II - Weekend Forklift

Ozark Automotive Distributors, Inc. • Des Moines (IA)

On-site
USD 32,000 - 42,000
Competitive wages
Paid time off
Stock purchase plan
+5