Service Reliability Engineer, G&A Solutions Engineering

Apple

Austin (TX)

On-site

USD 140,000 - 200,000

Full time

13 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Apple's G&A Solutions Engineering team in Austin is seeking a Service Reliability Engineer to ensure the reliability, scalability, and performance of mission-critical services.

You will monitor, lead incident response, automate repetitive tasks, and work with engineers, DBAs, data engineers, and network specialists to improve reliability.

Qualifications

  • 4+ years in SRE/production support for large-scale services.
  • Strong programming skills in Python/Java/Go and scripting (Bash/PowerShell).
  • Experience with cloud platforms and Kubernetes/Docker.
  • Hands-on monitoring/alerting with Prometheus, Grafana, Splunk, or Datadog.

Responsibilities

  • Proactively monitor service performance and identify bottlenecks.
  • Lead incident response and conduct RCA to prevent recurrence.
  • Develop automation to reduce manual tasks and improve resilience.
  • Collaborate with developers and DBAs to design for operability and scale.
  • Create and maintain runbooks, SLIs, and documentation.

Skills

Languages (Python/Java/Go)
Scripting (Bash/PowerShell)
Cloud (AWS/Azure/GCP)
Kubernetes
Docker
Monitoring (Prometheus/Grafana/Datadog
Linux admin

Education

Bachelor's degree in CS or related

Tools

Prometheus
Grafana
Splunk
Datadog

Job description

Summary

Do you have a passion for ensuring the reliability, scalability, and performance of critical services? Are you a highly motivated and expert engineer with a strong understanding of Site Reliability Engineering (SRE) principles and a desire to automate and improve processes? Join Apple's General and Administrative (G&A) Solutions Engineering team as a Service Reliability Engineer and play a vital role in supporting our global, mission-critical production systems.

Description

You'll be at the forefront of maintaining the health, stability, and efficiency of our services, working with a diverse range of technologies and platforms. You will collaborate with Engineers, Data Engineers, DBAs, and network specialists to proactively identify and resolve potential issues, automate repetitive tasks, and drive continuous improvement initiatives. Your expertise will directly impact the reliability of our systems, enabling Apple to deliver innovative products and services to our customers.

Key Responsibilities
  • Proactively monitor service performance, identify potential bottlenecks, and implement solutions to optimize efficiency and resilience
  • Lead incident response efforts, driving rapid resolution and conducting thorough root cause analysis (RCA)
  • Develop and implement automation strategies to streamline operational tasks, improve service resilience, and reduce manual intervention
  • Apply SRE principles to maintain highly reliable and scalable service infrastructure
  • Collaborate closely with development teams to ensure that new services are designed for operational perfection, incorporating best practices for monitoring, alerting, and scalability
  • Contribute to the creation and maintenance of comprehensive documentation, including run-books, service level objectives (SLOs)
  • Participate in on-call rotations, providing 24/7 support for critical services and responding to incidents with a sense of urgency
  • Find opportunities for process improvement and drive initiatives to enhance the efficiency and effectiveness of the service reliability team
  • Champion a culture of continuous learning and knowledge sharing within the team
  • Define and supervise key service level indicators (SLIs) to measure and improve service reliability
Minimum Qualifications
  • 4+ years of experience in a Site Reliability Engineering, production support or related role, supporting large-scale, enterprise-level services
  • Strong proficiency in at least one programming language (e.g., Python, Java, Go) and scripting languages (e.g., Bash, PowerShell)
  • Experience with cloud platforms (e.g., AWS, Azure, GCP) and cloud-native technologies (e.g., Kubernetes, Docker)
  • Hands-on experience with monitoring and alerting tools (e.g., Prometheus, Grafana, Splunk, Datadog)
  • Bachelor's degree in Computer Science or work related equivalent experience
Preferred Qualifications
  • Familiarity with CI/CD pipelines and DevOps practices
  • Experience with database technologies (e.g., MySQL, PostgreSQL, NoSQL databases)
  • Knowledge of ITIL frameworks and incident management processes
  • Experience with vibe coding
  • Understanding of Linux/Unix system administration
  • Experience with configuration management tools (Ansible, Chef, Puppet)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Service Reliability Engineer, G&A Solutions Engineering
Service Reliability Engineer, G&A Solutions Engineering

Apple Inc. • Austin (TX)

On-site
USD 130,000 - 180,000
Teamcenter Site Reliability Engineer, Enterprise Technology Services
Teamcenter Site Reliability Engineer, Enterprise Technology Services

Apple • Austin (TX)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer, Infrastructure Services
Sr. Site Reliability Engineer, Infrastructure Services

Apple • Sunnyvale (CA)

On-site
USD 120,000 - 180,000
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 120,000 - 190,000
Service Reliability Engineer (SRE)
Service Reliability Engineer (SRE)

Apple Inc. • Seattle (WA)

On-site
USD 142,300 - 263,300
Medical and Dental coverage
Retirement benefits
Employee stock purchase plan
+2
Security Site Reliability Engineer - Apple Service Engineering
Security Site Reliability Engineer - Apple Service Engineering

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 309,000
Site Reliability Engineer, Apple Data Platform
Site Reliability Engineer, Apple Data Platform

Apple Inc. • Austin (TX), Northern (KY)

On-site
USD 150,000 - 190,000
Senior Site Reliability Engineer - Apple Services Engineering (ASE) / iCloud at Apple Culver Ci[...]
Senior Site Reliability Engineer - Apple Services Engineering (ASE) / iCloud at Apple Culver Ci[...]

null • Culver City (CA)

On-site
USD 167,000 - 251,000
Comprehensive medical coverage
Dental coverage
Retirement benefits
+2
Site Reliability Engineer (Edge Services), Infrastructure Services
Site Reliability Engineer (Edge Services), Infrastructure Services

Apple Inc. • Elk Grove (CA)

On-site
USD 132,100 - 244,600
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
Senior Site Reliability Engineer - Apple Services Engineering (ASE) / iCloud at Apple Cupertino, CA
Senior Site Reliability Engineer - Apple Services Engineering (ASE) / iCloud at Apple Cupertino, CA

null • Cupertino (CA)

On-site
USD 175,800 - 264,200
Employee stock programs
Medical and dental coverage
Tuition reimbursement
+1