Site Reliability Engineer - Kafka

Socket.dev

Seattle (WA)

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple Service Engineering – Data Streaming SRE team is seeking Site Reliability Engineers to design and automate distributed systems in production, focusing on low latency data services and scalable platforms.

You will contribute to Kafka deployment infrastructure, build tooling for automation, monitoring, and on-call incident response while collaborating with cross-functional teams across Apple services to improve reliability and performance.

Qualifications

  • 5+ years of experience supporting internet-facing production services and distributed systems via deployments, On Call and Incident Management.
  • 5+ years of experience running large scale infrastructure with heavy automation tooling.
  • 5+ years of experience troubleshooting and performance deep dive analysis.
  • Experience with Kubernetes in production.
  • Experience deploying in and running on Datacenter and Cloud architectures; design of multi-datacenter systems and WANs.
  • Proactive, self-motivated, quick to learn new technologies.
  • Experience developing and troubleshooting distributed systems and database storage engines.
  • Experience with AWS, GCP and Terraform.

Responsibilities

  • Develop and maintain Kafka deployment infrastructure and related tooling.
  • Build automation for deployment, monitoring, and alerting dashboards.
  • Collaborate across teams to define metrics, targets, and optimization opportunities.
  • Contribute to safety, stability, performance, and scaling of data services.

Skills

Distributed systems
Incident management
Automation tooling
Troubleshooting
On-call experience
Self-motivation / fast learner
Cross-functional collaboration

Tools

Kafka
Terraform
AWS
GCP
Go
Java
Python

Job description

The Apple Service Engineering – Data Streaming SRE team is looking for Site Reliability Engineers with experience developing processes, tools, and automation for managing distributed systems in production environments. Our SRE team combines software engineering, systems engineering, and Devops practices to build and run large-scale, massively distributed, fault-tolerant systems. Our software ensures that Apple's services are reliable, scalable, and secure, and we leverage both open-source and homegrown technologies to provide managed data infrastructure services. You will help build next-generation Kafka infrastructure and platform services, collaborating cross-functionally with various ASE teams—from store and commerce to search and recommendations. You'll create platforms that can rapidly scale to serve data with very low latencies. You should be someone who isn't afraid to question assumptions, thrives as a collaborative partner under tight deadlines, and tackles complex problems with elegant technical solutions.

DESCRIPTION

The Data Service SRE team develops applications and tooling that are safe, reliable, scalable, and fast. This work requires an innovative spirit and an extraordinary degree of care and difficulty in engineering. Team members contribute to all major components of Kafka deployment infrastructure, including maintenance automation, control plane enhancements, monitoring and alerting tooling/dashboards, advanced deployment architecture, focused on safety, stability, performance, and scaling. Come join us at Apple Services Engineering and help us deliver services and applications that are fluid and responsive. You will collaborate with engineers from across Apple to define the metrics, set targets, uncover optimization opportunities, and ship a service that will delight our customers. This role is for engineers who enjoy deep technical engineering that spans large cross-organizational projects. Your openness to learning and implementing new technologies will contribute to the continuous evolution of our organization. Good ideas are valued and rewarded.

MINIMUM QUALIFICATIONS
  • 5 or more years of experience in support of internet-facing production services and distributed systems via deployments, On Call and Incident Management.
  • 5 or more years of experience running large scale infrastructure with a heavy reliance on automation tooling
  • 5 or more years of experience troubleshooting and performance deep dive analysis
  • Real operational experience managing services at scale on Kubernetes
  • Proficient in one or more of the following programming languages: Java, Go (golang), Python
  • Operational experience deploying in and running on Datacenter and Cloud architectures (networking topologies, host placement strategies, and failure modes); design of multi-datacenter systems; failure domains; and wide-area networking
  • Self motivated, inquisitive with an aptitude to learn new technologies quickly and effectively
  • Demonstrated expertise developing and troubleshooting distributed systems and database storage engines
  • Experience developing critical internet services and/or platform infrastructure
  • Experience with AWS, GCP and IaC such as Terraform
PREFERRED QUALIFICATIONS
  • Experience managing messaging services such as Kafka or other Data services
  • Proficient in Java, Go (golang) & Python
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Kafka Platform SRE for Large-Scale Data Streaming
Kafka Platform SRE for Large-Scale Data Streaming

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 190,000
Site Reliability Engineer, Apple Data Platform / Big Data Platform
Site Reliability Engineer, Apple Data Platform / Big Data Platform

Socket.dev • Austin (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Socket.dev • Austin (TX)

On-site
USD 140,000 - 220,000
Site Reliability Engineer, Apple Data Platform
Site Reliability Engineer, Apple Data Platform

Socket.dev • Austin (TX)

On-site
USD 150,000 - 230,000
Site Reliability Engineer, Apple Data Platform - AI/ML Platform
Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Apple • Austin (TX)

On-site
USD 120,000 - 180,000
SRE Kafka Lead
SRE Kafka Lead

Compunnel, Inc. • Orlando (FL)

Hybrid
USD 120,000 - 150,000
Site Reliability Engineer, AiDP Production Engineering
Site Reliability Engineer, AiDP Production Engineering

Apple Inc. • Austin (TX)

On-site
USD 140,000 - 170,000
Senior Kafka SRE Engineer
Senior Kafka SRE Engineer

Charles Schwab Inc. • Austin (TX)

On-site
USD 140,000 - 180,000
Senior Software Engineer, Apple Data Platform
Senior Software Engineer, Apple Data Platform

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 190,000
Apple Services Engineering (ASE) Compute - Software Engineering Manager
Apple Services Engineering (ASE) Compute - Software Engineering Manager

Socket.dev • Cupertino (CA)

On-site
USD 190,000 - 240,000