Kafka Platform SRE for Large-Scale Data Streaming

Socket.dev

Seattle (WA)

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple Service Engineering – Data Streaming SRE team is seeking Site Reliability Engineers to design and automate distributed systems in production, focusing on low latency data services and scalable platforms.

You will contribute to Kafka deployment infrastructure, build tooling for automation, monitoring, and on-call incident response while collaborating with cross-functional teams across Apple services to improve reliability and performance.

Qualifications

  • 5+ years of experience supporting internet-facing production services and distributed systems via deployments, On Call and Incident Management.
  • 5+ years of experience running large scale infrastructure with heavy automation tooling.
  • 5+ years of experience troubleshooting and performance deep dive analysis.
  • Experience with Kubernetes in production.
  • Experience deploying in and running on Datacenter and Cloud architectures; design of multi-datacenter systems and WANs.
  • Proactive, self-motivated, quick to learn new technologies.
  • Experience developing and troubleshooting distributed systems and database storage engines.
  • Experience with AWS, GCP and Terraform.

Responsibilities

  • Develop and maintain Kafka deployment infrastructure and related tooling.
  • Build automation for deployment, monitoring, and alerting dashboards.
  • Collaborate across teams to define metrics, targets, and optimization opportunities.
  • Contribute to safety, stability, performance, and scaling of data services.

Skills

Distributed systems
Incident management
Automation tooling
Troubleshooting
On-call experience
Self-motivation / fast learner
Cross-functional collaboration

Tools

Kafka
Terraform
AWS
GCP
Go
Java
Python

Job description

Apple Service Engineering – Data Streaming SRE team is seeking Site Reliability Engineers to design and automate distributed systems in production, focusing on low latency data services and scalable platforms.

You will contribute to Kafka deployment infrastructure, build tooling for automation, monitoring, and on-call incident response while collaborating with cross-functional teams across Apple services to improve reliability and performance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Kafka
Site Reliability Engineer - Kafka

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 190,000
Senior Kafka SRE: Real-Time Data Reliability & Automation
Senior Kafka SRE: Real-Time Data Reliability & Automation

Charles Schwab • Austin (TX)

On-site
USD 140,000 - 180,000
On-Site Kafka SRE Engineer: Real-Time Data Reliability
On-Site Kafka SRE Engineer: Real-Time Data Reliability

Charles Schwab Inc. • Austin (TX)

On-site
USD 140,000 - 180,000
Kafka SRE Lead: Reliability & Platform Architecture
Kafka SRE Lead: Reliability & Platform Architecture

Compunnel, Inc. • Orlando (FL)

Hybrid
USD 120,000 - 150,000
SRE Kafka Lead
SRE Kafka Lead

Compunnel, Inc. • Orlando (FL)

Hybrid
USD 120,000 - 150,000
Data Platform SRE — Spark, Flink & Cloud
Data Platform SRE — Spark, Flink & Cloud

Socket.dev • Austin (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Veriipro • San Jose (CA)

On-site
USD 100,000 - 140,000
Remote Real-Time Data Engineer (Kafka/SRE)
Remote Real-Time Data Engineer (Kafka/SRE)

Bright Vision Technologies • Flower Mound (TX)

On-site
USD 135,000 - 160,000
Data-Streaming Software Engineer (Kafka/MSK)
Data-Streaming Software Engineer (Kafka/MSK)

Amazon Web Services (AWS) • Santa Monica (CA)

On-site
USD 143,000 - 195,000
Health insurance
401(k) matching
Paid time off
+1
Senior Software Engineer - Data Streaming & Kafka
Senior Software Engineer - Data Streaming & Kafka

Amazon • Seattle (WA)

On-site
USD 143,700 - 194,400