Distributed Systems Engineer - Fault-Tolerant Infrastructure

JobCubby

California, Northern (MO, KY)

Hybrid

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple is seeking an experienced software engineer to join a team focused on building fault‑tolerant, hyper‑scale cloud network systems. You will design and implement sensing, decision, and remediation layers to detect anomalies and automate recovery in real time.

The role emphasizes strong fundamentals in distributed systems, control theory, graph algorithms, and hands‑on experience with production‑grade infrastructure. Early‑career engineers with deep technical depth are welcome.

Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
  • 4–6+ years of professional software engineering experience designing, building, and operating production‑grade distributed systems and backend infrastructure.
  • Deep foundation in computer science fundamentals, including distributed systems architecture, concurrency models, graph algorithms, and systems design.
  • Strong proficiency in at least one systems‑level or high‑performance language, such as Go, C++, Rust, or Python.
  • Direct experience designing and building fault‑tolerant mechanisms, including automated self‑healing, active remediation, circuit breaking, load shedding, and blast‑radius mitigation for network services.
  • Demonstrated ability to model complex failure domains, handle network partitions and split‑brain scenarios, and reason rigorously about system behavior under extreme load and degradation.
  • Practical experience with chaos engineering, fault injection, simulation‑based testing, and stress testing in production or staging environments.
  • Track record of technical ownership, including authoring design documents, driving code reviews, and leading post‑mortem root cause analyses.

Responsibilities

  • Build fault‑tolerant, self‑healing, scalable network services for massive deployments.
  • Design sensing, decision, and remediation layers to detect and remediate failures with minimal human intervention.
  • Develop automated remediation actions in live production infrastructure without introducing new risk.
  • Collaborate across teams to translate research ideas into production systems that run at enormous scale.

Skills

Go
C++
Rust
Python
Distributed systems
Chaos engineering
Fault-tolerant design
Graph algorithms
Concurrency models

Education

Bachelor’s degree in CS/CE/EE or equivalent practical experience
Master’s or Ph.D. in CS/Distributed Systems/Networking

Tools

eBPF
XDP
OVS

Job description

Apple is seeking an experienced software engineer to join a team focused on building fault‑tolerant, hyper‑scale cloud network systems. You will design and implement sensing, decision, and remediation layers to detect anomalies and automate recovery in real time.

The role emphasizes strong fundamentals in distributed systems, control theory, graph algorithms, and hands‑on experience with production‑grade infrastructure. Early‑career engineers with deep technical depth are welcome.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Fault-Tolerant Infra Systems Engineer
Fault-Tolerant Infra Systems Engineer

Socket.dev • Cupertino (CA)

On-site
USD 210,000 - 280,000
Cloud Fault-Tolerant Network Engineer
Cloud Fault-Tolerant Network Engineer

Apple Inc. • San Francisco (CA)

On-site
USD 185,000 - 325,000
Comprehensive medical and dental
Retirement benefits
Stock programs
+1
Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)
Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)

JobCubby • California (MO), Northern (KY)

Hybrid
USD 150,000 - 210,000
Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)
Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)

Socket.dev • Cupertino (CA)

On-site
USD 210,000 - 280,000
Network Reliability Engineer - Global Cloud Infra
Network Reliability Engineer - Global Cloud Infra

Socket.dev • California (MO)

On-site
USD 160,000 - 210,000
Network Reliability Engineer, Infrastructure Services
Network Reliability Engineer, Infrastructure Services

Socket.dev • California (MO)

On-site
USD 160,000 - 210,000
Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)
Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)

Apple Inc. • San Francisco (CA)

On-site
USD 185,000 - 325,000
Comprehensive medical and dental
Retirement benefits
Stock programs
+1
Senior Network Reliability Engineer, Cloud & Global Scale
Senior Network Reliability Engineer, Cloud & Global Scale

Apple Inc. • San Francisco (CA)

On-site
USD 185,000 - 325,000
Stock programs
Relocation assistance
Comprehensive benefits
Senior Data Center Network Architect - Scale & Reliability
Senior Data Center Network Architect - Scale & Reliability

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 325,000
Comprehensive medical and dental
Employee stock programs
Relocation assistance
Senior Distributed Systems Engineer - Scale & Resiliency
Senior Distributed Systems Engineer - Scale & Resiliency

Apple Inc. • Cupertino (CA)

On-site
USD 147,400 - 272,100
Comprehensive medical and dental coverage
Employee stock programs
Educational reimbursement