Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)

Apple Inc.

San Francisco (CA)

On-site

USD 185,000 - 325,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive medical and dental
Retirement benefits
Stock programs
Tuition reimbursement

Job summary

Apple Inc. is seeking a Software Engineer for Infrastructure Services (Cloud Network Fault Tolerance) in the San Francisco Bay Area. Join a team building fault-tolerant, distributed systems that automate detection, diagnosis, and remediation across a massive cloud network.

You will design sensing, decision, and remediation layers for real-time resilience at scale. The role emphasizes deep formal reasoning about systems, with opportunities to work on research-inspired problems, develop live

Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
  • 4 to 6+ years of professional software engineering experience designing, building, and operating production-grade distributed systems and backend infrastructure.
  • Foundations in distributed systems architecture, concurrency, graph algorithms, and systems design.

Responsibilities

  • Design and build automated detection, diagnosis, and remediation systems for network faults across a hyperscale cloud environment.
  • Develop the decision logic and safety mechanisms for automated corrective actions on production infrastructure.
  • Analyze failure patterns and systemic weaknesses to eliminate or mitigate them.
  • Build simulation, chaos testing, and fault-injection tooling to surface weaknesses before incidents occur.
  • Apply formal reasoning about distributed systems and control theory to production software.

Skills

Distributed systems
Fault-tolerant design
Go
C++
Rust
Python
Chaos engineering
System design
Code reviews

Education

Bachelor’s degree in Computer Science or related field
Master’s or Ph.D. preferred

Tools

eBPF/XDP
Linux networking
OVS
Raft/Paxos

Job description

Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)

San Francisco Bay Area, California, United States Software and Services

Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization.IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless.”Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here.Infrastructure Services is part of IS&T and the foundation of Apple's global network operations — managing data center equipment and systems to deliver compute, storage, and networking services for teams across Apple, including its internal developer community. From individual facilities to a worldwide network, Infrastructure Services ensures the technology underneath everything works without question.This is not a traditional network engineering role. This is a software engineering role for builders who love deep technical challenges, have strong fundamentals in distributed systems and fault-tolerant design, and want to turn cutting-edge ideas into production systems running at massive scale. We welcome early-career engineers with exceptional technical depth, if you have the intellectual horsepower and hunger to solve problems that don't have textbook answers yet, we want to talk to you.

Description

You will join a team chartered to make Apple's hyper-scale cloud network dramatically more fault-tolerant by building systems that detect and remediate failures automatically, with minimal or no human intervention. This means designing and building the sensing, decision-making, and remediation layers that allow the network to identify anomalies, localize root cause, and take corrective action in real time.You'll work on problems like: how do you detect a degrading network path before it causes an outage? How do you safely automate remediation actions on live production infrastructure without introducing new risk? How do you build a system that learns from past incidents to prevent recurrence? These are open, hard problems, and you will be expected to bring rigorous thinking, creativity, and strong engineering execution to solve them.This role is ideal for someone who loves taking a deep, formal understanding of distributed systems, control theory, graph algorithms, or fault-tolerant design patterns and turning it into resilient, real-world software running in production at enormous scale. You should be comfortable reading research papers and translating relevant ideas into working systems, while also being pragmatic about the constraints of operating in a live, hyper-scale environment.If you have exceptional technical depth from your academic background, competitive programming, research experience, or personal projects, and you're hungry to work on problems most companies aren't yet equipped to solve, this team is built for you.

Responsibilities
  • Design and build automated detection, diagnosis, and remediation systems for network faults across a hyperscale cloud environment
  • Develop the decision logic and safety mechanisms that allow automated systems to take corrective action on production infrastructure with confidence
  • Analyze failure patterns and systemic weaknesses across the network and design software solutions to eliminate or mitigate them
  • Build simulation, chaos testing, and fault-injection tooling to proactively surface weaknesses before they cause real incidents
  • Apply formal reasoning about distributed systems, failure modes, and control systems to build robust, production-grade software
  • Translate emerging research and industry innovation in resilience engineering into practical systems that work at Apple's scale
  • Bring new ideas and approaches to the team, this is a group expected to push the boundaries of what's possible in network resiliency
  • Prototype and experiment rapidly, validating ideas against real production data and failure scenarios
  • Partner with Cloud Network production and reliability teams to understand real-world failure modes and operational constraints
  • Work with peer infrastructure teams to ensure self-healing systems integrate safely and effectively across the broader platform
Minimum Qualifications
  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience
  • 4 to 6+ years of professional software engineering experience designing, building, and operating production-grade distributed systems and backend infrastructure
  • Deep foundation in computer science fundamentals, including distributed systems architecture, concurrency models, graph algorithms, and systems design
  • Strong proficiency in at least one systems-level or high-performance language, such as Go, C++, Rust, or Python
  • Direct experience designing and building fault-tolerant mechanisms, including automated self-healing, active remediation, circuit breaking, load shedding, and blast-radius mitigation for network services
  • Demonstrated ability to model complex failure domains, handle network partitions and split-brain scenarios, and reason rigorously about system behavior under extreme load and degradation
  • Practical experience with chaos engineering, fault injection, simulation-based testing, and stress testing in production or staging environments
  • Track record of technical ownership, including authoring design documents, driving code reviews, and leading post-mortem root cause analyses
Preferred Qualifications
  • Master’s or Ph.D. in Computer Science, Distributed Systems, Networking, or a related technical discipline
  • Deep domain knowledge in cloud networking architectures, Software-Defined Networking (SDN) control planes, L3/L4 routing protocols, overlay networks, and Linux networking constructs (such as eBPF, XDP, or OVS)
  • Hands-on experience implementing or tuning distributed consensus algorithms (such as Raft or Paxos) and closed-loop control systems or reconciliation controllers
  • Experience architecting high-throughput telemetry pipelines and automated anomaly detection systems for real-time network health analysis
  • Proven ability to digest academic papers and industry research, translating state-of-the-art resilience concepts into production systems
  • History of notable open-source contributions, technical publications, or patent filings in distributed networking and systems reliability
At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $324,800, and your base pay will depend on your skills, qualifications, experience, and location. Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Reliability Engineer, Infrastructure Services
Network Reliability Engineer, Infrastructure Services

Apple Inc. • San Francisco (CA)

On-site
USD 185,000 - 325,000
Stock programs
Relocation assistance
Comprehensive benefits
Sr Software Engineer, Infrastructure Services
Sr Software Engineer, Infrastructure Services

Apple Inc. • Sunnyvale (CA)

Hybrid
USD 216,000 - 293,000
Discretionary RSU awards
Employee stock purchase plan discount
Medical and dental coverage
+2
Data Center Network Architect
Data Center Network Architect

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 325,000
Comprehensive medical and dental
Employee stock programs
Relocation assistance
Senior Security Engineer, Data Center Network
Senior Security Engineer, Data Center Network

Apple Inc. • San Francisco (CA)

On-site
USD 176,000 - 312,000
Sr. Software Release Engineer, Infrastructure Services
Sr. Software Release Engineer, Infrastructure Services

Apple Inc. • San Francisco (CA)

On-site
USD 185,000 - 325,000
Comprehensive medical & dental
Employee stock programs
Relocation assistance
Sr Software Engineer, Apple Cloud Networking
Sr Software Engineer, Apple Cloud Networking

Apple Inc. • Sunnyvale (CA)

On-site
USD 184,000 - 325,000
Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)
Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)

Socket.dev • Cupertino (CA)

On-site
USD 210,000 - 280,000
Cloud Network Platform Software Engineer
Cloud Network Platform Software Engineer

Apple Inc. • Sunnyvale (CA)

On-site
USD 150,400 - 277,600
Stock options
Medical and dental coverage
Retirement benefits
+2
Sr. Linux Engineer - DNS and Global Server Load Balancing, Infrastructure Services
Sr. Linux Engineer - DNS and Global Server Load Balancing, Infrastructure Services

Apple Inc. • Sunnyvale (CA)

Hybrid
USD 216,000 - 325,000
Stock options
Medical coverage
Retirement benefits
+1
Network Architect Lead, Infrastructure Engineering
Network Architect Lead, Infrastructure Engineering

Apple Inc. • Cupertino (CA)

On-site
USD 207,400 - 311,700
Stock options
Employee Stock Purchase Plan
Medical & dental coverage
+2