Staff Software Engineer, Reliability

LinkedIn

Mountain View (CA)

Hybrid

USD 156,000 - 255,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

LinkedIn is building a world-class Site Health Platform within its Reliability Infrastructure group. The role is based in Mountain View, CA with a hybrid work arrangement and exposure to diverse technologies and data centers globally.

You will own incident lifecycle tooling, drive proactive reliability, and translate incident data into actionable guidance for engineering and leadership. This role emphasizes scalable platforms and measurable improvements.

Qualifications

  • 6+ years in software development or reliability engineering.
  • Experience leading technical projects and architectural leadership.
  • Experience with Go, Python or Java and production systems.
  • Full-stack experience with React.js and data pipelines.

Responsibilities

  • Design and evolve core incident management platforms across many services.
  • Serve on-call, triage incidents and coordinate rapid resolutions.
  • Transform incident data into actionable intelligence using AI/LLM.
  • Build analytics to surface systemic reliability risks and dependencies.
  • Enable fleet-wide stress testing of data centers and capacity.
  • Raise reliability standards and ensure durable, scalable improvements.
  • Influence architecture, SLOs and reliability practices at scale.

Skills

Go
Python
Java
React.js
Distributed Systems
Reliability Engineering

Education

Bachelor’s in CS/Engineering
MS/PhD preferred for Staff

Tools

Vector databases
LLM-based systems

Job description

LinkedIn is the world's largest professional network, built to create economic opportunity for every member of the global workforce. Our products help people make powerful connections, discover exciting opportunities, build necessary skills, and gain valuable insights every day. We're also committed to providing transformational opportunities for our own employees by investing in their growth. We aspire to create a culture that's built on trust, care, inclusion, and fun – where everyone can succeed.

Join us to transform the way the world works.

Job Description

This role will be based in Mountain View, CA.

At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team.

Site Health Platform sits at the core of LinkedIn’s Reliability Infrastructure organization, with a primary focus on the end-to-end incident management ecosystem. Our mission is for every member and customer to experience LinkedIn as \"always on\", every engineer to benefit from a more insightful and proactive site-wide reliability ecosystem, and every business and product owner to be well-informed about service disruptions as they occur.

We own the full incident lifecycle across thousands of services and multiple regions, from incident response and mitigation, through problem management and post-incident learning. The platforms we build are the backbone of how LinkedIn detects issues, coordinates incident response, captures context, and turns outages and near misses into structured, actionable insights.

By transforming incidents into data and learnings, we enable teams to systematically improve reliability over time. Our work informs engineering priorities, infrastructure investments, capacity planning, and executive decision-making, ensuring the network is dependable when it matters most.

You will be exposed to many different technologies, architectures, and systems hosted in state-of-the-art data centers across the globe.

Responsibilities
  • Designing and evolving the core incident management platforms that power LinkedIn’s full incident lifecycle, from detection and response to problem management and prevention, across thousands of services and teams.
  • Serving in a critical on-call rotation, providing expert incident triage and coordination during high-severity outages. Partnering closely with service owners and product teams to diagnose issues quickly, mitigate member impact, and drive timely resolution under pressure.
  • Transforming raw, unstructured incident data into clear, actionable intelligence using AI and LLM-based systems, including automated summarization, classification, root cause signals, and mitigation recommendations.
  • Building analytics and insights that surface systemic reliability risks, recurring failure patterns, and cross-service dependencies, enabling org-level prioritization rather than isolated, service-by-service fixes.
  • Building platforms and tools that enable realistic, fleet-wide stress testing of data center and regional capacity, validating incident readiness across dependencies, traffic patterns, and growth scenarios before they impact a significant production outage.
  • Driving consistency, clarity, and quality in how incidents are declared, managed, reviewed, and learned from, raising the reliability bar across a large, fast-moving engineering organization.
  • Influencing service architecture, SLOs, and reliability standards through platforms, data, and technical leadership, ensuring improvements are durable, measurable, and adopted at scale.
Basic Qualifications
  • Bachelor’s degree in Computer Science, Engineering, or related technical field or equivalent practical experience. Many postings also prefer or require an advanced degree (MS/PhD) for Staff-level roles.
  • 6+ years of professional experience in software development, distributed systems, or reliability engineering.
  • Experience leading technical projects/providing architectural leadership
  • Experience building products and operating large-scale distributed systems.
  • Experience with two or more backend languages such as Go, Python or Java with a track record of owning complex production systems.
  • Full-stack engineering experience, including building user-facing web applications and operational dashboards using modern frontend frameworks such as React.js, along with backend APIs and data pipelines.
  • Understanding of web development fundamentals including API design, performance, accessibility and building intuitive interfaces for engineers and operational users.
  • Understanding of reliability engineering principles, incident management, observability and operating systems under failure conditions.
  • Demonstrated ability to lead technical design across teams, influence architecture beyond direct ownership and drive adoption through well-designed platforms.
  • Debugging and root cause analysis skills, with the ability to communicate complex technical findings clearly to engineers, partners and leadership.
Preferred Qualifications
  • Experience applying AI or LLM-based techniques to operational or incident data, including automated summarization, classification, root cause hypothesis generation or reliability recommendations.
  • Familiarity with vector databases and retrieval-based systems used to power context-aware analytics, search or agentic workflows.
  • Frontend engineering experience beyond basic UI, including building data-dense, high-signal interfaces for engineers using React.js, modern state management and visualization libraries.
  • Experience designing end-to-end full-stack systems where frontend, backend, data and reliability concerns are considered holistically.
  • Background in building internal developer platforms, observability tools, or incident response systems used at scale.
  • A demonstrated ability to simplify complex workflows, reduce operational toil and replace manual processes with well-designed automation.
Suggested Skills
  • High Severity Incident Response
  • Production Troubleshooting & Root Cause Analysis
  • Distributed Systems and Linux fundamentals
  • Software Development (Go / Python / Java) & Architecture
  • Observability & Incident Detection
Additional Information

LinkedIn is committed to fair and equitable compensation practices.

The pay range for this role is $156,000 to $255,000. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience, certifications, and specific work location. This may be different in other locations due to differences in the cost of labor.

The total compensation package for this position may also include annual performance bonus, stock, benefits and/or other applicable incentive compensation plans. For more information, visit https://careers.linkedin.com/benefits.

Equal Opportunity Statement

We seek candidates with a wide range of perspectives and backgrounds and we are proud to be an equal opportunity employer. LinkedIn considers qualified applicants without regard to race, color, religion, creed, gender, national origin, age, disability, veteran status, marital status, pregnancy, sex, gender expression or identity, sexual orientation, citizenship, or any other legally protected class.

LinkedIn is committed to offering an inclusive and accessible experience for all job seekers, including individuals with disabilities. Our goal is to foster an inclusive and accessible workplace where everyone has the opportunity to be successful.

If you need a reasonable accommodation to search for a job opening, apply for a position, or participate in the interview process, connect with us at accommodations@linkedin.com and describe the specific accommodation requested for a disability-related limitation.

San Francisco Fair Chance Ordinance

Pursuant to the San Francisco Fair Chance Ordinance, LinkedIn will consider for employment qualified applicants with arrest and conviction records.

Pay Transparency Policy Statement

As a federal contractor, LinkedIn follows the Pay Transparency and non-discrimination provisions described at this link: https://lnkd.in/paytransparency.

Global Data Privacy Notice and Compliance Posters for Job Candidates

Please use this link to access documents that provide information about how LinkedIn handles the personal data of employees and job applicants, as well as the E-Verify Participation Notice and the Department of Justice Immigrant and Employee Rights Section Right to Work posters: https://www.linkedin.com/legal/candidate-portal.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer - Applications
Staff Software Engineer - Applications

LinkedIn • Mountain View (CA)

Hybrid
USD 156,000 - 255,000
Manager, Software Engineering, Development Infrastructure
Manager, Software Engineering, Development Infrastructure

LinkedIn • Bellevue (WA)

On-site
USD 200,000 - 268,000
Principal Staff Software Engineer - Compute Infrastructure
Principal Staff Software Engineer - Compute Infrastructure

LinkedIn • California (MO)

On-site
USD 231,000 - 378,000
Senior Software Engineer - Systems and Infrastructure
Senior Software Engineer - Systems and Infrastructure

LinkedIn • Mountain View (CA)

On-site
USD 144,000 - 236,000
Software Engineer - Applications
Software Engineer - Applications

LinkedIn • California (MO)

Hybrid
USD 109,000 - 147,000
Health & wellness programs
Paid time off
Stock options
Senior Software Engineer - Applications
Senior Software Engineer - Applications

LinkedIn • California (MO)

On-site
USD 129,000 - 212,000
Stock options
Benefits
Performance bonus
Software Engineer - Systems and Infrastructure
Software Engineer - Systems and Infrastructure

LinkedIn • California (MO)

On-site
USD 120,000 - 195,000
Annual bonus
Stock options
Benefits
Sr. Software Engineer, Development Infrastructure
Sr. Software Engineer, Development Infrastructure

LinkedIn • Bellevue (WA)

On-site
USD 139,000 - 229,000
Hybrid work model
Health and wellness programs
Stock options and incentive plans
Software Engineer - Web Infrastructure
Software Engineer - Web Infrastructure

LinkedIn • California (MO)

Hybrid
USD 104,000 - 171,000
Health & wellness programs
Time off & leave benefits
Senior Software Engineer, Data Science
Senior Software Engineer, Data Science

LinkedIn • California (MO)

Hybrid
USD 127,000 - 212,000