Staff Engineer, Hardware Reliability

LinkedIn

Sunnyvale (CA)

Hybrid

USD 156,000 - 255,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

LinkedIn seeks a Staff Engineer for the Hardware Capacity Engineering team in Sunnyvale, CA. The role focuses on scaling and sustaining on‑prem data center hardware, building automation, and collaborating with SRE, software, and hardware vendors. Hybrid work model with office days.

Responsibilities include designing tests, benchmarking, onboarding new platforms, cost analysis, and leading fleet upgrades to ensure high reliability and observability at scale.

Qualifications

  • BS in Computer Science, Computer Engineering, or related field, or equivalent practical experience.
  • 6+ years working in Linux-based infrastructure, systems, or hardware engineering.
  • 4+ years of hardware troubleshooting, systems engineering, and performance analysis.
  • Experience developing software or automation for infrastructure at scale.

Responsibilities

  • Collaborate with LinkedIn engineering teams to select hardware platforms for apps.
  • Design test environments; benchmark compute, storage, and power; report results.
  • Qualify and integrate new server platforms end-to-end with vendors.
  • Drive cost analysis and define SLAs with partner teams.
  • Qualify BIOS, BMC, and firmware; lead fleet upgrade programs.
  • Improve fleet reliability via fault detection and remediation.
  • Design and own automation for hardware qualification and fleet health monitoring.
  • Contribute to AI/ML infrastructure performance for GPU platforms.
  • Troubleshoot complex hardware issues and lead incident response.

Skills

Linux
Hardware Troubleshooting
Infrastructure Automation
Software Development

Education

BS in Computer Science, Computer Engineering, or related field

Job description

LinkedIn is the world's largest professional network, built to create economic opportunity for every member of the global workforce. Our products help people make powerful connections, discover exciting opportunities, build necessary skills, and gain valuable insights every day. We're also committed to providing transformational opportunities for our own employees by investing in their growth. We aspire to create a culture that's built on trust, care, inclusion, and fun – where everyone can succeed.

Join us to transform the way the world works.

Job Description

At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team.

This role will be based in Sunnyvale, CA.

We are looking for a highly skilled, self-motivated Staff Engineer to join our Hardware Capacity Engineering (HCE) team and help us scale and sustain the infrastructure that powers LinkedIn. HCE qualifies, integrates, and operates the full range of hardware in our on-prem data centers , such as general-purpose compute, GPU/accelerator, storage, and networking platforms, across a large-scale, multi-vendor, multi-generation fleet. This role spans both bringing new platforms into production and keeping our existing fleet healthy, performant, and reliable, backed by the software, firmware automation, and fleet-health systems the team builds.

In this role, you will identify requirements and the best-suited hardware platform or solution, integrate that solution into our on-prem data center environment, and help operate and continuously improve the existing fleet at scale. You will build software and automation that make the fleet observable, performant, and reliable, and partner closely with SRE, software engineering, AI/ML, and hardware vendors.

Responsibilities

  • Collaborate with LinkedIn engineering teams to collect requirements for LinkedIn applications and provide guidance on selecting the best-suited hardware platforms and solutions across general-purpose compute, GPU/accelerator, storage, and networking.
  • Design test environments and testing scenarios; benchmark compute, storage, and power (e.g., SPEC, SPECpower, etc) and provide detailed analysis of qualification and performance results for varied audiences, including engineers and senior leadership.
  • Qualify and integrate new server platforms and components end-to-end, working with hardware vendors on optimal configurations and driving the full integration process.
  • Work jointly with other teams on cost and TCO analysis for proposed solutions, present them to decision makers, and define SLAs and technical standards with partner teams.
  • Qualify BIOS, BMC, and component firmware. Drive fleet-wide upgrade programs.
  • Support and improve the reliability of our existing large-scale, diverse fleet including fault detection and remediation, firmware management, and OS and security compliance.
  • Design, build (AI-assisted) and own automation for hardware qualification, provisioning, lifecycle, and fleet health monitoring and telemetry analysis.
  • Contribute to AI/ML infrastructure performance and reliability of GPU platforms, InfiniBand/RDMA as part of the team’s broader scope.
  • Troubleshoot complex hardware, firmware, kernel, and platform issues across the fleet, and lead critical production incident response.
Qualifications

Basic Qualifications

  • BS in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
  • 6+ years of experience working in Linux-based infrastructure, systems, or hardware engineering.
  • 4+ years of experience with hardware troubleshooting, systems engineering, and performance analysis.
  • Experience developing software or automation (AI-assisted or otherwise) for infrastructure at scale.

Preferred Qualifications

  • Experience with x86 server architecture and multi-vendor hardware, BMC/BIOS and firmware (IPMI/Redfish).
  • Experience qualifying and integrating new hardware platforms and operating them across a large-scale, diverse, multi-generation fleet.
  • Experience with hardware provisioning, imaging/OS, and lifecycle or inventory systems
  • Experience building fleet health, observability, or reliability tooling (fault detection, SMART and telemetry analysis, data-driven thresholds).
  • Experience benchmarking with common tools such as SPEC, SPECpower, fio, unixbench, and similar.
  • Experience with storage devices and performance engineering (NVMe/SSD/HDD) and/or distributed/parallel filesystems (GPFS, HDFS).
  • Experience with GPU/accelerator platforms and the ML training stack (NCCL/collective communications, CUDA) and high-performance networking (InfiniBand/RDMA, RoCE).
  • Experience with Kubernetes and containerized workloads.
  • Experience working with hardware vendors on both designing a solution and troubleshooting issues.
  • Demonstrated experience putting together summary reports and visual presentations of the results of benchmarks and performance tests for technical and executive audiences.
  • Familiarity with HPC/Machine Learning environments and solutions.

Suggested Skills

  • Hardware Qualification & Integration
  • Firmware / BMC (BIOS, IPMI/Redfish)
  • Fleet Reliability & Observability
  • GPU / AI Infrastructure

You will Benefit from our Culture:

We strongly believe in the well-being of our employees and their families. That is why we offer generous health and wellness programs and time away for employees of all levels.

LinkedIn is committed to fair and equitable compensation practices.

The pay range for this role is $156,000 to $255,000. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience, certifications, and specific work location. This may be different in other locations due to differences in the cost of labor.

The total compensation package for this position may also include annual performance bonus, stock, benefits and/or other applicable incentive compensation plans. For more information, visit https://careers.linkedin.com/benefits.

Additional Information

Equal Opportunity Statement

We seek candidates with a wide range of perspectives and backgrounds and we are proud to be an equal opportunity employer. LinkedIn considers qualified applicants without regard to race, color, religion, creed, gender, national origin, age, disability, veteran status, marital status, pregnancy, sex, gender expression or identity, sexual orientation, citizenship, or any other legally protected class.

LinkedIn is committed to offering an inclusive and accessible experience for all job seekers, including individuals with disabilities. Our goal is to foster an inclusive and accessible workplace where everyone has the opportunity to be successful.

Reasonable accommodations are modifications or adjustments to the application or hiring process that would enable you to fully participate in that process. Examples of reasonable accommodations include but are not limited to:

  • Documents in alternate formats or read aloud to you
  • Having interviews in an accessible location
  • Being accompanied by a service dog
  • Having a sign language interpreter present for the interview

A request for an accommodation will be responded to within three business days. However, non-disability related requests, such as following up on an application, will not receive a response.

San Francisco Fair Chance Ordinance

Pursuant to the San Francisco Fair Chance Ordinance, LinkedIn will consider for employment qualified applicants with arrest and conviction records.

Pay Transparency Policy Statement

As a federal contractor, LinkedIn follows the Pay Transparency and non-discrimination provisions described at this link: https://lnkd.in/paytransparency.

Global Data Privacy Notice and Compliance Posters for Job Candidates

Please use this link to access documents that provide information about how LinkedIn handles the personal data of employees and job applicants, as well as the E-Verify Participation Notice and the Department of Justice Immigrant and Employee Rights Section Right to Work posters: https://www.linkedin.com/legal/candidate-portal.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer - Compute Infrastructure
Staff Software Engineer - Compute Infrastructure

LinkedIn • California (MO)

Hybrid
USD 175,000 - 287,000
Health and wellness programs
Competitive compensation
Stock options
Principal Staff Software Engineer, Systems Infrastructure
Principal Staff Software Engineer, Systems Infrastructure

LinkedIn • California (MO)

Hybrid
USD 226,000 - 369,000
Principal Staff Software Engineer, Physical Infrastructure
Principal Staff Software Engineer, Physical Infrastructure

LinkedIn • California (MO)

Hybrid
USD 231,000 - 378,000
Health and wellness programs
Stock options / equity
Staff Network Engineer
Staff Network Engineer

LinkedIn • California (MO)

Hybrid
USD 152,000 - 248,000
Generous health and wellness programs
Annual performance bonus
Staff Software Engineer
Staff Software Engineer

LinkedIn • Mountain View (CA)

Hybrid
USD 152,000 - 248,000
Sr. Technical Program Manager
Sr. Technical Program Manager

LinkedIn • Mountain View (CA)

Hybrid
USD 117,000 - 193,000
Senior Software Engineering Manager, Reliability
Senior Software Engineering Manager, Reliability

LinkedIn • California (MO)

Hybrid
USD 182,000 - 304,000
Sr. Staff Software Engineer, Systems Infrastructure
Sr. Staff Software Engineer, Systems Infrastructure

LinkedIn • California (MO)

Hybrid
USD 198,000 - 326,000
Annual performance bonus
Stock options
Comprehensive benefits package
Staff Software Engineer - Systems Infrastructure
Staff Software Engineer - Systems Infrastructure

LinkedIn • California (MO)

Hybrid
USD 175,000 - 287,000
Generous health and wellness programs
Flexible work arrangements
Inclusive culture and support for personal growth
Associate Engineer, Data Center
Associate Engineer, Data Center

LinkedIn • Manassas (VA)

On-site
USD 74,000 - 122,000