Principal Engineer, System Reliability

Ayar Labs

San Jose (CA)

On-site

USD 185,000 - 290,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Ayar Labs is seeking a Principal Engineer, System Reliability to own reliability qualification at the system and fleet level in San Jose, CA. The role focuses on building test infrastructure, generating evidence, and partnering with customers for joint qualification programs.

You will drive reliability planning, standards compliance, and staged qualification through lab and pilot deployments, while translating findings into commitments for partners and customers.

Qualifications

  • 5+ years in fleet/systems reliability engineering for large-scale infrastructure.
  • BS in Electrical Engineering, Computer Engineering, or a related field.
  • Defined and executed statistical reliability demonstration plans.
  • Experience with die-to-die, chip-to-chip, or board-level interconnect standards.

Responsibilities

  • Design, build, and operate test infrastructure that emulates real system-level conditions and workload profiles.
  • Define statistical sample sizes and test durations to demonstrate reliability targets.
  • Ensure test infrastructure and telemetry align with industry standards for external transfer of evidence.
  • Execute staged qualification gates from lab testing to live pilot deployment.
  • Design and run fault-injection and failure-mode tests with customer and partner teams.
  • Provide fleet reliability data for cost-of-ownership and incident response.

Skills

Fleet reliability
System reliability engineering
Test infrastructure
Statistical reliability planning
Customer engagement
Written communication

Education

BS in Electrical/Computer Engineering or related field
MS in Electrical/Computer Engineering or related field

Tools

FPGA-based test systems

Job description

Position: Principal Engineer, System Reliability

Location: San Jose, CA

Job Id: 650

# of Openings: 0

Principle Systems Reliability Engineer

Location: San Jose, CA (On-site)

Ayar Labs is shattering AI data bottlenecks by moving data at the speed of light. As pioneers of co-packaged optics (CPO), we are using light instead of electricity to move data faster, further, and with a fraction of the energy needed to fuel the explosive growth of AI models.

Backed by industry giants like NVIDIA, AMD, Mediatek and Intel and manufactured in partnership with the world’s leading semiconductor ecosystem, Ayar Labs’ co-packaged optics solution is key to unleashing next-generation AI scale-up architectures.

About the Role

Ayar Labs builds optical I/O technology for hyperscale AI infrastructure. This role owns reliability qualification at the system and fleet level: building the test infrastructure and evidence base that proves our products are ready for large-scale deployment, and then partnering directly with customers and integration partners to carry that evidence through joint qualification programs and pilot deployments. You will be responsible for both the engineering (test infrastructure, statistical reliability planning, standards compliance) and the relationship (translating technical evidence into the specific commitments our partners need to move forward).

What You'll Own
  • Test infrastructure and evidence generation: Design, build, and operate custom test infrastructure that emulates real system-level electrical, thermal, and workload conditions ahead of, or independent of, any single customer's specific hardware. This includes sourcing or generating representative workload and power profiles and using them to drive continuous, long-duration test campaigns.
  • Reliability statistics and demonstration planning: Define statistical sample sizes and test durations needed to demonstrate specific reliability and confidence targets, and apply the right methodology to each failure population in the system rather than a single blanket target.
  • Standards and compliance: Ensure that test infrastructure, interfaces, and telemetry are built against relevant industry interconnect and management standards, so evidence generated internally holds up when reviewed by, or transferred to, an external partner's platform.
  • Staged qualification execution: Execute staged qualification gates, from lab-level interoperability and functional testing through environmental and stress testing to live pilot deployment, applying a reliability / availability / serviceability lens throughout rather than treating any one dimension in isolation.
  • Fault injection and lifecycle monitoring: Design and run fault-injection and failure-mode testing, including scenarios that exercise field-serviceable components under representative operating conditions, in partnership with customer and integration-partner operations teams.
  • Economic modeling and fleet reporting: Provide reliability and performance data into cost-of-ownership and total-cost modeling shared with customers, and own ongoing fleet reliability reporting and incident response once products reach production.
  • Cross-functional partnership: Serve as the primary technical point of contact with Tier-1 integration partners and hyperscale customer engineering teams across the qualification lifecycle, from early technical evidence through pilot sign-off and into steady-state operations.
Basic Qualifications
  • 5+ years in fleet/systems reliability engineering for large-scale infrastructure, with a BS in Electrical Engineering, Computer Engineering, or a related field.
  • You've built or operated test infrastructure that validates a component or subsystem's behavior before it is deployed into a customer's actual system, using representative rather than production hardware.
  • You've defined and executed statistical reliability demonstration plans (sample sizes, test durations, confidence and reliability targets) for hardware components, and can explain the methodology behind them, not just apply a lookup table.
  • You've worked with die-to-die, chip-to-chip, or board-level interconnect standards and telemetry or management interfaces relevant to high-speed data center hardware.
  • You've partnered directly with external OEM or hyperscale customer engineering teams on a joint qualification or certification program, and are comfortable owning that relationship technically.
  • Comfort operating a live, continuously running test or fleet environment: on-call coverage, incident response, and building the telemetry pipeline that turns raw sensor data into fleet-level reliability statistics.
  • Strong written and verbal communication skills; you will regularly translate internal engineering data into evidence and documentation for external partners.
Preferred Qualifications
  • MS in Electrical Engineering, Computer Engineering, or a related field.
  • Experience with FPGA-based test or signal-generation systems.
  • Experience characterizing or emulating real compute or network workload behavior for use in a test or validation environment.
  • Background in data-center fleet operations, site-reliability engineering, or customer-facing qualification and certification processes.
  • Familiarity with co-packaged optics, silicon photonics, or other emerging interconnect packaging technologies.

Salary Range: $185,000 - $290,000

Ayar Labs is an Equal Opportunity Employer and is strongly committed to all policies which will afford equal opportunity employment to all qualified persons without regard to age, sex, national origin, race, color, ethnicity, creed, religion, gender identity, sexual orientation, disability, veteran status, or any other characteristic protected by law. It is the policy of Ayar Labs to provide reasonable accommodation when requested by a qualified applicant or employee with a disability, unless such accommodation would cause an undue hardship. Veterans are more than welcome and encouraged to apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Photonics Reliability Architect
Photonics Reliability Architect

Ayar Labs • San Jose (CA)

On-site
Sr. Engineer, Photonics Reliability
Sr. Engineer, Photonics Reliability

Ayar Labs • San Jose (CA)

On-site
Principal Engineer, Photonics Reliability
Principal Engineer, Photonics Reliability

Ayar Labs • San Jose (CA)

On-site
USD 200,000 - 240,000
Principal Engineer, Silicon Validation
Principal Engineer, Silicon Validation

Ayar Labs • San Jose (CA)

On-site
USD 200,000 - 255,000
Sr. Engineer, Silicon Validation
Sr. Engineer, Silicon Validation

Ayar Labs • San Jose (CA)

On-site
USD 160,000 - 192,000
Principal Engineer, SerDes Validation
Principal Engineer, SerDes Validation

Ayar Labs • San Jose (CA)

On-site
USD 180,000 - 230,000
Sr. Engineer, System Validation
Sr. Engineer, System Validation

Ayar Labs • San Jose (CA)

On-site
USD 160,000 - 192,000
Principal Engineer, Digital Verification and Emulation
Principal Engineer, Digital Verification and Emulation

Ayar Labs • San Jose (CA)

On-site
Engineer, Hardware Test Software
Engineer, Hardware Test Software

Ayar Labs • San Jose (CA)

On-site
USD 130,000 - 160,000
Sr. Staff Engineer, Package and Interconnect Reliability
Sr. Staff Engineer, Package and Interconnect Reliability

Ayar Labs • San Jose (CA)

On-site
USD 180,000 - 223,000