Senior AI Compute Systems Engineer – Hardware Reliability

Graphcore

West of England

Hybrid

GBP 65,000 - 90,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Flexible working
Generous leave
Pension support
Phantom equity
On-site barista and free food
Life insurance and income protection
Private medical insurance
Cycle to work scheme

Job summary

Graphcore is looking for a hands-on engineer to own reliability across AI compute hardware. You will lead diagnostic and engineering support in lab and data centre environments, improving bring-up, validation and troubleshooting for AI compute platforms.

The role requires strong debugging of server hardware, experience with HPC/rack-scale infra, and the ability to drive structured root cause analysis across cross-functional teams. Applicants should communicate clearly and own issues end-to-end.

Qualifications

  • Strong knowledge of server hardware architectures and board-level debugging.
  • Experience isolating failures using system logs, telemetry, power data and thermal metrics.
  • Hands-on experience with HPC systems, AI compute platforms or rack-scale infrastructure.
  • Ability to lead structured root cause analysis and propose practical corrective actions.
  • Confidence working across engineering, platform and data centre teams.
  • Clear written and verbal communication, with the judgement to guide junior engineers.

Responsibilities

  • Lead advanced operational, diagnostic and engineering support across lab and data centre environments.
  • Improve hardware bring-up, validation and troubleshooting for AI compute platforms.
  • Diagnose failures across server blades, racks, power systems, thermal behaviour, network configuration and BIOS/BMC issues.
  • Turn root causes into corrective actions and better operating practice.
  • Support new platforms from early bring-up through deployment.

Skills

Server hardware
Board-level debugging
Telemetry data
HPC systems
AI compute platforms
Root cause analysis
Cross-team collaboration
Communication

Job description

Graphcore is looking for a hands-on engineer to own reliability across AI compute hardware. You will lead diagnostic and engineering support in lab and data centre environments, improving bring-up, validation and troubleshooting for AI compute platforms.

The role requires strong debugging of server hardware, experience with HPC/rack-scale infra, and the ability to drive structured root cause analysis across cross-functional teams. Applicants should communicate clearly and own issues end-to-end.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Compute Systems Reliability Engineer
Senior AI Compute Systems Reliability Engineer

Applied Methods Ltd • West of England

Hybrid
GBP 65,000 - 95,000
Senior AI Compute Reliability Engineer
Senior AI Compute Reliability Engineer

Jackalope Digital LLC • West of England

Hybrid
GBP 70,000 - 90,000
Flexible working
Generous leave
Pension matching
+5
Senior AI Compute Systems Engineer
Senior AI Compute Systems Engineer

Graphcore • West of England

On-site
GBP 75,000 - 110,000
Flexible working
Generous leave
Pension matching
+7
Staff Hardware Engineer — AI Compute & Diagnostics Lead
Staff Hardware Engineer — AI Compute & Diagnostics Lead

Graphcore • West of England

On-site
GBP 85,000 - 120,000
Flexible working
Generous leave
Retirement planning support
+5
Hardware Reliability Engineer – AI Compute
Hardware Reliability Engineer – AI Compute

Graphcore • West of England

Hybrid
GBP 80,000 - 110,000
Flexible working
Generous leave
Pension matching
+4
Senior AI Compute Hardware Engineer
Senior AI Compute Hardware Engineer

Applied Methods Ltd • West of England

Hybrid
GBP 60,000 - 90,000
Lead Hardware Diagnostics Engineer for AI Compute Platforms
Lead Hardware Diagnostics Engineer for AI Compute Platforms

Graphcore • West of England

On-site
GBP 70,000 - 110,000
Flexible working
Generous leave
Pension matching
+5
Senior Semiconductor Reliability Engineer for AI Silicon
Senior Semiconductor Reliability Engineer for AI Silicon

Graphcore • West of England

On-site
GBP 95,000 - 130,000
Flexible working
Generous leave
Pension matching
+3
Staff Hardware Engineer New Bristol, UK
Staff Hardware Engineer New Bristol, UK

Graphcore • West of England

Hybrid
GBP 80,000 - 110,000
Flexible working
Generous leave
Pension matching
+4
Senior Systems Engineer
Senior Systems Engineer

Jackalope Digital LLC • West of England

Hybrid
GBP 70,000 - 90,000
Flexible working
Generous leave
Pension matching
+5