Principal Engineer

Graphcore

Milpitas, Austin (CA, TX)

On-site

USD 180,000 - 260,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision
401(k) retirement plan
Commuter benefits
Flexible working
Wellness programs

Job summary

Graphcore in Milpitas, CA, is seeking an experienced Principal Engineer to join the System Management team and lead development of interfaces used to manage system state.

You will provide technical leadership, mentor engineers, guide architecture, and drive reliable deployment and operation at scale across hardware and software.

This hands-on role requires strong judgment, cross-team collaboration and a track record of delivering complex engineering initiatives.

Qualifications

  • Bachelor's degree or equivalent practical experience in a relevant subject.
  • Substantial Linux-based infrastructure or distributed systems experience.
  • Experience leading complex engineering initiatives across multiple teams.
  • Ability to influence architecture across teams without formal authority.
  • Strong experience designing RESTful APIs and coding in Go, with Bash/Python for automation.
  • Deep Kubernetes expertise and production workloads experience.
  • Hands-on with IaC, version control and CI/CD tools like Terraform, Ansible, GitLab, and GitHub Actions.
  • Experience with hardware-management interfaces such as Redfish or IPMI.
  • Strong Linux troubleshooting and operational debugging skills.
  • Proven mentoring and coaching across engineers and teams.
  • Excellent communication to align stakeholders on technical outcomes.
  • Experience using AI coding assistants in engineering workflows.
  • Experience building Kubernetes operators and CRs.
  • Familiarity with HPC environments and SLURM/LSF.
  • Experience with virtualization tech and distributed storage.

Responsibilities

  • Translate System Management direction into technical plans and deliverables.
  • Provide architecture and design leadership across teams, documenting trade-offs.
  • Serve as technical authority for System Management initiatives and coordinate plans, risks and decisions.
  • Mentor engineers in system management, deployment automation and production operations.
  • Own key technical outcomes across the full software lifecycle including testing and deployment.
  • Identify reliability and operability improvements across the platform.
  • Collaborate with Hardware, Firmware, Platform Software and Datacenter Ops to diagnose system-level issues.
  • Improve CI/CD, IaC, automated testing and release safety practices.
  • Act as escalation point while creating reusable tooling and automation.

Skills

RESTful APIs
Go
Bash
Python
Linux systems
Technical leadership
Communication

Education

Bachelor's degree or equivalent practical experience

Tools

Kubernetes
Container runtimes
Terraform/OpenTofu
Ansible
GitLab
GitHub Actions
Git
OpenTelemetry
Redfish/IPMI
VSphere/Open vSwitch/KVM/QEMU

Job description

Austin, Texas, United States; Milpitas, California, United States

About us

Graphcore is one of the world's leading innovators in Artificial Intelligence compute.

It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry.

As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world's most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone.

Graphcore's teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives, spanning AI research specialists, silicon designers, software engineers and systems architects.

Job Summary

We are looking for an experienced Principal Engineer to join our System Management team and help lead the development of critical interfaces used by internal and external customers to manage system state. You will provide technical leadership within assigned areas of System Management, guide architecture and implementation choices, mentor engineers and translate broader technical direction into effective execution. This is a hands-on engineering role for someone who can lead complex technical work, improve reliability and operational readiness, and collaborate effectively across multiple engineering disciplines.

The Team

The System Management team sits within the Software Platform group and helps build Graphcore products into large-scale AI solutions for our customers.

The team is responsible for developing the interfaces between hardware, AI software and frameworks, as well as providing interfaces for public and private cloud environments. This includes system management capabilities that abstract complex hardware administration and enable reliable deployment and operation at scale.

As one of the first teams to work with new hardware and software, we regularly solve complex system-level problems in environments where components and interfaces are still evolving. The role requires strong technical judgement, adaptability and an ability to work effectively across engineering teams.

Responsibilities and Duties
  • Convert agreed System Management direction into technical plans, engineering priorities and deliverable work for assigned areas.
  • Provide technical leadership for architecture and design decisions, building alignment across collaborating teams and documenting important technical trade‑offs.
  • Act as a technical authority for assigned areas of System Management, leading the delivery of large and complex engineering initiatives and coordinating technical plans, dependencies, risks and decisions.
  • Provide technical direction and mentoring to engineers working across system management, hardware lifecycle management, deployment automation and production operations.
  • Take responsibility for key technical outcomes across the full software lifecycle, including design, implementation, automated testing, integration, deployment, observability and production readiness.
  • Identify systemic reliability, scalability and operability issues and lead practical improvements across the platform.
  • Collaborate with Hardware, Firmware, Platform Software and Datacenter Operations teams to diagnose system-level issues and improve end-to-end product behaviour.
  • Improve engineering standards and working practices, including CI/CD, Infrastructure-as-Code, automated testing, release safety and learning from operational incidents.
  • Act as a senior technical escalation point for complex issues while creating reusable knowledge, tooling and automation that reduce future operational effort.
Candidate Profile
  • Bachelor’s degree or equivalent practical experience in a relevant subject.
  • Substantial experience designing, building and operating Linux-based infrastructure or distributed systems.
  • Demonstrated experience providing technical leadership for complex engineering initiatives involving multiple teams or stakeholder groups.
  • Experience influencing architecture and technical decisions across team boundaries without relying on formal authority.
  • Experience translating broad technical goals into scoped plans, milestones, technical decisions, risks and delivery priorities.
  • Strong experience developing RESTful APIs and programming in Go, with Bash and Python used for systems automation.
  • Deep practical experience with Kubernetes, container runtimes and operating production workloads.
  • Hands-on experience with Infrastructure-as-Code, source control and CI/CD technologies such as Terraform/OpenTofu, Ansible, GitLab, GitHub Actions and Git.
  • Experience with hardware-management interfaces such as Redfish, IPMI or equivalent management systems.
  • Strong Linux systems engineering, troubleshooting and operational debugging capability.
  • Demonstrated ability to develop other engineers through technical mentoring, design reviews and coaching.
  • Clear communication skills, with the ability to persuade, build alignment and bring stakeholders together around practical technical outcomes.
  • Experience using AI coding assistants effectively within professional engineering workflows.
  • Experience developing Kubernetes operators and custom resources.
  • Experience with High Performance Computing environments using SLURM, LSF or similar workload-management systems.
  • Experience with virtualisation technologies such as Open vSwitch, KVM and QEMU.
  • Experience with distributed object, block and file storage technologies such as Ceph.
  • Experience with monitoring and observability platforms such as Grafana, Prometheus, OpenSearch/Elasticsearch, Loki, Mimir or OpenTelemetry.
  • Experience configuring managed network switches using technologies such as EOS, SONiC or DNOS.
  • Experience supporting AI infrastructure or PyTorch workloads.

In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.

As set forth in Graphcore's Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Engineer
Principal Engineer

Cerebras • United States

On-site
USD 180,000 - 240,000
Medical coverage
Dental coverage
Vision coverage
+4
Staff Engineer
Staff Engineer

Graphcore • Town of Texas (WI), Northern (KY)

Hybrid
USD 150,000 - 230,000
Medical, dental, and vision coverage
FSAs/HSAs
Disability and life insurance
+4
Senior Systems Engineer
Senior Systems Engineer

Graphcore • Austin (TX)

On-site
USD 140,000 - 190,000
Medical insurance
Dental insurance
Vision insurance
+4
Principal Engineer
Principal Engineer

EngineersOfAI • Town of Texas (WI), Northern (KY)

Hybrid
USD 150,000 - 230,000
Principal Engineer, Power Electronics US - Milpitas
Principal Engineer, Power Electronics US - Milpitas

graphcore • Town of Texas (WI)

Hybrid
USD 110,000 - 150,000
Systems Engineer
Systems Engineer

Cerebras • Milpitas (CA)

On-site
USD 150,000 - 210,000
Medical, dental and vision coverage
Flexible Spending Accounts (FSAs)
Health Savings Accounts (HSAs)
+5
Principal Engineer, Power Electronics
Principal Engineer, Power Electronics

Graphcore • Milpitas (CA)

Hybrid
USD 120,000 - 150,000
Principal Electrical Engineer
Principal Electrical Engineer

Graphcore • Austin (TX)

On-site
USD 170,000 - 230,000
Medical coverage
401(k) retirement plan
Commuter benefits
+2
Principal Engineer, Power Electronics
Principal Engineer, Power Electronics

Cerebras • Milpitas (CA)

Hybrid
USD 130,000 - 180,000
Staff Hardware Engineer
Staff Hardware Engineer

Graphcore • Austin (TX)

On-site
USD 140,000 - 210,000
Medical coverage
Dental coverage
Vision coverage
+4