Staff Engineer at Graphcore. About the role Graphcore is a SoftBank Group company focused on building the hardware and software infrastructure required for advanced artificial intelligence. This role sits within the System Management team, where you will develop interfaces that bridge the gap between physical hardware and customer AI workloads. You will lead the engineering of rack management solutions while ensuring high-performance operation across internal and external datacenters.
Location: Austin, Texas, United States; Milpitas, California, United States
Engagement: Full-time
Team: System Management (Software Platform group)
What you'll do
- Manage the entire software development life cycle for rack management systems, from initial design through to production deployment and automated testing.
- Collaborate across internal departments to resolve infrastructure issues and maintain system reliability.
- Utilize Infrastructure-as-code and Continuous Deployment to configure and validate new AI hardware in production environments.
- Partner with Datacenter Operations to monitor fleet performance and execute corrective maintenance.
- Lead technical planning and scoping within an Agile and Scrum framework, identifying project risks and constraints.
Requirements
- Bachelor degree in a relevant field or equivalent professional experience.
- Proficiency in Go programming and Linux systems engineering, including Bash and Python scripting.
- Experience developing RESTful APIs.
- Background in managing production Kubernetes clusters and containerized workloads using Docker or Podman.
- Hands-on experience with Infrastructure-as-code tools such as Terraform, OpenTofu, or Ansible, along with CI/CD platforms like GitLab or GitHub Actions.
- Knowledge of Redfish for datacenter hardware provisioning, telemetry, and control.
- Ability to manage project work plans, priorities, and technical documentation.
Nice to have
- Familiarity with Kubernetes operator development and custom resources.
- Experience with High Performance Computing environments, specifically SLURM.
- Knowledge of virtualized deployments using KVM, QEMU, or Open vSwitch.
- Experience with distributed storage systems like Ceph.
- Proficiency in monitoring and observability stacks including Prometheus, Grafana, OpenSearch, Loki, Mimir, OpenTelemetry, Fluentd, or Kafka.
- Experience configuring managed switches such as SONiC, EOS, or DNOS.
- Familiarity with PyTorch for AI workloads.
- Experience using AI coding assistants.
- Background in end-to-end pipeline automation for build, test, and deployment.
Skills & tools
- Go, Python, Bash
- Kubernetes, Docker, Podman
- Terraform, OpenTofu, Ansible
- GitLab, GitHub Actions
- Redfish
- Linux Administration
- Agile/Scrum
Benefits include medical, dental, and vision coverage, 401(k) retirement plans, disability and life insurance, commuter benefits, and wellness services.
- Flexible spending accounts and health savings accounts are available.
The company provides an equal opportunity hiring process and supports reasonable adjustments for candidates during the interview stage.