We are looking for an HPC Data Center Developer to build the automation and tooling that supports large-scale high-performance computing infrastructure.
This is a development-focused role for an engineer who understands both software and data-centre hardware. You will automate the onboarding, provisioning, monitoring and lifecycle management of servers, network equipment, power systems, cooling infrastructure and other critical data-centre assets.
The ideal candidate will have strong, hands-on experience with Golang and a proven track record of automating physical hardware and infrastructure at scale.
Responsibilities
- Design, develop and maintain production-grade automation in Golang.
- Automate the full lifecycle of HPC data-centre hardware, from initial discovery and provisioning through to deployment, monitoring, maintenance and decommissioning.
- Build workflows for servers, GPUs, network switches, rack PDUs, CDUs, environmental sensors and related infrastructure.
- Develop end-to-end processes that take hardware from racked and cabled through configuration, validation and production readiness with minimal manual intervention.
- Integrate with hardware management interfaces, APIs and protocols such as Redfish, IPMI, SNMP and vendor-specific platforms.
- Build tooling for hardware health monitoring, diagnostics, alerting and automated recovery.
- Integrate telemetry from compute, networking, power, cooling and environmental systems into centralised monitoring platforms.
- Support capacity planning across power, cooling, rack space, networking and compute resources.
- Develop operational tools for inventory management, spares, change management, troubleshooting and hardware lifecycle tracking.
- Work closely with HPC engineering, data-centre operations, network, systems and infrastructure teams.
- Translate manual operational processes and pain points into reliable, maintainable automation.
- Own the reliability, documentation and ongoing improvement of the systems and tools you develop.
- Support large-scale infrastructure deployments, maintenance activities and production incidents when required.
Required experience
- Strong professional experience developing software in Golang.
- Demonstrable experience automating physical hardware or data-centre infrastructure.
- Experience building production automation for server, GPU, network, storage, power or cooling systems.
- Strong Linux and systems engineering knowledge.
- Understanding of server components, rack-scale infrastructure and data-centre operations.
- Experience developing reliable, scalable and observable production systems.
- Ability to work effectively with infrastructure, hardware and operations teams.
- Strong troubleshooting, debugging and problem-solving skills.
- Experience supporting HPC, AI/ML, GPU or high-density compute environments.
- Experience with data-centre hardware such as GPUs, servers, switches, rack PDUs, CDUs and environmental monitoring systems.
- Knowledge of Kubernetes, Slurm or other cluster-management and workload-orchestration technologies.
- Experience with Python, Bash or another systems programming language.
- Familiarity with configuration management, infrastructure as code and CI/CD.
- Experience with Prometheus, Grafana or other observability platforms.
- Exposure to power, cooling and data-centre capacity planning.
- Experience in a high-availability, trading, cloud, hyperscale or other performance-critical environment.
What we are looking for
This role would suit a software engineer, infrastructure developer, systems engineer or data-centre automation engineer who enjoys working close to the hardware.