HPC Developer

Autonomai Recruitment

New York (NY)

On-site

USD 130,000 - 195,000

Full time

13 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Autonomai Recruitment seeks an HPC Data Center Developer to build automation and tooling for large-scale HPC infrastructure. You will automate onboarding, provisioning, monitoring and lifecycle management of servers, network equipment, power systems, cooling infrastructure and other data-centre assets.

The ideal candidate will have strong Golang development experience and a proven ability to automate physical hardware and infrastructure at scale, collaborating with HPC engineering and operations

Qualifications

  • Strong professional experience developing software in Golang.
  • Demonstrable experience automating physical hardware or data-centre infrastructure.
  • Experience building production automation for server, GPU, network, storage, power or cooling systems.
  • Strong Linux and systems engineering knowledge.
  • Understanding of server components, rack-scale infrastructure and data-centre operations.
  • Experience developing reliable, scalable and observable production systems.
  • Ability to work effectively with infrastructure, hardware and operations teams.
  • Strong troubleshooting, debugging and problem-solving skills.
  • Experience supporting HPC, AI/ML, GPU or high-density compute environments.
  • Experience with data-centre hardware such as GPUs, servers, switches, rack PDUs, CDUs and environmental monitoring systems.
  • Knowledge of Kubernetes, Slurm or other cluster-management and workload-orchestration technologies.
  • Experience with Python, Bash or another systems programming language.
  • Familiarity with configuration management, infrastructure as code and CI/CD.
  • Experience with Prometheus, Grafana or other observability platforms.
  • Exposure to power, cooling and data-centre capacity planning.
  • Experience in a high-availability, trading, cloud, hyperscale or other performance-critical environment.

Responsibilities

  • Design, develop and maintain production-grade automation in Golang.
  • Automate the full lifecycle of HPC data-centre hardware, from initial discovery and provisioning through to deployment, monitoring, maintenance and decommissioning.
  • Build workflows for servers, GPUs, network switches, rack PDUs, CDUs, environmental sensors and related infrastructure.
  • Develop end-to-end processes that take hardware from racked and cabled through configuration, validation and production readiness with minimal manual intervention.
  • Integrate with hardware management interfaces, APIs and protocols such as Redfish, IPMI, SNMP and vendor-specific platforms.
  • Build tooling for hardware health monitoring, diagnostics, alerting and automated recovery.
  • Integrate telemetry from compute, networking, power, cooling and environmental systems into centralised monitoring platforms.
  • Support capacity planning across power, cooling, rack space, networking and compute resources.
  • Develop operational tools for inventory management, spares, change management, troubleshooting and hardware lifecycle tracking.
  • Work closely with HPC engineering, data-centre operations, network, systems and infrastructure teams.
  • Translate manual operational processes and pain points into reliable, maintainable automation.
  • Own the reliability, documentation and ongoing improvement of the systems and tools you develop.
  • Support large-scale infrastructure deployments, maintenance activities and production incidents when required.

Skills

Golang
Linux
Observability
Infrastructure automation
Kubernetes
Python

Tools

Redfish
IPMI
Prometheus
Grafana
SNMP
Ansible
Terraform

Job description

We are looking for an HPC Data Center Developer to build the automation and tooling that supports large-scale high-performance computing infrastructure.

This is a development-focused role for an engineer who understands both software and data-centre hardware. You will automate the onboarding, provisioning, monitoring and lifecycle management of servers, network equipment, power systems, cooling infrastructure and other critical data-centre assets.

The ideal candidate will have strong, hands-on experience with Golang and a proven track record of automating physical hardware and infrastructure at scale.

Responsibilities

  • Design, develop and maintain production-grade automation in Golang.
  • Automate the full lifecycle of HPC data-centre hardware, from initial discovery and provisioning through to deployment, monitoring, maintenance and decommissioning.
  • Build workflows for servers, GPUs, network switches, rack PDUs, CDUs, environmental sensors and related infrastructure.
  • Develop end-to-end processes that take hardware from racked and cabled through configuration, validation and production readiness with minimal manual intervention.
  • Integrate with hardware management interfaces, APIs and protocols such as Redfish, IPMI, SNMP and vendor-specific platforms.
  • Build tooling for hardware health monitoring, diagnostics, alerting and automated recovery.
  • Integrate telemetry from compute, networking, power, cooling and environmental systems into centralised monitoring platforms.
  • Support capacity planning across power, cooling, rack space, networking and compute resources.
  • Develop operational tools for inventory management, spares, change management, troubleshooting and hardware lifecycle tracking.
  • Work closely with HPC engineering, data-centre operations, network, systems and infrastructure teams.
  • Translate manual operational processes and pain points into reliable, maintainable automation.
  • Own the reliability, documentation and ongoing improvement of the systems and tools you develop.
  • Support large-scale infrastructure deployments, maintenance activities and production incidents when required.

Required experience

  • Strong professional experience developing software in Golang.
  • Demonstrable experience automating physical hardware or data-centre infrastructure.
  • Experience building production automation for server, GPU, network, storage, power or cooling systems.
  • Strong Linux and systems engineering knowledge.
  • Understanding of server components, rack-scale infrastructure and data-centre operations.
  • Experience developing reliable, scalable and observable production systems.
  • Ability to work effectively with infrastructure, hardware and operations teams.
  • Strong troubleshooting, debugging and problem-solving skills.
  • Experience supporting HPC, AI/ML, GPU or high-density compute environments.
  • Experience with data-centre hardware such as GPUs, servers, switches, rack PDUs, CDUs and environmental monitoring systems.
  • Knowledge of Kubernetes, Slurm or other cluster-management and workload-orchestration technologies.
  • Experience with Python, Bash or another systems programming language.
  • Familiarity with configuration management, infrastructure as code and CI/CD.
  • Experience with Prometheus, Grafana or other observability platforms.
  • Exposure to power, cooling and data-centre capacity planning.
  • Experience in a high-availability, trading, cloud, hyperscale or other performance-critical environment.

What we are looking for

This role would suit a software engineer, infrastructure developer, systems engineer or data-centre automation engineer who enjoys working close to the hardware.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Developer
HPC Developer

Autonomai Recruitment • Chicago (IL)

On-site
USD 110,000 - 190,000
HPC Data Center Developer
HPC Data Center Developer

P2P • New York (NY)

On-site
USD 100,000 - 130,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
HPC Data Center Automation Engineer
HPC Data Center Automation Engineer

P2P • New York (NY)

On-site
USD 100,000 - 130,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

Cloudjobs • New York (NY)

On-site
USD 120,000 - 160,000
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 150,000
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 140,000 - 190,000