HPC Developer

Autonomai Recruitment

Chicago (IL)

On-site

USD 110,000 - 190,000

Full time

14 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Autonomai Recruitment seeks a software-focused HPC Data Center Developer to automate onboarding, provisioning, monitoring and lifecycle management of data-centre assets. You will design production-grade automation in Golang and work across servers, GPUs, switches, PDUs and environmental sensors.

The role demands strong Linux systems engineering, experience automating physical hardware, and familiarity with Kubernetes, Prometheus and Grafana for observability.

Qualifications

  • Strong Golang development experience in production
  • Proven automation of physical hardware or data-centre infra
  • Experience building automation for servers, GPUs, switches, PDUs and cooling
  • Strong Linux and systems engineering knowledge
  • Understanding of rack-scale infrastructure and data-centre operations
  • Experience with observability platforms (Prometheus, Grafana)
  • Familiarity with configuration management, infrastructure as code and CI/CD
  • Experience in HPC, AI/ML, GPU or high-density compute environments

Responsibilities

  • Design, develop and maintain production-grade automation in Golang
  • Automate full lifecycle of HPC data-centre hardware from discovery to decommissioning
  • Build workflows for servers, GPUs, switches, rack PDUs, CDUs and environmental sensors
  • Develop end-to-end processes from racked hardware to production readiness
  • Integrate with hardware management interfaces and vendor platforms
  • Build tooling for health monitoring, diagnostics and automated recovery
  • Integrate telemetry into centralised monitoring platforms
  • Support capacity planning across power, cooling, rack space and compute
  • Develop inventory, spares, change management and lifecycle tracking
  • Collaborate with HPC engineering, data-centre operations and network teams
  • Translate manual processes into reliable automation and maintain it
  • Support large-scale deployments and production incidents

Skills

Golang
Linux
Kubernetes
Prometheus
Grafana
Python
CI/CD

Tools

Redfish
IPMI
SNMP
Vendor APIs
Shell

Job description

We are looking for an HPC Data Center Developer to build the automation and tooling that supports large-scale high-performance computing infrastructure.

This is a development-focused role for an engineer who understands both software and data-centre hardware. You will automate the onboarding, provisioning, monitoring and lifecycle management of servers, network equipment, power systems, cooling infrastructure and other critical data-centre assets.

The ideal candidate will have strong, hands-on experience with Golang and a proven track record of automating physical hardware and infrastructure at scale.

Responsibilities

  • Design, develop and maintain production-grade automation in Golang.
  • Automate the full lifecycle of HPC data-centre hardware, from initial discovery and provisioning through to deployment, monitoring, maintenance and decommissioning.
  • Build workflows for servers, GPUs, network switches, rack PDUs, CDUs, environmental sensors and related infrastructure.
  • Develop end-to-end processes that take hardware from racked and cabled through configuration, validation and production readiness with minimal manual intervention.
  • Integrate with hardware management interfaces, APIs and protocols such as Redfish, IPMI, SNMP and vendor-specific platforms.
  • Build tooling for hardware health monitoring, diagnostics, alerting and automated recovery.
  • Integrate telemetry from compute, networking, power, cooling and environmental systems into centralised monitoring platforms.
  • Support capacity planning across power, cooling, rack space, networking and compute resources.
  • Develop operational tools for inventory management, spares, change management, troubleshooting and hardware lifecycle tracking.
  • Work closely with HPC engineering, data-centre operations, network, systems and infrastructure teams.
  • Translate manual operational processes and pain points into reliable, maintainable automation.
  • Own the reliability, documentation and ongoing improvement of the systems and tools you develop.
  • Support large-scale infrastructure deployments, maintenance activities and production incidents when required.

Required experience

  • Strong professional experience developing software in Golang.
  • Demonstrable experience automating physical hardware or data-centre infrastructure.
  • Experience building production automation for server, GPU, network, storage, power or cooling systems.
  • Strong Linux and systems engineering knowledge.
  • Understanding of server components, rack-scale infrastructure and data-centre operations.
  • Experience developing reliable, scalable and observable production systems.
  • Ability to work effectively with infrastructure, hardware and operations teams.
  • Strong troubleshooting, debugging and problem-solving skills.
  • Experience supporting HPC, AI/ML, GPU or high-density compute environments.
  • Experience with data-centre hardware such as GPUs, servers, switches, rack PDUs, CDUs and environmental monitoring systems.
  • Knowledge of Kubernetes, Slurm or other cluster-management and workload-orchestration technologies.
  • Experience with Python, Bash or another systems programming language.
  • Familiarity with configuration management, infrastructure as code and CI/CD.
  • Experience with Prometheus, Grafana or other observability platforms.
  • Exposure to power, cooling and data-centre capacity planning.
  • Experience in a high-availability, trading, cloud, hyperscale or other performance-critical environment.

What we are looking for

This role would suit a software engineer, infrastructure developer, systems engineer or data-centre automation engineer who enjoys working close to the hardware.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Developer
HPC Developer

Autonomai Recruitment • New York (NY)

On-site
USD 130,000 - 195,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
HPC Data Center Developer
HPC Data Center Developer

P2P • New York (NY)

On-site
USD 100,000 - 130,000
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
HPC Data Center Automation Engineer
HPC Data Center Automation Engineer

P2P • New York (NY)

On-site
USD 100,000 - 130,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

Cloudjobs • New York (NY)

On-site
USD 120,000 - 160,000
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 150,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Staff HPC Software Engineer
Staff HPC Software Engineer

San Diego Stealth Startup • San Diego (CA)

On-site
USD 140,000 - 210,000