HPC Data Center Developer

Autonomai Recruitment

Chicago (IL)

On-site

USD 140,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Autonomai Recruitment in Chicago/New York is partnering with a global trading firm to hire an HPC Data Center Developer to build the software and automation behind its large-scale compute infrastructure. This is a development-heavy role with ownership of production systems spanning hardware provisioning, infrastructure monitoring, capacity planning and failure simulation.

You'll write substantial code and work closely with HPC engineering and operations teams to solve problems across the

Qualifications

  • 5+ years in infrastructure software development, production engineering, automation or SRE.
  • Track record of shipping reliable production tooling.
  • Strong Go and Python development experience.

Responsibilities

  • Build end-to-end hardware provisioning workflows across servers, switches, power and cooling.
  • Develop tools to model power and cooling capacity, forecast demand and identify constraints.
  • Create failure simulation tooling to assess outages and validate resilience.
  • Integrate hardware telemetry into central monitoring platforms with dashboards and alerts.
  • Automate hardware lifecycle tracking, inventory management, diagnostics and workflows.
  • Own the reliability, maintenance and ongoing development of the systems you build.
  • Use AI development tools daily to support coding, debugging, analysis and documentation.

Skills

Go
Python
Linux
Redfish
IPMI/BMC
SNMP
Networking
Grafana
Prometheus
InfluxDB
CI/CD
APIs
Automation

Tools

Grafana
Prometheus
InfluxDB
CI/CD tooling

Job description

Chicago or New York |

We're partnering with a global trading firm to hire an HPC Data Center Developer to build the software and automation behind its large-scale compute infrastructure.

This is a development-heavy role with ownership of production systems spanning hardware provisioning, infrastructure monitoring, capacity planning and failure simulation. You'll write substantial amounts of code and work closely with HPC engineering and operations teams to solve problems across the physical infrastructure powering research and trading.

The role
  • Build end-to-end hardware provisioning workflows covering discovery, configuration, validation and deployment across servers, switches, power and cooling equipment.
  • Develop tools to model power and cooling capacity, forecast demand and identify infrastructure constraints.
  • Create failure simulation tooling to assess the impact of outages and validate infrastructure resilience.
  • Integrate hardware telemetry and external data feeds into central monitoring platforms, building dashboards, exporters and alerting.
  • Automate hardware lifecycle tracking, inventory management, diagnostics and operational workflows.
  • Own the reliability, maintenance and ongoing development of the systems you build.
  • Use AI development tools daily to support coding, debugging, analysis and documentation.
What we’re looking for
  • 5+ years in infrastructure software development, production engineering, automation or SRE, ideally within HPC or large-scale data center environments.
  • A track record of shipping reliable, maintainable production tooling.
  • Strong programming skills in Go and Python, or comparable infrastructure development experience.
  • Deep Linux knowledge, including networking, storage, system administration and troubleshooting.
  • Experience automating hardware provisioning and working with interfaces such as Redfish, IPMI/BMC, SNMP or vendor APIs.
  • An understanding of data center power, cooling and physical infrastructure, including air and liquid cooling.
  • Experience with observability platforms such as Grafana, Prometheus or InfluxDB, alongside configuration management and infrastructure-as-code tools.
  • Strong networking fundamentals and confidence integrating APIs, databases and infrastructure metrics.

Experience with ClickHouse, MySQL, Arista/Cisco networking and CI/CD workflows would also be valuable.

You'll suit this role if you enjoy understanding how infrastructure works from the hardware through to the software, investigating root causes and turning operational challenges into well-engineered systems. Participation in coordinated evening and weekend maintenance is required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Data Center Developer
HPC Data Center Developer

Autonomai Recruitment • New York (NY)

On-site
USD 140,000 - 210,000
HPC Data Center Engineer (GoLang) - Quant Trading
HPC Data Center Engineer (GoLang) - Quant Trading

Radley James • New York (NY)

On-site
USD 120,000 - 180,000
HPC & Compute Engineering Lead
HPC & Compute Engineering Lead

Autonomai Recruitment • Chicago (IL)

On-site
USD 180,000 - 250,000
HPC Data Center Automation Engineer
HPC Data Center Automation Engineer

Autonomai Recruitment • Chicago (IL)

On-site
USD 140,000 - 190,000
HPC Data Center Production Engineer (Automation) - Banking & Finance
HPC Data Center Production Engineer (Automation) - Banking & Finance

Hamilton Barnes • Illinois

On-site
USD 175,000 - 235,000
Premium health, dental, and vision
Performance-based bonus
Meals and high-end office amenities
+2
HPC Developer
HPC Developer

Autonomai Recruitment • New York (NY)

On-site
USD 130,000 - 195,000
HPC Developer
HPC Developer

Autonomai Recruitment • Chicago (IL)

On-site
USD 110,000 - 190,000
Data Center Engineer
Data Center Engineer

Acquire Me • Illinois

On-site
USD 120,000 - 190,000
HPC Data Center Automation Engineer – On-site
HPC Data Center Automation Engineer – On-site

Jump Trading, LLC. • New York (NY), Chicago (IL)

On-site
USD 150,000 - 200,000
Discretionary bonus eligibility
Medical, dental, and vision insurance
HSA, FSA, and Dependent Care options
+6
HPC Data Center Automation Engineer (Go/Python)
HPC Data Center Automation Engineer (Go/Python)

Radley James • New York (NY)

On-site
USD 120,000 - 180,000