HPC Data Center Production Engineer
Chicago, IL or New York, NY – On-site 5 days/week.
Jump Trading Group is a global organization of engineers who architect, build, and maintain world‑class trading infrastructure. We empower exceptional talents in Mathematics, Physics, and Computer Science to push scientific boundaries and apply cutting‑edge research to global financial markets. Our culture demands fearlessness, creativity, intellectual honesty, and relentless competitiveness.
We are looking for an HPC Data Center Production Engineer to build and own the automation and tooling that powers Jump’s HPC data‑center operations. This development‑heavy role focuses on automating the onboarding and lifecycle management of data‑center hardware—servers, switches, rack PDUs, CDUs, and environmental sensors—and building tools for capacity planning, outage simulation, monitoring, and metrics integration.
Responsibilities
- Hardware Onboarding Automation – Design, develop, and maintain automation to onboard new hardware devices—including servers, network switches, rack PDUs, CDUs, and environmental sensors—taking them from racked and cabled to discovery, configuration, validation, and production‑ready state with minimal manual intervention.
- Data Center Tooling Development – Build tools for power and cooling capacity planning, outage simulation, and day‑to‑day operational support such as hardware lifecycle tracking, inventory management, change management, and diagnostics.
- Monitoring & Metrics Integration – Pull telemetry from all infrastructure components into centralized observability platforms, integrate provider metrics feeds, and implement the monitoring and alerting strategy.
- Cross‑Team Collaboration – Work closely with HPC Planning, Engineering, and Operations leads to translate tooling and monitoring needs into production‑ready systems, and partner with HPC Engineering on integration points with compute, storage, and network provisioning.
- Systems Maintenance & Reliability – Own reliability and lifecycle of all developed systems, monitor for failures, respond to incidents, iterate based on feedback, maintain comprehensive documentation, and participate in large maintenance operations including evenings and weekends.
- AI‑Driven Development – Use AI tools daily for code, debugging, documentation, and accelerate development velocity. Identify opportunities to apply AI to data‑center operations such as anomaly detection and predictive planning.
- Perform additional duties as assigned or needed.
Qualifications
- 5+ years of professional experience in production engineering, infrastructure automation, or site reliability engineering, preferably in HPC or large‑scale data‑center environments.
- Proven track record of building and shipping production automation and tooling that is maintained and reliable.
- Experience automating hardware provisioning and lifecycle management for servers, network devices, and power/cooling infrastructure.
- Strong understanding of data‑center infrastructure: power distribution, cooling systems (air and liquid), environmental monitoring, and structured cabling.
- Experience integrating with hardware management interfaces (IPMI, BMC, Redfish, SNMP, vendor APIs) for discovery, configuration, and telemetry collection.
- Result‑driven, high‑energy professional able to work under pressure with tight deadlines.
- Excellent written and verbal communication skills.
- Reliable and predictable availability, including ability to work evenings and weekends as required.
- Bachelor’s degree preferred.
Technical Skills
- High proficiency in Golang and at least one additional language (e.g., Python).
- Strong Linux systems knowledge, including system administration, networking, storage, process management, log analysis, and troubleshooting at the OS level.
- Experience with Grafana dashboards, Prometheus, InfluxDB, or similar observability platforms and building custom integrations or exporters.
- Experience with configuration‑management and infrastructure‑as‑code tools such as SaltStack, Ansible, Terraform.
- Solid networking knowledge: L2/L3 protocols, VLANs, BGP, SNMP, and switch/router configuration (Arista, Cisco).
- Experience consuming vendor APIs, normalizing heterogeneous data sources, building data pipelines for metrics and reporting.
- Experience with ClickHouse and MySQL, writing queries, designing schemas, and building tooling that reads from and writes to these databases.
- Experience with GitHub for version control, code review, CI/CD workflows, and collaborative development.
- Demonstrated heavy use of AI tools such as LLM‑based coding assistants and AI‑driven analytics in a professional setting.
- Compulsion to perform root‑cause analysis.
- Extremely high personal standards for work quality.
Benefits
- Discretionary bonus eligibility
- Medical, dental, and vision insurance
- HSA, FSA, and Dependent Care options
- Employer‑paid Group Term Life and AD&D insurance
- Voluntary Life & AD&D insurance
- Paid vacation plus paid holidays
- Retirement plan with employer match
- Paid parental leave
- Wellness programs
Salary
Annual Base Salary Range: $150,000 – $200,000 USD