Sr HPC Hardware Engineer

Career Techniques

Dallas (TX)

Hybrid

USD 120,000 - 180,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Career Techniques is seeking an experienced Infrastructure Engineer to design, deploy, and manage a large-scale HPC/AI compute fleet. You will own the firmware and BIOS lifecycle, lead hardware troubleshooting, and drive automation across GPU and CPU nodes.

You will collaborate with vendors, implement security hardening, and mentor junior engineers to elevate team practices in a fast-paced environment.

Qualifications

  • Bachelor’s degree in Electrical Engineering, Computer Engineering, or related field, or equivalent hands‑on experience.
  • 8+ years of experience managing large-scale HPC or AI compute infrastructure in a production environment.
  • Deep knowledge of server hardware architecture including processors, memory, storage, networking, power systems, and thermal management.
  • Hands‑on experience with bare-metal provisioning and hardware automation tools (Ansible, Puppet, Chef).
  • Proficiency with Redfish API and BMC/IPMI tooling (iDRAC, iLO) for remote hardware management.
  • Ability to troubleshoot complex hardware issues across GPU/CPU nodes with NVSMI and diagnostics.
  • Experience with hardware monitoring, performance tuning, and capacity planning at scale.
  • Scripting in Python, Bash, or PowerShell for infra automation.
  • Experience with OpenStack (Ironic) or similar provisioning platforms is strongly preferred.
  • Strong cross-functional communication and prior technical leadership.

Responsibilities

  • Design, configure, and manage a high-performance compute fleet of GPU and CPU nodes.
  • Own firmware and BIOS lifecycle from baselining to rollout and maintenance.
  • Lead troubleshooting of CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, and BMCs.
  • Automate health checks, onboarding workflows, and remediation to speed deployment.
  • Validate next-gen AI platforms for stability, performance, and production fitness.
  • Collaborate with vendors on firmware and hardware issues with clear repro steps.
  • Perform hardware performance analysis and capacity planning for scale-out.
  • Define security hardening for hardware infrastructure across the fleet.
  • Use IaC and scripting to drive repeatable infrastructure management.
  • Mentor junior engineers and drive team-wide best practices.

Skills

HPC infrastructure
Hardware troubleshooting
Firmware lifecycle
Automation tools
OpenStack Ironic
Redfish / iDRAC / iLO
Programming / scripting
GPU/CPU hardware
Infrastructure as Code
Mentoring / leadership

Education

Bachelor’s degree in Electrical or Computer Engineering or related field

Tools

Ansible
Puppet
Chef
Redfish
OpenStack Ironic
Python
PowerShell
Bash

Job description

RESPONSIBILITIES
  • Design, configure, and manage a high-performance compute fleet comprising large-scale GPU (NVIDIA V100/A100/H200/GB200) and CPU nodes across the firm's infrastructure.
  • Own the full firmware and BIOS lifecycle across the HPC/AI fleet — from establishing baselines and validation through rollout, compliance, and ongoing maintenance.
  • Lead troubleshooting of hardware components including CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, and BMCs; drive component replacement and configuration remediation.
  • Automate health checks, onboarding workflows, and recurring hardware issue remediation to accelerate safe deployment and reduce recovery time.
  • Validate and operationalize next-generation AI platforms (e.g., NVL72 / Grace Blackwell) from day one, ensuring stability, performance readiness, and production fitness.
  • Collaborate with vendors on firmware and hardware issues, providing clear reproduction cases, diagnostic logs, and business impact to drive timely resolution.
  • Perform hardware performance analysis, tuning, and capacity planning to ensure reliable scale-out of the compute environment.
  • Define and implement security hardening best practices for hardware infrastructure, maintaining platform integrity across the fleet.
  • Leverage Infrastructure as Code (IaC) methodologies and scripting to drive efficient, repeatable, and scalable infrastructure management.
  • Mentor junior engineers, act as a subject matter expert for infrastructure-related escalations, and champion a culture of continuous improvement across the team.
REQUIREMENTS
  • Bachelor’s degree in Electrical Engineering, Computer Engineering, or a related field, or equivalent hands‑on experience.
  • 8+ years of experience managing large-scale HPC or AI compute infrastructure in a production environment.
  • Deep knowledge of server hardware architecture, including processors, memory, storage, networking, power systems, and thermal management.
  • Hands‑on experience with bare-metal provisioning, firmware and BIOS lifecycle management, and hardware automation tools such as Ansible, Puppet, or Chef.
  • Proficiency with Redfish API and BMC/IPMI tooling (iDRAC, iLO) for remote hardware management and diagnostics.
  • Demonstrated ability to troubleshoot and resolve complex hardware issues across GPU and CPU nodes, including NVIDIA-SMI and GPU diagnostics.
  • Experience with hardware monitoring platforms, performance tuning, and capacity planning at scale.
  • Familiarity with Linux-based environments and scripting proficiency in Python, Bash, or PowerShell for infrastructure automation.
  • Experience with OpenStack (particularly Ironic) or equivalent cloud/bare-metal provisioning platforms is strongly preferred.
  • Strong cross-functional communication skills and proven ability to collaborate effectively with software, networking, and vendor teams.
  • Prior technical leadership experience, including mentoring engineers and driving team-wide best practices.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 150,000 - 200,000
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior HPC Hardware Architect & Infra Automation Lead
Senior HPC Hardware Architect & Infra Automation Lead

Career Techniques • Dallas (TX)

Hybrid
USD 120,000 - 180,000
High Performance Computing Hardware / Systems Engineer
High Performance Computing Hardware / Systems Engineer

KLA • Ann Arbor Charter Township (MI)

On-site
USD 120,000 - 180,000
Infrastructure Hardware Specialist
Infrastructure Hardware Specialist

Blue Signal Search • Austin (TX)

On-site
USD 90,000 - 130,000
Competitive base salary plus equity
Health/dental/vision coverage
Retirement benefits
+1
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Infrastructure Hardware Specialist
Infrastructure Hardware Specialist

Blue Signal Search • Phoenix (AZ)

On-site
USD 90,000 - 130,000
Equity
Health/Dental/Vision
Retirement benefits
+2
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000