Senior HPC Hardware Architect & Infra Automation Lead

Career Techniques

Dallas (TX)

Hybrid

USD 120,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Career Techniques is seeking an experienced Infrastructure Engineer to design, deploy, and manage a large-scale HPC/AI compute fleet. You will own the firmware and BIOS lifecycle, lead hardware troubleshooting, and drive automation across GPU and CPU nodes.

You will collaborate with vendors, implement security hardening, and mentor junior engineers to elevate team practices in a fast-paced environment.

Qualifications

  • Bachelor’s degree in Electrical Engineering, Computer Engineering, or related field, or equivalent hands‑on experience.
  • 8+ years of experience managing large-scale HPC or AI compute infrastructure in a production environment.
  • Deep knowledge of server hardware architecture including processors, memory, storage, networking, power systems, and thermal management.
  • Hands‑on experience with bare-metal provisioning and hardware automation tools (Ansible, Puppet, Chef).
  • Proficiency with Redfish API and BMC/IPMI tooling (iDRAC, iLO) for remote hardware management.
  • Ability to troubleshoot complex hardware issues across GPU/CPU nodes with NVSMI and diagnostics.
  • Experience with hardware monitoring, performance tuning, and capacity planning at scale.
  • Scripting in Python, Bash, or PowerShell for infra automation.
  • Experience with OpenStack (Ironic) or similar provisioning platforms is strongly preferred.
  • Strong cross-functional communication and prior technical leadership.

Responsibilities

  • Design, configure, and manage a high-performance compute fleet of GPU and CPU nodes.
  • Own firmware and BIOS lifecycle from baselining to rollout and maintenance.
  • Lead troubleshooting of CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, and BMCs.
  • Automate health checks, onboarding workflows, and remediation to speed deployment.
  • Validate next-gen AI platforms for stability, performance, and production fitness.
  • Collaborate with vendors on firmware and hardware issues with clear repro steps.
  • Perform hardware performance analysis and capacity planning for scale-out.
  • Define security hardening for hardware infrastructure across the fleet.
  • Use IaC and scripting to drive repeatable infrastructure management.
  • Mentor junior engineers and drive team-wide best practices.

Skills

HPC infrastructure
Hardware troubleshooting
Firmware lifecycle
Automation tools
OpenStack Ironic
Redfish / iDRAC / iLO
Programming / scripting
GPU/CPU hardware
Infrastructure as Code
Mentoring / leadership

Education

Bachelor’s degree in Electrical or Computer Engineering or related field

Tools

Ansible
Puppet
Chef
Redfish
OpenStack Ironic
Python
PowerShell
Bash

Job description

Career Techniques is seeking an experienced Infrastructure Engineer to design, deploy, and manage a large-scale HPC/AI compute fleet. You will own the firmware and BIOS lifecycle, lead hardware troubleshooting, and drive automation across GPU and CPU nodes.

You will collaborate with vendors, implement security hardening, and mentor junior engineers to elevate team practices in a fast-paced environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Hardware Architect – AI Compute & Scale
Senior HPC Hardware Architect – AI Compute & Scale

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 140,000 - 190,000
Lunch stipend
Company-paid medical benefits
Dental and Vision for employees and 가족
+6
Sr HPC Hardware Engineer
Sr HPC Hardware Engineer

Career Techniques • Dallas (TX)

Hybrid
USD 120,000 - 180,000
Senior HPC Deployment & Validation Lead Equity
Senior HPC Deployment & Validation Lead Equity

NVIDIA • South Carolina

On-site
USD 230,000 - 360,000
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2
Senior HPC Hardware Engineer: Lead the AI Compute Fleet
Senior HPC Hardware Engineer: Lead the AI Compute Fleet

NorthMark Strategies • Town of Texas (WI)

On-site
USD 140,000 - 190,000
Lunch stipend
Medical benefits (employer-paid)
Parental leave 16 weeks
+2
Senior HPC Hardware Architect - AI & GPU Compute
Senior HPC Hardware Architect - AI & GPU Compute

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 150,000 - 200,000
Senior HPC Architect: Large-Scale GPU AI Infra (Equity)
Senior HPC Architect: Large-Scale GPU AI Infra (Equity)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior HPC Architect: At-Scale GPU Deployment & Automation
Senior HPC Architect: At-Scale GPU Deployment & Automation

NVIDIA • New Mexico

On-site
USD 184,000 - 288,000