Senior HPC Hardware Architect for AI Compute

NorthMark Compute & Cloud

Dallas (TX)

On-site

USD 140,000 - 220,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Lunch stipend
Medical benefits - employer-paid
Dental & Vision benefits
Paid parental leave
Life insurance
Disability insurance
401(k) company match
Flexible benefits
Paid time off

Job summary

NorthMark Compute & Cloud (NMC²) is seeking a Senior HPC Hardware Engineer in Dallas to own the full hardware lifecycle for a large GPU/CPU fleet. This hands-on role emphasizes development, validation, and production readiness across NVIDIA platforms and next-gen AI hardware.

You will lead troubleshooting, automation, capacity planning, and security hardening while mentoring engineers and collaborating with software, networking, and vendor teams.

Qualifications

  • Bachelor’s degree in Electrical Engineering, Computer Engineering, or related field, or equivalent hands-on experience.
  • 5+ years of experience managing large-scale HPC or AI compute infrastructure in a production environment.
  • Deep knowledge of server hardware architecture, including processors, memory, storage, networking, power systems, and thermal management.
  • Hands-on experience with bare-metal provisioning, firmware and BIOS lifecycle management, and hardware automation tools such as Ansible, Puppet, or Chef.
  • Proficiency with Redfish API and BMC/IPMI tooling (iDRAC, iLO) for remote hardware management and diagnostics.
  • Demonstrated ability to troubleshoot and resolve complex hardware issues across GPU and CPU nodes, including NVIDIA-SMI and GPU diagnostics.
  • Experience with hardware monitoring platforms, performance tuning, and capacity planning at scale.
  • Familiarity with Linux-based environments and scripting proficiency in Python, Bash, or PowerShell for infrastructure automation.
  • Experience with OpenStack (particularly Ironic) or equivalent cloud/bare-metal provisioning platforms is strongly preferred.
  • Strong cross-functional communication skills and proven ability to collaborate effectively with software, networking, and vendor teams.
  • Prior technical leadership experience, including mentoring engineers and driving team-wide best practices.

Responsibilities

  • Design, configure, and manage a high-performance compute fleet comprising large-scale GPU (NVIDIA V100/A100/H200/GB200) and CPU nodes across NMC²’s Dallas infrastructure.
  • Own the full firmware and BIOS lifecycle across the HPC/AI fleet — from establishing baselines and validation through rollout, compliance, and ongoing maintenance.
  • Lead troubleshooting of hardware components including CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, and BMCs; drive component replacement and configuration remediation.
  • Automate health checks, onboarding workflows, and recurring hardware issue remediation to accelerate safe deployment and reduce recovery time.
  • Validate and operationalize next-generation AI platforms (e.g., NVL72 / Grace Blackwell) from day one, ensuring stability, performance readiness, and production fitness.
  • Collaborate with vendors on firmware and hardware issues, providing clear reproduction cases, diagnostic logs, and business impact to drive timely resolution.
  • Perform hardware performance analysis, tuning, and capacity planning to ensure reliable scale-out of the compute environment.
  • Define and implement security hardening best practices for hardware infrastructure, maintaining platform integrity across the fleet.
  • Leverage Infrastructure as Code (IaC) methodologies and scripting to drive efficient, repeatable, and scalable infrastructure management.
  • Mentor junior engineers, act as a subject matter expert for infrastructure-related escalations, and champion a culture of continuous improvement across the team.

Skills

HPC hardware knowledge
Troubleshooting hardware
Cross-functional collaboration
Leadership & mentoring
Scripting (Python/Bash/PowerShell)

Education

Bachelor’s degree in Electrical/Computer Engineering or equivalent

Tools

Redfish/iDRAC/iLO
Ansible/Puppet/Chef
OpenStack Ironic

Job description

NorthMark Compute & Cloud (NMC²) is seeking a Senior HPC Hardware Engineer in Dallas to own the full hardware lifecycle for a large GPU/CPU fleet. This hands-on role emphasizes development, validation, and production readiness across NVIDIA platforms and next-gen AI hardware.

You will lead troubleshooting, automation, capacity planning, and security hardening while mentoring engineers and collaborating with software, networking, and vendor teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Hardware Engineer & Automation Lead
Senior HPC Hardware Engineer & Automation Lead

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior HPC Hardware Engineer: Lead the AI Compute Fleet
Senior HPC Hardware Engineer: Lead the AI Compute Fleet

NorthMark Strategies • Town of Texas (WI)

On-site
USD 140,000 - 190,000
Lunch stipend
Medical benefits (employer-paid)
Parental leave 16 weeks
+2
Senior HPC Hardware Architect – AI Compute & Scale
Senior HPC Hardware Architect – AI Compute & Scale

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 140,000 - 190,000
Lunch stipend
Company-paid medical benefits
Dental and Vision for employees and 가족
+6
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior HPC Architect: Large-Scale GPU AI Infra (Equity)
Senior HPC Architect: Large-Scale GPU AI Infra (Equity)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
HPC Systems Engineer — Cross‑Domain Platform Architect
HPC Systems Engineer — Cross‑Domain Platform Architect

NorthMark Compute and Cloud LLC • Dallas (TX)

On-site
USD 140,000 - 210,000
Lunch stipend
Medical, dental and vision coverage
Parental leave 16 weeks
+6
HPC Systems Engineer - Cross-Domain Platform Architect
HPC Systems Engineer - Cross-Domain Platform Architect

northmark • Dallas (TX)

On-site
USD 120,000 - 180,000
Senior HPC Architect: At-Scale GPU Deployment & Automation
Senior HPC Architect: At-Scale GPU Deployment & Automation

NVIDIA • New Mexico

On-site
USD 184,000 - 288,000
Senior HPC Hardware Architect & Infra Automation Lead
Senior HPC Hardware Architect & Infra Automation Lead

Career Techniques • Dallas (TX)

Hybrid
USD 120,000 - 180,000
Senior HPC Architect - GPU Compute, Equity Eligible
Senior HPC Architect - GPU Compute, Equity Eligible

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Inclusive work environment
Comprehensive benefits