Senior HPC Hardware Engineer

NorthMark Strategies LLC

Dallas, Northern (TX, KY)

Hybrid

USD 140,000 - 190,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Lunch stipend
Company-paid medical benefits
Dental and Vision for employees and 가족
16 weeks of Paid Parental Leave
Employee Assistance Program
Life insurance
Disability insurance (Short/Long Term)
401(k) company match up to 6%
Optional employee-paid benefits (PPO,H

Job summary

NorthMark Strategies LLC in Dallas, TX seeks a Senior HPC Hardware Engineer to join the Compute Engineering team. You will own the hardware lifecycle for a large-scale HPC/AI compute fleet, spanning GPU and CPU nodes, and drive automation and reliability across the environment.

The role demands hands-on leadership, experience with firmware, BIOS, Redfish, and OpenStack/Ironic, and strong collaboration with software, networking, and vendor teams. 5+ years in HPC infra preferred.

Qualifications

  • Bachelor's degree or equivalent hands-on experience.
  • 5+ years managing large-scale HPC/AI compute infra in production.
  • Deep server hardware architecture knowledge across CPUs/GPUs.
  • Hands-on with bare-metal provisioning, firmware, BIOS lifecycle and automation tools.
  • Proficiency with Redfish API and BMC/IPMI tooling for remote management.
  • Experience with Linux-based environments and scripting for automation.
  • OpenStack Ironic or similar provisioning platform preferred.
  • Strong cross-functional communication and leadership.

Responsibilities

  • Design, configure, and manage a high-performance compute fleet (GPU/CPU) across Dallas.
  • Own firmware/BIOS lifecycle from baselining to rollout and maintenance.
  • Lead hardware troubleshooting across CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, and BMCs.
  • Automate health checks, onboarding, and hardware issue remediation to accelerate deployment.
  • Validate and operationalize AI platforms (NVL72/Grace) from day one.
  • Collaborate with vendors on firmware/hardware issues with clear repro cases.
  • Perform hardware performance analysis, tuning, and capacity planning.
  • Define security hardening for hardware infrastructure and IaC workflows.
  • Mentor junior engineers and drive team-wide best practices.

Skills

server hardware architecture
bare-metal provisioning
firmware and BIOS lifecycle
hardware automation tools
Redfish API / BMC tooling
GPU/CPU diagnostics
Linux scripting (Python, Bash, Power?)
OpenStack Ironic or similar
IaC / automation
mentoring engineers

Education

Bachelor’s degree in Electrical or Computer Engineering

Tools

Ansible
Puppet
Chef
iDRAC / iLO
GPU/CPU diagnostics tooling

Job description

THE COMPANY NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients’ research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision‑making, accelerating discovery and driving faster innovation. THE POSITION NMC² is seeking a Senior HPC Hardware Engineer to join the Compute Engineering team based at our Dallas, TX offices at Victory Commons. This is a hands‑on role at the center of one of the most demanding and rapidly scaling HPC environments in operation, spanning a large fleet of GPU and CPU nodes built on the latest NVIDIA platforms including H200 and GB200/NVL72 architectures. You will own the full hardware lifecycle for NMC²’s compute fleet — from bare‑metal provisioning and firmware baseline management through production validation, troubleshooting, and capacity planning. You will be the go‑to expert on server hardware architecture, driving standards and automation that ensure the fleet operates at peak performance and availability. Your work directly enables the research and delivery workloads that NMC²’s clients depend on. This role requires close collaboration with Software Engineering, Networking, and Vendor teams, and involves mentoring junior engineers. The ideal candidate is a technically deep infrastructure leader who thrives in fast‑paced environments, brings a strong automation mindset, and has a proven track record managing large‑scale HPC or AI compute infrastructure.

RESPONSIBILITIES
  • Design, configure, and manage a high‑performance compute fleet comprising large‑scale GPU (NVIDIA V100/A100/H200/GB200) and CPU nodes across NMC²’s Dallas infrastructure.
  • Own the full firmware and BIOS lifecycle across the HPC/AI fleet — from establishing baselines and validation through rollout, compliance, and ongoing maintenance.
  • Lead troubleshooting of hardware components including CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, and BMCs; drive component replacement and configuration remediation.
  • Automate health checks, onboarding workflows, and recurring hardware issue remediation to accelerate safe deployment and reduce recovery time.
  • Validate and operationalize next‑generation AI platforms (e.g., NVL72 / Grace Blackwell) from day one, ensuring stability, performance readiness, and production fitness.
  • Collaborate with vendors on firmware and hardware issues, providing clear reproduction cases, diagnostic logs, and business impact to drive timely resolution.
  • Perform hardware performance analysis, tuning, and capacity planning to ensure reliable scale‑out of the compute environment.
  • Define and implement security hardening best practices for hardware infrastructure, maintaining platform integrity across the fleet.
  • Leverage Infrastructure as Code (IaC) methodologies and scripting to drive efficient, repeatable, and scalable infrastructure management.
  • Mentor junior engineers, act as a subject matter expert for infrastructure‑related escalations, and champion a culture of continuous improvement across the team.
REQUIREMENTS
  • Bachelor’s degree in Electrical Engineering, Computer Engineering, or a related field, or equivalent hands‑on experience.
  • 5+ years of experience managing large‑scale HPC or AI compute infrastructure in a production environment.
  • Deep knowledge of server hardware architecture, including processors, memory, storage, networking, power systems, and thermal management.
  • Hands‑on experience with bare‑metal provisioning, firmware and BIOS lifecycle management, and hardware automation tools such as Ansible, Puppet, or Chef.
  • Proficiency with Redfish API and BMC/IPMI tooling (iDRAC, iLO) for remote hardware management and diagnostics.
  • Demonstrated ability to troubleshoot and resolve complex hardware issues across GPU and CPU nodes, including NVIDIA‑SMI and GPU diagnostics.
  • Experience with hardware monitoring platforms, performance tuning, and capacity planning at scale.
  • Familiarity with Linux‑based environments and scripting proficiency in Python, Bash, or PowerShell for infrastructure automation.
  • Experience with OpenStack (particularly Ironic) or equivalent cloud/bare‑metal provisioning platforms is strongly preferred.
  • Strong cross‑functional communication skills and proven ability to collaborate effectively with software, networking, and vendor teams.
  • Prior technical leadership experience, including mentoring engineers and driving team‑wide best practices.

Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future.

BENEFITS & PERKS
  • Company-Paid Lunch Stipend: Lunch is provided via GrubHub
  • Company-Paid Benefits: 100% Employer-Paid Medical in our High Deductible Health Plan, Dental and Vision benefits for employees and their families
  • 16 weeks of Paid Parental Leave
  • Employee Assistance Program
  • Life insurance
  • Short-Term Disability and Long-Term Disability
  • 401(k): Company will match 100% of your contributions up to 6%
  • Optional Employee-Paid Benefits: Medical insurance in our PPO plan and a variety of other benefits such as Health Savings Accounts (with Company Contribution!), Flexible Spending Accounts, Supplemental Life Insurance, Wellhub and more.
  • Time Off: 25 days of Paid Time Off plus 12 company holidays
EQUAL OPPORTUNITY EMPLOYER

NORTHMARK STRATEGIES LLC IS AN EQUAL EMPLOYMENT OPPORTUNITY EMPLOYER. THE COMPANY'S POLICY IS NOT TO DISCRIMINATE AGAINST ANY APPLICANT OR EMPLOYEE BASED ON RACE, COLOR, RELIGION, NATIONAL ORIGIN, GENDER, AGE, SEXUAL ORIENTATION, GENDER IDENTITY OR EXPRESSION, MARITAL STATUS, MENTAL OR PHYSICAL DISABILITY, AND GENETIC INFORMATION, OR ANY OTHER BASIS PROTECTED BY APPLICABLE LAW. THE FIRM ALSO PROHIBITS HARASSMENT OF APPLICANTS OR EMPLOYEES BASED ON ANY OF THESE PROTECTED CATEGORIES.

NorthMark Strategies is a leading strategic capital firm, combining capital, innovation, and engineering to drive long‑term value. From operating complex businesses to backing breakthrough technologies, our mission is to deploy capital and build enduring businesses. Our team combines intelligent risk‑taking, operational excellence, exceptional talent, and world‑class computing capacity to create shareholder value. Our company offers a dynamic environment where individuals have the freedom to lead companies toward bold achievements by embracing innovation, leveraging technology, and fostering differentiated business strategies. Our values are Integrity, Ability, and Energy, and the company aims to hire individuals who possess those qualities. At NorthMark Strategies, we believe the future isn’t something to hope for, it’s something to build. We don’t just invest, we create. Bringing together strategic insight and technical horsepower to deliver outcomes that endure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NorthMark Compute and Cloud LLC • Town of Texas (WI), Fort Worth (TX)

On-site
USD 150,000 - 230,000
Lunch stipend
Medical benefits
Dental & Vision
+7
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NorthMark Strategies • Town of Texas (WI)

On-site
USD 140,000 - 190,000
Lunch stipend
Medical benefits (employer-paid)
Parental leave 16 weeks
+2
VP of Security Operations
VP of Security Operations

NorthMark Compute and Cloud LLC • Town of Texas (WI), Fort Worth (TX)

On-site
USD 180,000 - 240,000
Company-Paid Benefits: Medical, Dental
Life insurance
401(k) company match
+2
VP of Security Operations
VP of Security Operations

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 180,000 - 280,000
Health insurance
401(k) match
Paid parental leave
+3
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NorthMark Strategies LLC • Dallas (TX)

On-site
USD 150,000 - 210,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical
401(k) Company Match
+1
Director, Strategic Infrastructure Partnerships
Director, Strategic Infrastructure Partnerships

NorthMark Strategies • Dallas (TX)

On-site
USD 140,000 - 190,000
Lunch stipend
Employer-paid medical insurance
Dental and Vision benefits
+3
Emerging Network Architect
Emerging Network Architect

NorthMark Strategies LLC • Dallas (TX)

On-site
USD 140,000 - 220,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical
401(k) matching
+2
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NorthMark Strategies • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Lunch stipend
Employer-paid health & dental & vision
Parental leave 16 weeks
+4
Software Engineer, HPC Scheduling
Software Engineer, HPC Scheduling

NorthMark Strategies LLC • Dallas (TX)

On-site
USD 90,000 - 140,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical
401(k) match up to 6%
+2
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

northmark • United States

On-site
USD 130,000 - 190,000