Lead HPC Systems & Performance Engineer

Stanford Black Limited

Dallas (TX)

On-site

USD 140,000 - 210,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Stanford Black Limited seeks a Lead HPC Systems & Performance Engineer to take ownership of compute performance in large-scale HPC environments. You will evaluate CPU, GPU and accelerator technologies, benchmark hardware, identify bottlenecks, and translate test data into architecture decisions.

Working with storage, networking, vendors, and customers, you''ll help build and optimize end-to-end HPC platforms across the compute stack, applying deep expertise in performance tuning and AI/ML

Qualifications

  • Masters or PhD in Computer Science, Systems Engineering, or similar field.
  • Strong background in HPC systems, performance engineering or systems architecture.
  • Deep understanding of CPU, GPU and accelerator architectures.
  • Strong knowledge of memory hierarchy, NUMA and system performance.
  • Hands-on experience with Linux tuning, optimisation and performance profiling.
  • Experience designing or optimising HPC clusters and parallel computing environments.
  • Understanding of InfiniBand/RoCE and their impact on compute performance.
  • Experience benchmarking and evaluating new hardware.
  • Knowledge of distributed/parallel storage and its relationship with compute performance.
  • Experience supporting demanding AI/ML, scientific computing or simulation workloads.

Responsibilities

  • Own compute performance across large-scale HPC environments.
  • Evaluate emerging CPU, GPU and accelerator technologies and benchmark new hardware.
  • Collaborate with storage/networking, vendors, and internal teams to optimize end-to-end HPC platforms.

Skills

HPC systems
Performance engineering
Systems architecture
CPU architectures
GPU architectures
Memory hierarchy
NUMA
Linux tuning
Performance profiling
HPC clusters
Parallel computing
InfiniBand
RoCE
AI/ML workloads
NVIDIA DCGM
Nsight
MLPerf
GPU/CPU/DPU clusters

Education

Masters
PhD

Tools

DCGM
Nsight
MLPerf

Job description

Lead HPC Systems & Performance Engineer (Dallas, Texas):

The Company:

My client is a rapidly growing HPC and AI Supercompute Lab operating at the forefront of hyperscale compute.

With significant investment behind its growth, the business is building next-generation infrastructure supporting demanding AI, scientific research, simulation, and data-intensive workloads.

The Role:

They are looking for a Lead HPC Systems & Performance Engineer to take ownership of compute performance across large-scale HPC environments.

You'll evaluate emerging CPU, GPU and accelerator technologies, benchmark new hardware, identify performance bottlenecks, and turn real-world test data into architecture decisions.

Working across the compute stack, you'll collaborate with storage and networking specialists, hardware vendors, internal engineering teams and customers to build and optimise end-to-end HPC platforms.

Required Skills:

  • Masters &/or PhD in Computer Science, Systems Engineering, or similar subject(s).
  • Strong background in HPC systems, performance engineering or systems architecture.
  • Deep understanding of CPU, GPU and accelerator architectures.
  • Strong knowledge of memory hierarchy, NUMA and system performance.
  • Hands-on experience with Linux tuning, optimisation and performance profiling.
  • Experience designing or optimising HPC clusters and parallel computing environments.
  • Understanding of InfiniBand/RoCE and their impact on compute performance.
  • Experience benchmarking and evaluating new hardware.
  • Knowledge of distributed/parallel storage and its relationship with compute performance.
  • Experience supporting demanding AI/ML, scientific computing or simulation workloads.
  • NVIDIA GPU ecosystem and tools such as DCGM, Nsight or MLPerf.
  • Experience with large-scale GPU/CPU/DPU clusters & emerging/pre-production hardware.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Dallas-Based Lead HPC Systems & Performance Engineer
Dallas-Based Lead HPC Systems & Performance Engineer

Stanford Black Limited • Dallas (TX)

On-site
USD 140,000 - 210,000
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 140,000 - 190,000
Sr HPC Hardware Engineer
Sr HPC Hardware Engineer

Career Techniques • Dallas (TX)

Hybrid
USD 120,000 - 180,000
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 150,000 - 200,000
HPC Architect Leader
HPC Architect Leader

VC5 Consulting • Houston (TX)

On-site
USD 140,000 - 210,000
High-Performance Computing Talent Community
High-Performance Computing Talent Community

CGG Services SAS • Houston (TX)

On-site
USD 90,000 - 130,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
System Engineer
System Engineer

Acceler8 Talent • Fremont (CA)

On-site
USD 135,000 - 165,000
Comprehensive benefits
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000