HPC Performance and Validation Engineer

NorthMark Compute & Cloud

Dallas (TX)

On-site

USD 110,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NorthMark Compute & Cloud is seeking an HPC Validation and Performance Engineer to own validation and optimization of HPC CPU/GPU farms. You will develop a validation and baselining framework for AI/ML workloads, drive performance benchmarking, and implement advanced tooling to scale our datacenter footprint.

You will lead cross-functional efforts, build automation using Python/Go/Kubernetes, and implement modern monitoring with Prometheus and Grafana to deliver data-driven insights for

Qualifications

  • Bachelor's degree or equivalent experience required.
  • Experience profiling and tuning with large GPU clusters.
  • Deep knowledge of NVIDIA ClusterKit, Nsight, MLPerf and DCGM.
  • Networking & storage perf; InfiniBand/RoCe profiling and optimization.
  • Automation tooling and micro-benchmarking using Python and Go; Kubernetes in Ubuntu.
  • Experience with HPC workloads across distributed global locations.

Responsibilities

  • Architect and implement a validation framework for GPU nodes in a large HPC environment.
  • Define methods to continually assess performance and optimize AI/ML workloads.
  • Develop and execute comprehensive performance tests using benchmarks for HPC compute, storage and networking.
  • Contribute to research reports describing benchmarking findings and HW performance.
  • Lead bottleneck debugging and resolution efforts in system performance.
  • Build scalable tools for automated validation and testing using Python, Go, Kubernetes and CI/CD.
  • Implement monitoring with Prometheus, Grafana, OTEL and related technologies for real-time health.
  • Define observability strategy for HPC validation and performance monitoring.
  • Stay informed on industry trends to ensure long-term strategic alignment.
  • Collaborate cross-functionally with engineering, infra and research teams.

Skills

Accelerator performance
Profiling & tuning
Automation tooling
Leadership
Cross-functional collaboration

Education

Bachelor's degree or equivalent experience

Tools

NVIDIA ClusterKit
Nsight
MLPerf
DCGM
Phoronix Test Suite
Python
Go
Kubernetes
iPerf
Prometheus
Grafana

Job description

NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients' research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis and decision-making, accelerating discovery and driving faster innovation.

The Position

As an HPC Validation and Performance Engineer at NMC², you will take ownership of the validation and optimization of our HPC CPU and GPU calc farms. This critical role will involve developing a validation and performance baselining framework, which ensures system readiness for AI/ML and HPC workloads across multiple architectures. Your role will be essential in providing continuous performance benchmarking, real-time observability and long-term strategic readiness. You will drive the implementation of advanced tooling and frameworks, maintaining an infrastructure that is crucial to our cutting-edge research efforts. You will be accountable for providing data-driven performance metrics to support architectural design choices as we continue to globally scale our datacenter footprint. We are looking for someone with deep technical expertise in compute, storage or networking optimizations and performance engineering who can develop solutions that scale with our growing infrastructure. This role demands a forward-thinking engineer who can anticipate industry trends and adopt emerging architectures and strategies to keep NMC² at the forefront of innovation.

Responsibilities
  • Architecting and implementing a validation framework to certify the readiness and utilization of GPU nodes across a large, distributed HPC environment.
  • Defining methodologies to continually assess performance and optimize infrastructure across AI/ML workloads.
  • Developing and executing comprehensive performance testing using industry and customer specific benchmarks, ensuring optimal performance across HPC compute, storage and networking.
  • Contribute to research reports that will describe the discoveries of the benchmarking, evaluating the complete HW performance and efficiency.
  • Leading efforts to debug, identify and then resolve bottlenecks in system performance.
  • Building robust, scalable tools for automated validation and testing, utilising Python, Go, Kubernetes and CI/CD pipelines to streamline continuous validation and benchmarking processes.
  • Implementing monitoring solutions using Prometheus, Grafana and other modern monitoring technologies to track performance metrics and real-time health of the cluster.
  • Defining and implementing best practice for continuous performance validation, ensuring that the infrastructure remains reliable and efficient as new technologies emerge.
  • Staying informed on industry trends and advancements to ensure long-term strategic alignment.
  • Working cross-functionally with engineering, infrastructure and research teams to align validation efforts with the broader business objectives, ensuring that the platform meets evolving research demands.
Requirements
  • Bachelor's Degree or equivalent experience.
  • Accelerator performance experience, including profiling and tuning with large-scale GPU clusters.
  • In-depth understanding of NVIDIA ClusterKit, Nsight and Validation Suite, MLPerf and DCGM tools for GPU and DPUs.
  • Networking & storage performance experience, including profiling and optimisation with NVIDIA ClusterKit, iPerf or equivalent across InfiniBand/RoCe network implementations.
  • System benchmarking experience across Linux and familiarity with the Phronix suite or equivalent.
  • Experience with HPC workloads across distributed global locations, bringing data-driven performance data to compliment key architectural decisions.
  • Strong proficiency in developing automation tools and micro benchmarking frameworks for validation using Python, Go and Kubernetes in a Ubuntu Linux environment.
  • Expertise with key monitoring platforms including OTEL, Prometheus, ELK and Grafana and in definition and implementing the overall observability strategy for HPC validation and performance monitoring.
  • A deep understanding of emerging technologies, architectures and strategies, with the ability to assess their potential impact on infrastructure and adopt them as part of a long-term plan.
  • Proven ability to lead complex technical projects, influence decisions and engage with stakeholders across technical and research teams.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Compute Platform Engineer
Compute Platform Engineer

NMC2 • Dallas (TX)

On-site
USD 100,000 - 130,000
HPC Performance & Validation Architect
HPC Performance & Validation Architect

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 110,000 - 170,000
Compute Platform Engineer
Compute Platform Engineer

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 120,000 - 180,000
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 150,000
Senior Data Center Performance Engineer - Benchmarking and Optimization
Senior Data Center Performance Engineer - Benchmarking and Optimization

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Equity
Benefits
HPC Operations Engineer
HPC Operations Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Equity eligibility
Diverse work environment
Comprehensive benefits
Emerging Network Architect
Emerging Network Architect

NMC2 • Dallas (TX)

On-site
USD 100,000 - 140,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NMC2 • Dallas (TX)

On-site
USD 120,000 - 160,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 140,000 - 210,000
Principal Software Engineer – E2E Performance, Goodput
Principal Software Engineer – E2E Performance, Goodput

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000