HPC Performance and Validation Engineer

Gosnaphop

Dallas (TX)

Hybrid

USD 180,000 - 260,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

100% paid medical, dental, vision
401(k)
25 days PTO
HSA contribution
Lunch on office days
Gym membership

Job summary

Gosnaphop is seeking an HPC Performance and Validation Engineer in Dallas, TX to measure, validate, and improve distributed HPC system performance. You will build repeatable GPU readiness tests for production workloads, investigate bottlenecks, and guide infrastructure decisions with data.

The role uses Python/Go, Kubernetes, and CI/CD to create validation workflows, with hybrid work and relocation options. A strong focus on performance, profiling, and observability is required.

Qualifications

  • Experience benchmarking, profiling, and tuning large-scale GPU clusters and HPC workloads.
  • Familiar with accelerator performance tools and GPU debugging/test utilities.
  • Proficiency building automated validation tools with Python or Go on Linux; Kubernetes experience is needed.
  • Knowledge of observability tools (Prometheus, Grafana, OpenTelemetry, ELK) is important.

Responsibilities

  • Build automated tests that verify GPU node readiness, health, and utilization across a large HPC environment.
  • Benchmark compute, storage, and network performance with established tests.
  • Profile AI/research workloads; identify bottlenecks and drive improvements with teams.
  • Develop validation tools and CI/CD workflows using Python, Go, Kubernetes.
  • Establish monitoring and reporting to show cluster health and performance trends.
  • Document benchmark findings to guide architecture and capacity decisions.
  • Lead technical investigations and collaborate with infrastructure and research teams.

Skills

GPU benchmarking
Performance profiling
Technical leadership
Python programming
Go programming

Education

Bachelor's degree in a relevant field

Tools

NVIDIA Nsight
DCGM
ClusterKit
MLPerf
Prometheus
Grafana
OpenTelemetry
ELK stack
Kubernetes
Python
Go

Job description

Job Title: HPC Performance and Validation Engineer

Industry: High Performance Computing / AI Infrastructure

Location (city, state): Dallas, TX

Assignment Type: Direct hire

Pay: $180,000–$260,000 base salary, plus a potential $50,000–$100,000 bonus

Work Schedule: Hybrid; three days in the Dallas office and two days remote. The manager determines the in-office days.

Benefits: This position is eligible for 100% paid medical, dental, vision, and 401(k). Additional benefits include 25 days of PTO, an HSA contribution, lunch on office days, and a gym membership.

About The Company:

Our client is expanding the computing infrastructure used for large-scale AI, research, and simulation workloads.

Job Description:

We are seeking an engineer to measure, validate, and improve the performance of distributed HPC systems. You will create repeatable ways to confirm GPU clusters are ready for production workloads, investigate performance problems, and help technical teams make informed infrastructure decisions.

Key Responsibilities:

  • Build automated tests that verify GPU node readiness, health, and utilization across a large HPC environment.
  • Benchmark compute, storage, and network performance using established and workload-specific tests.
  • Profile AI and research workloads, identify bottlenecks, and work with engineering teams to improve results.
  • Develop validation tools and continuous testing workflows using Python, Go, Kubernetes, and CI/CD pipelines.
  • Establish monitoring and reporting that make cluster health and performance trends visible.
  • Document benchmark findings and use the results to guide architecture and capacity decisions.
  • Lead technical investigations and collaborate with infrastructure and research teams as the platform grows.

Qualifications:

  • Experience benchmarking, profiling, and tuning large-scale GPU clusters and HPC workloads.
  • Strong understanding of accelerator performance and tools such as NVIDIA Nsight, DCGM, ClusterKit, or MLPerf.
  • Experience testing network and storage performance, including InfiniBand or RoCE environments.
  • Proficiency building automated validation tools with Python or Go in Linux environments; Kubernetes experience is also needed.
  • Familiarity with observability tools such as Prometheus, Grafana, OpenTelemetry, or the ELK stack.
  • Ability to lead complex technical work, explain findings clearly, and influence decisions across teams.
  • A degree is preferred, but relevant experience is more important.

Additional Details:

Relocation assistance may be tailored to the candidate, and TN visa candidates may be considered. The anticipated interview process includes an HR screen, a hiring manager meeting, a discussion of technical experience and concepts, and an onsite visit.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Platform Engineer
HPC Platform Engineer

Gosnaphop • Dallas (TX)

Hybrid
USD 180,000 - 360,000
Medical
Dental
Vision
+5
HPC Platform Engineer
HPC Platform Engineer

Addison Group • Dallas (TX)

Hybrid
USD 180,000 - 260,000
Medical, dental, and vision insurance
401(k)
25 days PTO
+3
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
HPC Performance & Validation Engineer (Hybrid, Dallas)
HPC Performance & Validation Engineer (Hybrid, Dallas)

Gosnaphop • Dallas (TX)

Hybrid
USD 180,000 - 260,000
100% paid medical, dental, vision
401(k)
25 days PTO
+3
Systems Engineer – HPC & GPU Infrastructure
Systems Engineer – HPC & GPU Infrastructure

FiveInsights • Bethesda (MD)

On-site
USD 170,000 - 210,000
Systems Engineer – HPC & GPU Infrastructure
Systems Engineer – HPC & GPU Infrastructure

MAXISIQ, Inc. • Bethesda (MD)

On-site
USD 170,000 - 210,000
HPC Orchestration Architect
HPC Orchestration Architect

Gosnaphop • Dallas (TX)

Hybrid
USD 180,000 - 260,000
100% paid medical
Dental insurance
Vision insurance
+5
HPC & AI Solutions Architect
HPC & AI Solutions Architect

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 190,000
Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NMC2 • Dallas (TX), Northern (KY)

On-site
USD 140,000 - 190,000