HPC Platform Engineer

Gosnaphop

Dallas (TX)

Hybrid

USD 180,000 - 360,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical
Dental
Vision
401(k)
PTO 25 days
HSA contribution
Gym membership
Lunch on office days

Job summary

Gosnaphop seeks an experienced HPC Platform Engineer to design, implement, and operate the infrastructure behind large-scale computing workloads in a hybrid Dallas, TX environment. You will collaborate across compute, storage, networking, and platform software, guiding standards and mentoring teammates.

Responsibilities include capacity planning, performance tuning, and vendor evaluations while documenting architectures and procedures.

Qualifications

  • Five or more years of HPC engineering or closely related infrastructure experience.
  • Working knowledge of parallel computing, distributed storage, workload scheduling, high-speed networking, and GPU-based systems.
  • Experience with Slurm or PBS Pro; Lustre or GPFS; InfiniBand; OpenStack; and Docker or Kubernetes.
  • Ability to automate infrastructure work with Python and tools such as Ansible, Puppet, or Chef.
  • Experience using monitoring tools to diagnose and improve platform performance.
  • Strong troubleshooting, documentation, and communication skills.
  • Bachelor’s degree or equivalent practical experience. Experience with ZFS, NiFi, or computational fluid dynamics workloads is a plus.

Responsibilities

  • Design and implement HPC infrastructure across servers, storage, networking, and related data center systems.
  • Find performance bottlenecks and improve system throughput, reliability, and resource use.
  • Partner with engineering, operations, and research teams to install, configure, test, and support new systems.
  • Use monitoring data to troubleshoot complex issues and plan for future capacity needs.
  • Assess new technologies and vendor solutions, and recommend improvements to the platform.
  • Document architectures, configurations, and operating procedures while helping junior engineers develop their skills.

Skills

HPC experience
Parallel computing
System troubleshooting
Documentation
Communication

Education

Bachelor’s degree or equivalent

Tools

Slurm
PBS Pro
Lustre
GPFS
OpenStack
Docker
Kubernetes
Python
Ansible
Puppet
Chef
Monitoring tools

Job description

Job Title: HPC Platform Engineer

Industry: High Performance Computing / AI Infrastructure

Location (city, state): Dallas, TX

Assignment Type: Direct hire

Pay: $180,000–$260,000 base salary, plus a potential $50,000–$100,000 bonus

Work Schedule: Hybrid; three days in the Dallas office and two days remote. The team manager determines the in-office schedule.

Benefits: This position is eligible for medical, dental, vision, and 401(k). Additional benefits include 25 days of PTO, an HSA contribution, a gym membership, and lunch on office days.

About The Company:

Our client develops advanced computing and cloud infrastructure for demanding AI, research, and simulation workloads. The organization is investing in the systems and engineering teams needed to expand its computing capacity.

Job Description:

We are seeking an HPC Platform Engineer to build, operate, and improve the infrastructure behind large-scale computing workloads. This person will work across compute, storage, networking, and platform software, taking ownership from design and deployment through ongoing support and capacity planning. The role also provides an opportunity to guide technical standards and mentor other engineers.

Key Responsibilities:
  • Design and implement HPC infrastructure across servers, storage, networking, and related data center systems.
  • Find performance bottlenecks and improve system throughput, reliability, and resource use.
  • Partner with engineering, operations, and research teams to install, configure, test, and support new systems.
  • Use monitoring data to troubleshoot complex issues and plan for future capacity needs.
  • Assess new technologies and vendor solutions, and recommend improvements to the platform.
  • Document architectures, configurations, and operating procedures while helping junior engineers develop their skills.
Qualifications:
  • Five or more years of HPC engineering or closely related infrastructure experience.
  • Working knowledge of parallel computing, distributed storage, workload scheduling, high-speed networking, and GPU-based systems.
  • Experience with relevant technologies such as Slurm or PBS Pro; Lustre or GPFS; InfiniBand; OpenStack; and Docker or Kubernetes.
  • Ability to automate infrastructure work with Python and tools such as Ansible, Puppet, or Chef.
  • Experience using monitoring tools to diagnose and improve platform performance.
  • Strong troubleshooting, documentation, and communication skills.
  • Bachelor’s degree or equivalent practical experience. Experience with ZFS, NiFi, or computational fluid dynamics workloads is a plus.
Additional Details:

Relocation assistance may be tailored to the candidate. TN visa candidates may be considered. The interview process is expected to include an HR conversation, a hiring manager meeting, a technical discussion focused on prior experience and concepts, and an onsite visit.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Platform Engineer
HPC Platform Engineer

Addison Group • Dallas (TX)

Hybrid
USD 180,000 - 260,000
Medical, dental, and vision insurance
401(k)
25 days PTO
+3
HPC Performance and Validation Engineer
HPC Performance and Validation Engineer

Gosnaphop • Dallas (TX)

Hybrid
USD 180,000 - 260,000
100% paid medical, dental, vision
401(k)
25 days PTO
+3
HPC Orchestration Architect
HPC Orchestration Architect

Gosnaphop • Dallas (TX)

Hybrid
USD 180,000 - 260,000
100% paid medical
Dental insurance
Vision insurance
+5
Solutions Architect
Solutions Architect

Gosnaphop • Dallas (TX)

On-site
USD 180,000 - 260,000
Full medical, dental, vision
401(k)
PTO 25 days
+2
AI Systems Administrator
AI Systems Administrator

Murray Resources - Best Staffing Agency • Houston (TX)

On-site
USD 100,000 - 140,000
401K
Senior HPC Platform Engineer — Hybrid (Dallas)
Senior HPC Platform Engineer — Hybrid (Dallas)

Addison Group • Dallas (TX)

Hybrid
USD 180,000 - 260,000
Medical, dental, and vision insurance
401(k)
25 days PTO
+3
HPC Storage Engineer
HPC Storage Engineer

Selby Jennings • Dallas (TX)

On-site
USD 250,000 - 350,000
Sr. Software Engineer
Sr. Software Engineer

Addison Group • Dallas (TX)

Hybrid
USD 170,000 - 220,000
Competitive annual performance bonus
Employer-paid medical benefits for you
401(k) with generous company match
+7
HPC & AI Solutions Architect
HPC & AI Solutions Architect

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 190,000
HPC & Compute Engineering Lead
HPC & Compute Engineering Lead

Autonomai Recruitment • Chicago (IL)

On-site
USD 180,000 - 250,000