HPC Platform Engineer

Addison Group

Dallas (TX)

Hybrid

USD 180,000 - 260,000

Full time

35 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401(k)
25 days PTO
HSA contribution
Gym membership
Lunch on office days

Job summary

Addison Group is seeking an HPC Platform Engineer in Dallas to build and operate the infrastructure behind large-scale computing workloads. You will own design, deployment, and ongoing support across compute, storage, networking, and platform software.

You will guide standards and mentor engineers. The role requires 5+ years in HPC or related infrastructure, with strong skills in Slurm, Lustre/GPFS, InfiniBand, OpenStack, and Docker/Kubernetes.

Qualifications

  • Five+ years of HPC engineering or related infrastructure experience.
  • Experience with parallel computing, distributed storage, workload scheduling, high-speed networking.
  • Proficiency with Slurm/PBS Pro; Lustre/GPFS; InfiniBand; OpenStack; Docker or Kubernetes.
  • Automation using Python with Ansible, Puppet, or Chef.
  • Strong troubleshooting and documentation skills.

Responsibilities

  • Design and implement HPC infrastructure across servers, storage, networking, and data center systems.
  • Find performance bottlenecks and improve throughput, reliability, and resource use.
  • Partner with engineering, operations, and research teams to install, configure, test, and support new systems.
  • Use monitoring data to troubleshoot complex issues and plan for future capacity needs.
  • Assess new technologies and vendor solutions, and recommend platform improvements.
  • Document architectures, configurations, and operating procedures while mentoring junior engineers.

Skills

Python automation
GPU computing
High performance networking
Monitoring tools
Troubleshooting
Documentation

Education

Bachelor’s degree or equivalent

Tools

Slurm
PBS Pro
Lustre
GPFS
InfiniBand
OpenStack
Docker
Kubernetes

Job description

Job Title: HPC Platform Engineer

Industry: High Performance Computing / AI Infrastructure

Location (city, state): Dallas, TX

Assignment Type: Direct hire

Pay: $180,000–$260,000 base salary, plus a potential $50,000–$100,000 bonus

Work Schedule: Hybrid; three days in the Dallas office and two days remote. The team manager determines the in-office schedule.

Benefits: This position is eligible for medical, dental, vision, and 401(k). Additional benefits include 25 days of PTO, an HSA contribution, a gym membership, and lunch on office days.

About The Company

Our client develops advanced computing and cloud infrastructure for demanding AI, research, and simulation workloads. The organization is investing in the systems and engineering teams needed to expand its computing capacity.

Job Description

We are seeking an HPC Platform Engineer to build, operate, and improve the infrastructure behind large-scale computing workloads. This person will work across compute, storage, networking, and platform software, taking ownership from design and deployment through ongoing support and capacity planning. The role also provides an opportunity to guide technical standards and mentor other engineers.

Key Responsibilities
  • Design and implement HPC infrastructure across servers, storage, networking, and related data center systems.
  • Find performance bottlenecks and improve system throughput, reliability, and resource use.
  • Partner with engineering, operations, and research teams to install, configure, test, and support new systems.
  • Use monitoring data to troubleshoot complex issues and plan for future capacity needs.
  • Assess new technologies and vendor solutions, and recommend improvements to the platform.
  • Document architectures, configurations, and operating procedures while helping junior engineers develop their skills.
Qualifications
  • Five or more years of HPC engineering or closely related infrastructure experience.
  • Working knowledge of parallel computing, distributed storage, workload scheduling, high-speed networking, and GPU-based systems.
  • Experience with relevant technologies such as Slurm or PBS Pro; Lustre or GPFS; InfiniBand; OpenStack; and Docker or Kubernetes.
  • Ability to automate infrastructure work with Python and tools such as Ansible, Puppet, or Chef.
  • Experience using monitoring tools to diagnose and improve platform performance.
  • Strong troubleshooting, documentation, and communication skills.
  • Bachelor’s degree or equivalent practical experience. Experience with ZFS, NiFi, or computational fluid dynamics workloads is a plus.
Additional Details

Relocation assistance may be tailored to the candidate. TN visa candidates may be considered. The interview process is expected to include an HR conversation, a hiring manager meeting, a technical discussion focused on prior experience and concepts, and an onsite visit.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Platform Engineer
HPC Platform Engineer

Gosnaphop • Dallas (TX)

Hybrid
USD 180,000 - 360,000
Medical
Dental
Vision
+5
HPC Performance and Validation Engineer
HPC Performance and Validation Engineer

Gosnaphop • Dallas (TX)

Hybrid
USD 180,000 - 260,000
100% paid medical, dental, vision
401(k)
25 days PTO
+3
HPC Orchestration Architect
HPC Orchestration Architect

Gosnaphop • Dallas (TX)

Hybrid
USD 180,000 - 260,000
100% paid medical
Dental insurance
Vision insurance
+5
Solutions Architect
Solutions Architect

Gosnaphop • Dallas (TX)

On-site
USD 180,000 - 260,000
Full medical, dental, vision
401(k)
PTO 25 days
+2
AI Systems Administrator
AI Systems Administrator

Murray Resources - Best Staffing Agency • Houston (TX)

On-site
USD 100,000 - 140,000
401K
Senior HPC Platform Engineer — Hybrid (Dallas)
Senior HPC Platform Engineer — Hybrid (Dallas)

Addison Group • Dallas (TX)

Hybrid
USD 180,000 - 260,000
Medical, dental, and vision insurance
401(k)
25 days PTO
+3
HPC Storage Engineer
HPC Storage Engineer

Selby Jennings • Dallas (TX)

On-site
USD 250,000 - 350,000
Sr. Software Engineer
Sr. Software Engineer

Addison Group • Dallas (TX)

Hybrid
USD 170,000 - 220,000
Competitive annual performance bonus
Employer-paid medical benefits for you
401(k) with generous company match
+7
HPC & AI Solutions Architect
HPC & AI Solutions Architect

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 190,000
HPC & Compute Engineering Lead
HPC & Compute Engineering Lead

Autonomai Recruitment • Chicago (IL)

On-site
USD 180,000 - 250,000