HPC AI Systems Architect (On-Prem GPU Cluster)

MRE Consulting

Houston (TX)

On-site

USD 95,000 - 140,000

Full time

40 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MRE Consulting in Houston seeks a HPC AI Systems Administrator to architect a secure, scalable on-prem HPC compute platform, enabling your development teams to fine-tune and deploy production ML models. You will lead deployment, manage multi-GPU hardware, and oversee Linux, GPU drivers, CUDA, NCCL, containers, and orchestration tools, ensuring performance and security.

This role requires 3+ years in HPC or enterprise GPU infra, strong Linux, and experience with Slurm/Kubernetes, InfiniBand

Qualifications

  • 3+ years of dedicated systems administration experience managing Linux-based HPC environments or enterprise-scale GPU infrastructure.
  • Hands-on experience configuring and maintaining modern enterprise GPU hardware (NVIDIA Ampere or Hopper) in a data center context.
  • Deep expertise in Linux system engineering, container technologies (Docker, Apptainer/Singularity), and cluster resource management.
  • Solid baseline knowledge of high-throughput networking fabrics (InfiniBand/RoCE) and parallel or distributed enterprise storage systems.
  • Bachelor’s degree in computer science, Computer Engineering, System Administration, or equivalent practical industry experience.

Responsibilities

  • Infrastructure Architecture & Management: Lead deployment, bare-metal configuration, maintenance, and optimization of our on-premises HPC cluster and multi-GPU architecture.
  • Platform Enablement: Manage the end-to-end AI software stack, including Linux OS environments, specialized GPU drivers, runtime libraries (CUDA, NCCL), and containerization platforms.
  • Developer Sandbox Orchestration: Implement and maintain workload scheduling and orchestration systems (e.g., Kubernetes, Slurm, or equivalent enterprise platforms) to manage cluster resource allocation and job prioritization for engineering teams.
  • Monitoring & Performance Tuning: Establish automated telemetry and monitoring dashboards to track hardware utilization, thermal limits, and memory bandwidth, ensuring maximum compute efficiency.
  • Security & Governance Compliance: Operationalize strict data-at-rest and data-in-transit security baselines, ensuring the compute environment aligns with corporate Zero Trust network architecture and compliance mandates.
  • Vendor Relations & Support: Act as the primary technical interface for high-end hardware vendors and system integrators to manage system updates and platform maintenance.

Skills

Linux HPC admin
GPU infrastructure
Container tech
Cluster management
Networking InfiniBand
Storage systems

Education

Bachelor’s degree in CS/CE/Systems Admin

Tools

Docker
Apptainer/Singularity
Kubernetes
Slurm
NVIDIA GPUs

Job description

MRE Consulting in Houston seeks a HPC AI Systems Administrator to architect a secure, scalable on-prem HPC compute platform, enabling your development teams to fine-tune and deploy production ML models. You will lead deployment, manage multi-GPU hardware, and oversee Linux, GPU drivers, CUDA, NCCL, containers, and orchestration tools, ensuring performance and security.

This role requires 3+ years in HPC or enterprise GPU infra, strong Linux, and experience with Slurm/Kubernetes, InfiniBand

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits
AI/ML Engineer for HPC & Distributed Systems (Semi-Remote)
AI/ML Engineer for HPC & Distributed Systems (Semi-Remote)

Hewlett Packard Enterprise • Spring (TX)

On-site
USD 121,000 - 277,000
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters

Parallel Works • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
AI Infrastructure Engineer - GPU & HPC Expert
AI Infrastructure Engineer - GPU & HPC Expert

Socket.dev • Town of Florida (NY)

On-site
USD 87,000 - 266,000
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI/HPC Cluster Architect
AI/HPC Cluster Architect

Socket.dev • Austin (TX)

On-site
USD 140,000 - 230,000
AI/HPC Cluster Architect
AI/HPC Cluster Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits
AI/HPC Cluster Architect - Scalable Data Center Design
AI/HPC Cluster Architect - Scalable Data Center Design

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
AMD benefits
Equal opportunity employer
Visa sponsorship not available