HPC AI Systems Administrator

MRE Consulting

Houston (TX)

On-site

USD 95,000 - 140,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MRE Consulting in Houston seeks a HPC AI Systems Administrator to architect a secure, scalable on-prem HPC compute platform, enabling your development teams to fine-tune and deploy production ML models. You will lead deployment, manage multi-GPU hardware, and oversee Linux, GPU drivers, CUDA, NCCL, containers, and orchestration tools, ensuring performance and security.

This role requires 3+ years in HPC or enterprise GPU infra, strong Linux, and experience with Slurm/Kubernetes, InfiniBand

Qualifications

  • 3+ years of dedicated systems administration experience managing Linux-based HPC environments or enterprise-scale GPU infrastructure.
  • Hands-on experience configuring and maintaining modern enterprise GPU hardware (NVIDIA Ampere or Hopper) in a data center context.
  • Deep expertise in Linux system engineering, container technologies (Docker, Apptainer/Singularity), and cluster resource management.
  • Solid baseline knowledge of high-throughput networking fabrics (InfiniBand/RoCE) and parallel or distributed enterprise storage systems.
  • Bachelor’s degree in computer science, Computer Engineering, System Administration, or equivalent practical industry experience.

Responsibilities

  • Infrastructure Architecture & Management: Lead deployment, bare-metal configuration, maintenance, and optimization of our on-premises HPC cluster and multi-GPU architecture.
  • Platform Enablement: Manage the end-to-end AI software stack, including Linux OS environments, specialized GPU drivers, runtime libraries (CUDA, NCCL), and containerization platforms.
  • Developer Sandbox Orchestration: Implement and maintain workload scheduling and orchestration systems (e.g., Kubernetes, Slurm, or equivalent enterprise platforms) to manage cluster resource allocation and job prioritization for engineering teams.
  • Monitoring & Performance Tuning: Establish automated telemetry and monitoring dashboards to track hardware utilization, thermal limits, and memory bandwidth, ensuring maximum compute efficiency.
  • Security & Governance Compliance: Operationalize strict data-at-rest and data-in-transit security baselines, ensuring the compute environment aligns with corporate Zero Trust network architecture and compliance mandates.
  • Vendor Relations & Support: Act as the primary technical interface for high-end hardware vendors and system integrators to manage system updates and platform maintenance.

Skills

Linux HPC admin
GPU infrastructure
Container tech
Cluster management
Networking InfiniBand
Storage systems

Education

Bachelor’s degree in CS/CE/Systems Admin

Tools

Docker
Apptainer/Singularity
Kubernetes
Slurm
NVIDIA GPUs

Job description

We are seeking a high-caliber HPC AI Systems Administrator to serve as the foundational architect for our growing AI infrastructure. This role will be responsible for building a secure, scalable, and highly optimized environment to support our corporate data initiatives.

Operating at the critical intersection of infrastructure engineering and software application, you will design and maintain a robust compute platform. Your primary mission is to enable our Development Team to fine-tune and deploy production-level machine learning models smoothly, while ensuring the platform complies with enterprise-level security, governance, and data privacy policies.

Key Responsibilities
  • Infrastructure Architecture & Management: Lead the deployment, bare-metal configuration, maintenance, and optimization of our on-premises HPC cluster and multi-GPU architecture.
  • Platform Enablement: Manage the end-to-end AI software stack, including Linux OS environments, specialized GPU drivers, runtime libraries (CUDA, NCCL), and containerization platforms.
  • Developer Sandbox Orchestration: Implement and maintain workload scheduling and orchestration systems (e.g., Kubernetes, Slurm, or equivalent enterprise platforms) to manage cluster resource allocation and job prioritization for engineering teams.
  • Monitoring & Performance Tuning: Establish automated telemetry and monitoring dashboards to track hardware utilization, thermal limits, and memory bandwidth, ensuring maximum compute efficiency.
  • Security & Governance Compliance: Operationalize strict data-at-rest and data-in-transit security baselines, ensuring the compute environment aligns with corporate Zero Trust network architecture and compliance mandates.
  • Vendor Relations & Support: Act as the primary technical interface for high-end hardware vendors and system integrators to manage system updates and platform maintenance.
Professional Qualifications
  • Experience: 3+ years of dedicated systems administration experience managing Linux-based High-Performance Computing (HPC) environments or enterprise-scale GPU infrastructure.
  • Technical Fluency: Hands-on experience configuring and maintaining modern enterprise GPU hardware (such as NVIDIA Ampere or Hopper architecture) in a data center context.
  • Software & Tooling Mastery: Deep expertise in Linux system engineering, container technologies (Docker, Apptainer/Singularity), and cluster resource management.
  • Networking & Storage: Solid baseline knowledge of high-throughput networking fabrics (e.g., InfiniBand/RoCE) and parallel or distributed enterprise storage systems.
  • Education: Bachelor’s degree in computer science, Computer Engineering, System Administration, or equivalent practical industry experience.
Preferred Attributes
  • Relevant professional certifications in enterprise AI infrastructure, virtualization, or cloud/hybrid architecture solutions.
  • Familiarity with the infrastructure requirements supporting modern AI frameworks, machine learning lifecycles, or Large Language Model (LLM) fine-tuning pipelines.
What We Offer
  • Direct access to cutting-edge, top-tier enterprise compute infrastructure.
  • A collaborative environment with highly defined responsibilities and strategic, top-down execution support.
  • Competitive salary, comprehensive benefits package, and professional development suppport.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Solution Architect - AI Infrastructure
HPC Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Healthcare + Dental + Vision
401(k)
+1
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Director, Advanced Infrastructure Solutions (AI/HPC)
Director, Advanced Infrastructure Solutions (AI/HPC)

CSS • Austin (TX)

On-site
USD 140,000 - 180,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
AI and ML Infra Software Engineer, GPU Clusters
AI and ML Infra Software Engineer, GPU Clusters

Jobtailor • California (MO)

On-site
USD 120,000 - 190,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 176,000 - 334,000
AI Platform Engineer
AI Platform Engineer

Park Place Technologies in • Highland Heights (OH)

On-site
USD 90,000 - 130,000
AI HPC Infrastructure Engineer
AI HPC Infrastructure Engineer

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package