HPC AI Systems Administrator

MRE Consulting

Houston (TX)

On-site

USD 120,000 - 180,000

Full time

26 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive salary
Comprehensive benefits
Professional development support

Job summary

MRE Consulting is seeking a skilled HPC AI Systems Administrator to architect and optimize a secure on-prem AI compute environment. You will enable the Development Team to train and deploy production ML models while enforcing governance and data privacy policies.

The role covers deployment, bare-metal config, GPU driver management, and containerized workloads. Collaboration with Dev teams and vendors is essential to maintain peak performance and security.

Qualifications

  • 3+ years of systems administration in Linux-based HPC or enterprise GPU environments.
  • Experience with modern enterprise GPU hardware and data center contexts.
  • Deep expertise in Linux, container tech, and cluster resource management.

Responsibilities

  • Lead deployment, bare-metal configuration, and optimization of on-prem HPC cluster and multi-GPU systems.
  • Manage AI software stack, Linux environments, GPU drivers, CUDA/NCCL, and container platforms.
  • Implement workload scheduling/orchestration (Kubernetes, Slurm) for efficient resource allocation.
  • Establish telemetry/monitoring dashboards for hardware utilization and performance.
  • Ensure security/governance, data-at-rest/in-transit baselines, and compliance.
  • Coordinate with vendors and integrators for updates and platform maintenance.

Skills

Linux admin
HPC infrastructure
Containerization
Cluster scheduling
Networking
Security governance

Education

Bachelor's degree (CS/CE)

Tools

Docker
Kubernetes
Slurm
Apptainer/Singularity
CUDA drivers

Job description

We are seeking a high-caliber HPC AI Systems Administrator to serve as the foundational architect for our growing AI infrastructure. This role will be responsible for building a secure, scalable, and highly optimized environment to support our corporate data initiatives.

Operating at the critical intersection of infrastructure engineering and software application, you will design and maintain a robust compute platform. Your primary mission is to enable our Development Team to fine-tune and deploy production-level machine learning models smoothly, while ensuring the platform complies with enterprise-level security, governance, and data privacy policies.

Key Responsibilities
  • Infrastructure Architecture & Management: Lead the deployment, bare-metal configuration, maintenance, and optimization of our on-premises HPC cluster and multi-GPU architecture.
  • Platform Enablement: Manage the end-to-end AI software stack, including Linux OS environments, specialized GPU drivers, runtime libraries (CUDA, NCCL), and containerization platforms.
  • Developer Sandbox Orchestration: Implement and maintain workload scheduling and orchestration systems (e.g., Kubernetes, Slurm, or equivalent enterprise platforms) to manage cluster resource allocation and job prioritization for engineering teams.
  • Monitoring & Performance Tuning: Establish automated telemetry and monitoring dashboards to track hardware utilization, thermal limits, and memory bandwidth, ensuring maximum compute efficiency.
  • Security & Governance Compliance: Operationalize strict data-at-rest and data-in-transit security baselines, ensuring the compute environment aligns with corporate Zero Trust network architecture and compliance mandates.
  • Vendor Relations & Support: Act as the primary technical interface for high-end hardware vendors and system integrators to manage system updates and platform maintenance.
Professional Qualifications
  • Experience: 3+ years of dedicated systems administration experience managing Linux-based High-Performance Computing (HPC) environments or enterprise-scale GPU infrastructure.
  • Technical Fluency: Hands-on experience configuring and maintaining modern enterprise GPU hardware (such as NVIDIA Ampere or Hopper architecture) in a data center context.
  • Software & Tooling Mastery: Deep expertise in Linux system engineering, container technologies (Docker, Apptainer/Singularity), and cluster resource management.
  • Networking & Storage: Solid baseline knowledge of high-throughput networking fabrics (e.g., InfiniBand/RoCE) and parallel or distributed enterprise storage systems.
  • Education: Bachelor’s degree in computer science, Computer Engineering, System Administration, or equivalent practical industry experience.
Preferred Attributes
  • Relevant professional certifications in enterprise AI infrastructure, virtualization, or cloud/hybrid architecture solutions.
  • Familiarity with the infrastructure requirements supporting modern AI frameworks, machine learning lifecycles, or Large Language Model (LLM) fine-tuning pipelines.
What We Offer
  • Direct access to cutting-edge, top-tier enterprise compute infrastructure.
  • A collaborative environment with highly defined responsibilities and strategic, top-down execution support.
  • Competitive salary, comprehensive benefits package, and professional development suppport.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Systems Administrator
AI Systems Administrator

Murray Resources - Best Staffing Agency • Houston (TX)

On-site
USD 100,000 - 140,000
401K
AI & HPC Infrastructure Engineer
AI & HPC Infrastructure Engineer

Prodapt ASIC services (Formerly Innovative Logic) • San Jose (CA)

On-site
USD 150,000 - 210,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
AI/HPC Systems Engineer
AI/HPC Systems Engineer

Saige Partners • San Jose (CA)

On-site
USD 140,000 - 190,000
AI/HPC Systems Engineer
AI/HPC Systems Engineer

Saigepartners • San Jose (CA)

Hybrid
USD 120,000 - 180,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Infrastructure Engineer
Infrastructure Engineer

HCLTech • California (MO)

On-site
USD 150,000 - 210,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
AI HPC Infrastructure Engineer
AI HPC Infrastructure Engineer

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
HPC AI Infrastructure Architect
HPC AI Infrastructure Architect

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support