HPC AI Infrastructure Architect

MRE Consulting

Houston (TX)

On-site

USD 120,000 - 180,000

Full time

24 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive salary
Comprehensive benefits
Professional development support

Job summary

MRE Consulting is seeking a skilled HPC AI Systems Administrator to architect and optimize a secure on-prem AI compute environment. You will enable the Development Team to train and deploy production ML models while enforcing governance and data privacy policies.

The role covers deployment, bare-metal config, GPU driver management, and containerized workloads. Collaboration with Dev teams and vendors is essential to maintain peak performance and security.

Qualifications

  • 3+ years of systems administration in Linux-based HPC or enterprise GPU environments.
  • Experience with modern enterprise GPU hardware and data center contexts.
  • Deep expertise in Linux, container tech, and cluster resource management.

Responsibilities

  • Lead deployment, bare-metal configuration, and optimization of on-prem HPC cluster and multi-GPU systems.
  • Manage AI software stack, Linux environments, GPU drivers, CUDA/NCCL, and container platforms.
  • Implement workload scheduling/orchestration (Kubernetes, Slurm) for efficient resource allocation.
  • Establish telemetry/monitoring dashboards for hardware utilization and performance.
  • Ensure security/governance, data-at-rest/in-transit baselines, and compliance.
  • Coordinate with vendors and integrators for updates and platform maintenance.

Skills

Linux admin
HPC infrastructure
Containerization
Cluster scheduling
Networking
Security governance

Education

Bachelor's degree (CS/CE)

Tools

Docker
Kubernetes
Slurm
Apptainer/Singularity
CUDA drivers

Job description

MRE Consulting is seeking a skilled HPC AI Systems Administrator to architect and optimize a secure on-prem AI compute environment. You will enable the Development Team to train and deploy production ML models while enforcing governance and data privacy policies.

The role covers deployment, bare-metal config, GPU driver management, and containerized workloads. Collaboration with Dev teams and vendors is essential to maintain peak performance and security.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC AI Systems Architect – On-Prem, Secure & Scalable
HPC AI Systems Architect – On-Prem, Secure & Scalable

Murray Resources - Best Staffing Agency • Houston (TX)

On-site
USD 100,000 - 140,000
401K
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
Senior AI Infrastructure Engineer — HPC & MLOps
Senior AI Infrastructure Engineer — HPC & MLOps

United States Digital Space LLC • Town of Charlotte (NY)

On-site
USD 128,000 - 182,000
AI & HPC Infra Engineer: GPU Compute & Cloud Automation
AI & HPC Infra Engineer: GPU Compute & Cloud Automation

Prodapt ASIC services (Formerly Innovative Logic) • San Jose (CA)

On-site
USD 150,000 - 210,000
AI Systems Engineer: Cloud, GPUs & Automation
AI Systems Engineer: Cloud, GPUs & Automation

MCI • United States

On-site
USD 90,000 - 130,000
AI Systems Administrator
AI Systems Administrator

Murray Resources - Best Staffing Agency • Houston (TX)

On-site
USD 100,000 - 140,000
401K
AI HPC Systems Engineer: GPU Clusters & ML Platforms
AI HPC Systems Engineer: GPU Clusters & ML Platforms

AMD • San Jose (CA)

On-site
USD 140,000 - 210,000
HPC Deployment Manager — AI Data Center & Validation
HPC Deployment Manager — AI Data Center & Validation

NVIDIA Corporation • Northern (KY)

Hybrid
USD 216,000 - 397,000
Equity
Benefits
HPC Solutions Architecture Manager
HPC Solutions Architecture Manager

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 180,000 - 260,000
Lunch stipend
Employer-paid medical benefits
Parental leave (16 weeks)
+2