HPC AI Systems Architect – On-Prem, Secure & Scalable

Murray Resources - Best Staffing Agency

Houston (TX)

On-site

USD 100,000 - 140,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

401K

Job summary

Murray Resources is seeking an HPC AI Systems Administrator to architect and maintain a robust on-premises compute platform that enables the development team to fine-tune and deploy production‑level ML models, while enforcing enterprise security and data governance policies.

You will lead infrastructure design, containerization (Docker, Apptainer/Singularity), cluster management, and performance tuning on a multi-GPU HPC system, with emphasis on data privacy and Zero Trust standards.

Qualifications

  • 3+ years on Linux-based HPC or enterprise GPU infrastructure
  • Hands-on with modern enterprise GPU hardware in data centers
  • Deep Linux system engineering, containers, and cluster management
  • Familiarity with high-throughput networks and enterprise storage systems

Responsibilities

  • Infrastructure Architecture & Management: Lead deployment, bare-metal config, maintenance, and optimization of on-prem HPC cluster and multi-GPU architecture.
  • Platform Enablement: Manage end-to-end AI software stack, including Linux environments, GPU drivers, runtime libraries (CUDA, NCCL), and containers.
  • Developer Sandbox Orchestration: Implement workload scheduling/orchestration (Kubernetes, Slurm) for resource allocation and job prioritization.
  • Monitoring & Performance Tuning: Create telemetry and dashboards to track hardware utilization, thermal limits, memory bandwidth; optimize compute.
  • Security & Governance Compliance: Enforce data-at-rest and data-in-transit baselines; align with Zero Trust and compliance mandates.
  • Vendor Relations & Support: Interface with hardware vendors and integrators for updates and maintenance.

Skills

HPC administration
Linux
GPU infrastructure
Container technologies
Cluster resource management
InfiniBand/RoCE networking
Security governance
CUDA/NCCL

Education

Bachelor's degree in Computer Science/Engineering

Tools

Docker
Apptainer/Singularity
Kubernetes
Slurm
CUDA
NVIDIA drivers

Job description

Murray Resources is seeking an HPC AI Systems Administrator to architect and maintain a robust on-premises compute platform that enables the development team to fine-tune and deploy production‑level ML models, while enforcing enterprise security and data governance policies.

You will lead infrastructure design, containerization (Docker, Apptainer/Singularity), cluster management, and performance tuning on a multi-GPU HPC system, with emphasis on data privacy and Zero Trust standards.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC AI Infrastructure Architect
HPC AI Infrastructure Architect

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support
AI Systems Administrator
AI Systems Administrator

Murray Resources - Best Staffing Agency • Houston (TX)

On-site
USD 100,000 - 140,000
401K
AI & HPC Systems Administrator - Hybrid (3 onsite/2 remote)
AI & HPC Systems Administrator - Hybrid (3 onsite/2 remote)

100 Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 231,000
401(k)
Pension
Vacation benefits
+1
AI HPC Systems Engineer: GPU Clusters & ML Platforms
AI HPC Systems Engineer: GPU Clusters & ML Platforms

AMD • San Jose (CA)

On-site
USD 140,000 - 210,000
Senior AI Infrastructure Engineer — HPC & MLOps
Senior AI Infrastructure Engineer — HPC & MLOps

United States Digital Space LLC • Town of Charlotte (NY)

On-site
USD 128,000 - 182,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
HPC Solutions Architecture Manager
HPC Solutions Architecture Manager

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 180,000 - 260,000
Lunch stipend
Employer-paid medical benefits
Parental leave (16 weeks)
+2
AI/HPC Systems Engineer — GPU Cluster & Hybrid Cloud
AI/HPC Systems Engineer — GPU Cluster & Hybrid Cloud

Saige Partners • San Jose (CA)

On-site
USD 140,000 - 190,000