AI Systems Administrator

Murray Resources - Best Staffing Agency

Houston (TX)

On-site

USD 100,000 - 140,000

Full time

32 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

401K

Job summary

Murray Resources is seeking an HPC AI Systems Administrator to architect and maintain a robust on-premises compute platform that enables the development team to fine-tune and deploy production‑level ML models, while enforcing enterprise security and data governance policies.

You will lead infrastructure design, containerization (Docker, Apptainer/Singularity), cluster management, and performance tuning on a multi-GPU HPC system, with emphasis on data privacy and Zero Trust standards.

Qualifications

  • 3+ years on Linux-based HPC or enterprise GPU infrastructure
  • Hands-on with modern enterprise GPU hardware in data centers
  • Deep Linux system engineering, containers, and cluster management
  • Familiarity with high-throughput networks and enterprise storage systems

Responsibilities

  • Infrastructure Architecture & Management: Lead deployment, bare-metal config, maintenance, and optimization of on-prem HPC cluster and multi-GPU architecture.
  • Platform Enablement: Manage end-to-end AI software stack, including Linux environments, GPU drivers, runtime libraries (CUDA, NCCL), and containers.
  • Developer Sandbox Orchestration: Implement workload scheduling/orchestration (Kubernetes, Slurm) for resource allocation and job prioritization.
  • Monitoring & Performance Tuning: Create telemetry and dashboards to track hardware utilization, thermal limits, memory bandwidth; optimize compute.
  • Security & Governance Compliance: Enforce data-at-rest and data-in-transit baselines; align with Zero Trust and compliance mandates.
  • Vendor Relations & Support: Interface with hardware vendors and integrators for updates and maintenance.

Skills

HPC administration
Linux
GPU infrastructure
Container technologies
Cluster resource management
InfiniBand/RoCE networking
Security governance
CUDA/NCCL

Education

Bachelor's degree in Computer Science/Engineering

Tools

Docker
Apptainer/Singularity
Kubernetes
Slurm
CUDA
NVIDIA drivers

Job description

A fast-growing manufacturing company is seeking an HPC AI Systems Administrator to serve as the foundational architect for its growing AI infrastructure. This role will design and maintain a robust compute platform that enables the Development Team to fine-tune and deploy production-level machine learning models, while ensuring the platform complies with enterprise-level security, governance, and data privacy policies. The HPC AI Systems Administrator is responsible for building a secure, scalable, and highly optimized environment to support the company's corporate data initiatives.

Salary + Additional Benefits:
  • $100,000–$140,000 (Dependent on Experience)
  • 401K
Location:

Houston, TX

Type of Position:

Direct Hire

Responsibilities:
  • Infrastructure Architecture & Management: Lead the deployment, bare-metal configuration, maintenance, and optimization of our on-premises HPC cluster and multi-GPU architecture.
  • Platform Enablement: Manage the end-to-end AI software stack, including Linux OS environments, specialized GPU drivers, runtime libraries (CUDA, NCCL), and containerization platforms.
  • Developer Sandbox Orchestration: Implement and maintain workload scheduling and orchestration systems (e.g., Kubernetes, Slurm, or equivalent enterprise platforms) to manage cluster resource allocation and job prioritization for engineering teams.
  • Monitoring & Performance Tuning: Establish automated telemetry and monitoring dashboards to track hardware utilization, thermal limits, and memory bandwidth, ensuring maximum compute efficiency.
  • Security & Governance Compliance: Operationalize strict data-at-rest and data-in-transit security baselines, ensuring the compute environment aligns with corporate Zero Trust network architecture and compliance mandates.
  • Vendor Relations & Support: Act as the primary technical interface for high-end hardware vendors and system integrators to manage system updates and platform maintenance.
Requirements:
  • 3+ years of dedicated systems administration experience managing Linux-based High-Performance Computing (HPC) environments or enterprise-scale GPU infrastructure
  • Hands-on experience configuring and maintaining modern enterprise GPU hardware (such as NVIDIA Ampere or Hopper architecture) in a data center context
  • Deep expertise in Linux system engineering, container technologies (Docker, Apptainer/Singularity), and cluster resource management
  • Solid baseline knowledge of high-throughput networking fabrics (e.g., InfiniBand/RoCE) and parallel or distributed enterprise storage systems
  • Bachelor's degree in computer science, Computer Engineering, System Administration, or equivalent practical industry experience
  • Relevant professional certifications in enterprise AI infrastructure, virtualization, or cloud/hybrid architecture solutions (preferred)
  • Familiarity with the infrastructure requirements supporting modern AI frameworks, machine learning lifecycles, or Large Language Model (LLM) fine-tuning pipelines (preferred)

Due to the high volume of applications we typically receive, we regret that we are not able to personally respond to all applications. However, if you are invited to take the next step in the process, you will typically be contacted within one week of submitting your application.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
AI/HPC Systems Engineer
AI/HPC Systems Engineer

Saige Partners • San Jose (CA)

On-site
USD 140,000 - 190,000
HPC/AI Technical Solution Engineer
HPC/AI Technical Solution Engineer

VC5 Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer

NVIDIA • Colorado

On-site
USD 124,000 - 196,000
AI/HPC System Engineer
AI/HPC System Engineer

Norland Group • San Jose (CA)

On-site
USD 103,000 - 117,000
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 120,000 - 190,000
AMD benefits at a glance