GPU Infrastructure Engineer

TechClub Inc

Fort Worth (TX)

On-site

USD 120,000 - 180,000

Full time

47 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TechClub Inc is seeking a hands-on infrastructure professional to design, deploy, and maintain enterprise-scale Linux and Windows server infrastructure in a Fort Worth data center. You will support hyperscale GPU compute environments for AI workloads, troubleshoot GPU platforms, diagnose hardware, PCIe, and network issues, and develop automation to scale operations.

The role requires strong scripting, experience with Ansible or Chef, and collaboration across engineering teams, with on-call

Qualifications

  • 10+ years of experience in systems, infrastructure, production engineering, or related roles.
  • Strong Linux administration experience, preferably with RHEL and Ubuntu.
  • Hands-on experience with enterprise server and data center infrastructure.
  • Experience troubleshooting GPU-based infrastructure and high-performance compute environments.
  • Strong understanding of server hardware, PCIe, networking, storage, and OS diagnostics.
  • Programming/scripting experience with Python and Bash.
  • Experience with infrastructure automation tools such as Ansible or Chef.
  • Strong troubleshooting and root-cause-analysis skills.
  • Experience working in large-scale production environments.
  • Excellent communication and cross-functional collaboration skills.

Responsibilities

  • Design, deploy, and maintain enterprise-scale Linux and Windows server infrastructure.
  • Support hyperscale GPU compute environments used for AI workloads.
  • Troubleshoot NVIDIA/AMD GPU platforms and hardware faults.
  • Develop automation and tooling using Python, Bash, Ansible, Chef.
  • Monitor infrastructure health using Grafana, InfluxDB, Telegraf, Nagios, and related tools.
  • Collaborate with engineering and infrastructure teams to improve reliability and efficiency.
  • Participate in on-call rotations and incident response.
  • Administer and troubleshoot networking, switches, firewalls, and data center connectivity.

Skills

Linux admin
GPU infra
Python
Bash
Ansible
Chef
Windows servers
Networking
Troubleshooting
RHEL
Ubuntu

Tools

Ansible
Chef
VMware
Hyper-V
Grafana

Job description

  • Design, deploy, maintain, and troubleshoot enterprise-scale Linux and Windows server infrastructure.
  • Support hyperscale GPU and compute environments used for AI training and inference workloads.
  • Troubleshoot NVIDIA and AMD GPU platforms, including GPU memory errors, driver issues, CUDA/ROCm failures, PCIe problems, and system-level hardware faults.
  • Work with GPU technologies and related OCP-based hardware to optimize system performance.
  • Diagnose server, rack, power, networking, PCIe, GPU, NIC, and operating-system issues.
  • Utilize Linux kernel tools, NVIDIA/AMD diagnostic utilities, system logs, and custom troubleshooting tools to identify and resolve infrastructure problems.
  • Develop automation and infrastructure tooling using Python, Bash, Ansible, Chef, and related technologies.
  • Improve operational tools and automation used across large server fleets.
  • Support VMware and Hyper-V virtualization environments where required.
  • Administer and troubleshoot networking infrastructure, including switches, firewalls, and data center connectivity.
  • Monitor infrastructure health and performance using tools such as Grafana, InfluxDB, Telegraf, Nagios, and related monitoring platforms.
  • Participate in on-call rotations and provide critical incident response for production infrastructure.
  • Identify systemic infrastructure issues and develop scalable solutions to prevent recurrence.
  • Collaborate with engineering and infrastructure teams to improve reliability, availability, and operational efficiency.
.Required Qualifications
  • 10+ years of experience in systems, infrastructure, production engineering, or related roles.
  • Strong Linux administration and troubleshooting experience, preferably with RHEL and Ubuntu.
  • Hands-on experience with enterprise server and data center infrastructure.
  • Experience troubleshooting GPU-based infrastructure and high-performance compute environments.
  • Strong understanding of server hardware, PCIe, networking, storage, and operating-system diagnostics.
  • Programming/scripting experience with Python and Bash.
  • Experience with infrastructure automation tools such as Ansible or Chef.
  • Strong troubleshooting and root-cause-analysis skills.
  • Experience working in large-scale production environments.
  • Excellent communication and cross-functional collaboration skills.
Ideal Candidate Profile

The ideal candidate is a hands-on infrastructure professional who can operate at both the hardware and software layers, from diagnosing GPU/server failures and Linux kernel issues to developing automation that improves reliability across large-scale infrastructure. Experience supporting hyperscale AI/GPU environments, combined with strong systems engineering and production troubleshooting skills, is highly valued.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Linux Admin
Linux Admin

TechDigital Group • United States

On-site
USD 100,000 - 130,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior GPU Infrastructure Engineer
Senior GPU Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
System Engineer
System Engineer

Acceler8 Talent • Fremont (CA)

On-site
USD 135,000 - 165,000
Comprehensive benefits
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits