Infrastructure Engineer, GPU & Compute

Jobtailor

California (MO)

On-site

USD 110,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking an experienced Infrastructure Engineer to own and evolve image management, deployment, and validation across bare-metal infrastructure. You will run test clusters, validate firmware and OS images on GPU-enabled systems, and support hardware qualification for next-gen platforms.

You will diagnose complex issues across GPUs, drivers, OS, and hardware, analyze performance with NVIDIA DCGM, and build automation for provisioning and validation.

Qualifications

  • 5+ years of experience in infrastructure engineering or related roles.
  • Strong Linux experience in production environments.
  • Hands-on experience with GPU-enabled systems and NVIDIA DCGM.
  • Familiarity with bare-metal provisioning and system bring-up workflows.
  • Proficiency in Python or similar scripting/programming languages for automation.
  • Ability to debug complex issues across hardware, OS, GPUs, and system software.

Responsibilities

  • Own and evolve systems for image management, deployment, and validation across bare-metal infrastructure.
  • Run and maintain test clusters used for system validation, diagnostics, and bring-up.
  • Validate firmware, drivers, and OS images across compute and GPU-enabled systems.
  • Support hardware qualification efforts for next-generation platforms.
  • Own GPU diagnostics and validation workflows across large-scale infrastructure.
  • Diagnose and resolve complex issues across GPUs, drivers, OS, and hardware layers.
  • Analyze system and GPU performance using tools such as NVIDIA DCGM.
  • Identify failure patterns and drive improvements in system stability and validation coverage.
  • Build and maintain automation for provisioning, validation, and system bring-up.
  • Develop Python-based tools and workflows to improve efficiency and reduce manual operational overhead.
  • Improve the reliability, repeatability, and scalability of image pipelines and validation systems.
  • Manage and operate Linux-based systems in production and validation environments.
  • Manage virtualization technology.
  • Support bare-metal provisioning workflows, including PXE and image-based systems.
  • Interface with hardware management systems (e.g., IPMI, Redfish) for monitoring and debugging.
  • Partner with Infrastructure, Hardware, and Data Center teams on system bring-up and validation.
  • Collaborate with platform and ML teams to ensure systems meet workload requirements.
  • Contribute to best practices for provisioning, diagnostics, and lifecycle management of infrastructure

Skills

Linux
Python
GPU‑enabled systems
Debugging

Tools

NVIDIA DCGM
IPMI
Redfish
PXE

Job description

Responsibilities
  • Own and evolve systems for image management, deployment, and validation across bare-metal infrastructure
  • Run and maintain test clusters used for system validation, diagnostics, and bring-up
  • Validate firmware, drivers, and OS images across compute and GPU-enabled systems
  • Support hardware qualification efforts for next-generation platforms
  • Own GPU diagnostics and validation workflows across large-scale infrastructure
  • Diagnose and resolve complex issues across GPUs, drivers, OS, and hardware layers
  • Analyze system and GPU performance using tools such as NVIDIA DCGM
  • Identify failure patterns and drive improvements in system stability and validation coverage
  • Build and maintain automation for provisioning, validation, and system bring-up
  • Develop Python-based tools and workflows to improve efficiency and reduce manual operational overhead
  • Improve the reliability, repeatability, and scalability of image pipelines and validation systems
  • Manage and operate Linux-based systems in production and validation environments
  • Manage virtualization technology
  • Support bare-metal provisioning workflows, including PXE and image-based systems
  • Interface with hardware management systems (e.g., IPMI, Redfish) for monitoring and debugging
  • Partner with Infrastructure, Hardware, and Data Center teams on system bring-up and validation
  • Collaborate with platform and ML teams to ensure systems meet workload requirements
  • Contribute to best practices for provisioning, diagnostics, and lifecycle management of infrastructure
Requirements
  • 5+ years of experience in infrastructure engineering, systems engineering, or related roles
  • Strong Linux systems experience in production environments
  • Hands‑on experience with GPU‑enabled systems and tools such as NVIDIA DCGM
  • Familiarity with bare‑metal provisioning and system bring‑up workflows
  • Proficiency in Python or similar scripting/programming languages for automation
  • Ability to debug complex issues across hardware, OS, GPUs, and system software
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Linux Admin
Linux Admin

TechDigital Group • United States

On-site
USD 100,000 - 130,000
GPU & Compute Infra Engineer: Validation & Automation
GPU & Compute Infra Engineer: Validation & Automation

Jobtailor • California (MO)

On-site
USD 110,000 - 160,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Harrison Clarke • United States

On-site
USD 100,000 - 140,000
Staff Software Architect, Compute & GPU Infrastructure
Staff Software Architect, Compute & GPU Infrastructure

CoreWeave • New York (NY)

On-site
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Engineering Lab Technician
Engineering Lab Technician

SproutsAI • Santa Clara (CA)

On-site
USD 60,000 - 80,000
Senior Solutions Architect, AI Infrastructure
Senior Solutions Architect, AI Infrastructure

Jobtailor • California (MO)

On-site
USD 120,000 - 160,000
Compute Platform Engineer
Compute Platform Engineer

NMC2 • Dallas (TX)

On-site
USD 100,000 - 130,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000