GPU Infrastructure Support Engineer - Tier 2/3 (Remote)

Hydra Host

Miami (FL)

Remote

USD 90,000 - 130,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Hydra Host is seeking a Support Engineer for GPU infrastructure (Tier 2/3) to diagnose issues across Linux hosts, hardware, and network layers in a production environment. This remote role involves coordinating hands-on hardware steps and writing up findings to prevent repeats.

You will handle incidents end-to-end, contribute to runbooks, and work with data center partners to drive timely solutions. Expect a dynamic, less-structured setting with broad influence on how our support evolves.

Qualifications

  • Three+ years supporting production servers, data center infrastructure, or bare metal and cloud environments.
  • Strong hands-on Linux troubleshooting with logs, dmesg, systemd, storage tooling, and network utilities.
  • Experience with server hardware: CPU/memory, storage, RAID, PCIe, NICs, power, BIOS/UEFI, firmware and drivers.
  • Out-of-band management experience: IPMI, Redfish, iDRAC, iLO or similar.
  • Working TCP/IP knowledge; ability to verify host vs. network issues.
  • Reasoning in fault domains: impact of failures and remote fixability.
  • Experience with ticketing, monitoring, incident management, or infra management systems.
  • Clear written English; able to communicate well with customers and engineers.
  • Sound judgment in production; capable of making calls during high-stakes issues.

Responsibilities

  • Diagnose across the stack from Linux host to hardware and facility.
  • Troubleshoot Linux server issues: boot failures, kernel/driver problems, storage pressure, services, memory, CPU behavior.
  • Diagnose hardware failures using IPMI/Redfish/iDRAC, sensor data, POST/boot errors, SMART data, vendor tools.
  • Isolate server-side network problems: NICs, VLANs, routing, DNS, DHCP, bonding, packet captures.
  • Triage GPU server issues: GPU availability, driver/VBIOS problems, PCIe, XID errors.
  • Own incidents end-to-end: from detection to resolution or clean handoff.
  • Produce runbooks and post-incident reviews; contribute to RCA.

Skills

Linux troubleshooting
Server hardware experience
IPMI/Redfish/iDRAC/iLO knowledge
Networking TCP/IP
Incident management & ticketing

Tools

IPMI
Redfish
iDRAC
iLO

Job description

Hydra Host is seeking a Support Engineer for GPU infrastructure (Tier 2/3) to diagnose issues across Linux hosts, hardware, and network layers in a production environment. This remote role involves coordinating hands-on hardware steps and writing up findings to prevent repeats.

You will handle incidents end-to-end, contribute to runbooks, and work with data center partners to drive timely solutions. Expect a dynamic, less-structured setting with broad influence on how our support evolves.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Infra Support Engineer (Remote, Tier 2/3)
Senior GPU Infra Support Engineer (Remote, Tier 2/3)

Hydra Host, Inc. • Miami (FL), Northern (KY)

Hybrid
USD 90,000 - 130,000
Remote HPC Solutions Engineer: GPU Clusters & SLURM
Remote HPC Solutions Engineer: GPU Clusters & SLURM

Hydra Host • United States

Remote
USD 120,000 - 160,000
Competitive salary
Equity
Benefits
Senior GPU Infra Engineer — Remote
Senior GPU Infra Engineer — Remote

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Senior GPU Infrastructure Support Engineer
Senior GPU Infrastructure Support Engineer

Nscale • San Francisco (CA)

On-site
USD 120,000 - 170,000
Equity
Remote-friendly team
Flexible workplace
Senior GPU Compute Solutions Engineer
Senior GPU Compute Solutions Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 140,000 - 224,000
Equity
Benefits
Technical Lead - GPU Infrastructure (Fully Remote)
Technical Lead - GPU Infrastructure (Fully Remote)

Jobgether SRL • United States

Remote
USD 170,000 - 250,000
Remote GPU Infra NOC Engineer | Automation & AI Ops
Remote GPU Infra NOC Engineer | Automation & AI Ops

Orion Placement • Pittsburgh

On-site
USD 75,000 - 140,000
Dental insurance
Paid time off
Retirement plan
+2
Senior Storage Engineer
Senior Storage Engineer

Kindredventures • Miami (FL)

On-site
USD 120,000 - 150,000
Remote GPU Infra NOC Engineer — Automation & AI Ops
Remote GPU Infra NOC Engineer — Automation & AI Ops

Orionplacement • Pittsburgh

On-site
USD 75,000 - 140,000
Bonus and equity opportunities
Medical, dental, and vision insurance
401(k)
Senior AI/HPC Storage Architect
Senior AI/HPC Storage Architect

Hydra Host, Inc. • Miami (FL)

On-site
USD 120,000 - 150,000