Technical Support Engineer (Bare Metal)

Coreweave

United States

Remote

USD 110,000 - 170,000

Full time

10 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision fully covered
Life insurance and disability coverage
401(k) with employer match
Tuition reimbursement
Mental wellness benefits
Paid parental leave
Flexible paid time off
Casual, innovation-focused culture

Job summary

CoreWeave seeks a customer-facing technical support engineer to maintain reliability and performance of large bare-metal GPU fleets powering AI workloads across data centers. You will engage with customers, diagnose issues across firmware, drivers, and hardware, and influence product direction through field feedback.

You will troubleshoot, automate workflows with Python/Bash/Ansible, and collaborate with networking, hardware, and software teams to reduce recurring problems.

Qualifications

  • Hands-on experience in data centers, GPU clusters, large-scale server deployments, system administration, or hardware troubleshooting.

Responsibilities

  • Deliver high-level technical support for customers running on bare-metal GPU cloud infrastructure, diagnosing and triaging reported issues and high-priority incidents across firmware, drivers, and hardware layers.
  • Develop a deep understanding of customer workloads and use cases to provide tailored troubleshooting and guidance.
  • Coordinate remote troubleshooting and physical hardware interventions with on-site data center technicians.
  • Create and maintain internal documentation, including troubleshooting guides, best-practice articles, and knowledge-base entries.
  • Participate in an on-call rotation to support production GPU clusters and ensure operational reliability.
  • Build automation and scripts (Python, Bash, Ansible, or similar) to streamline repetitive support workflows.
  • Partner with networking, hardware, and software engineering teams to resolve complex issues and feed recurring problems back into product improvements.

Skills

Data center experience
Linux CLI
Networking fundamentals
Troubleshooting

Tools

Python
Bash
Ansible
Confluence

Job description

Role overview

A customer-facing technical support role focused on maintaining the reliability, performance, and scalability of large bare-metal GPU fleets that power AI workloads across multiple data centers. The position sits at the intersection of customer engineering, hardware operations, and infrastructure reliability, with opportunities to influence product and engineering direction through field feedback.

Responsibilities
  • Deliver high-level technical support for customers running on bare-metal GPU cloud infrastructure, diagnosing and triaging reported issues and high-priority incidents across firmware, drivers, and hardware layers.
  • Develop a deep understanding of customer workloads and use cases to provide tailored troubleshooting and guidance.
  • Coordinate remote troubleshooting and physical hardware interventions with on-site data center technicians.
  • Create and maintain internal documentation, including troubleshooting guides, best-practice articles, and knowledge-base entries.
  • Participate in an on-call rotation to support production GPU clusters and ensure operational reliability.
  • Build automation and scripts (Python, Bash, Ansible, or similar) to streamline repetitive support workflows.
  • Partner with networking, hardware, and software engineering teams to resolve complex issues and feed recurring problems back into product improvements.
Requirements
  • Hands-on experience in data centers, GPU clusters, large-scale server deployments, system administration, or hardware troubleshooting.
  • Intermediate proficiency with Linux (Ubuntu, CentOS, or similar) at the command line.
  • Working knowledge of GPU systems from vendors such as NVIDIA, server platforms such as SuperMicro and Dell, and high-performance computing environments.
  • Solid grasp of networking fundamentals (TCP/IP, VLANs, DNS, DHCP) and standard troubleshooting tools.
  • Experience with firmware updates, BIOS configuration, driver management, and multi-layer log analysis.
  • Familiarity with issue-tracking and documentation platforms (e.g., Jira, Confluence, Notion) plus scripting/automation experience.
Nice to have
  • Curiosity about Kubernetes, Docker, and other containerized infrastructure technologies.
  • Comfort working across cross-functional teams and driving continuous improvement in fast-changing environments.
Benefits and work setup
  • Medical, dental, and vision insurance fully covered for employees.
  • Company-paid life insurance plus short- and long-term disability coverage, flexible spending and health savings accounts.
  • 401(k) with employer match, employee stock purchase program eligibility, and tuition reimbursement.
  • Mental wellness benefits, family-forming support, paid parental leave, and full-service childcare assistance.
  • Flexible paid time off and a casual, innovation-focused work culture with catered meals at office and data center locations.
  • This position involves access to export-controlled information and is limited to U.S. persons as defined by applicable export regulations.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Technical Program Manager
Senior Technical Program Manager

nebius • United States

On-site
USD 115,000 - 275,000
Healthcare coverage
401(k) with company match
Parental leave (20 weeks primary, 12–?
+3
GPU Systems Engineer 3 with Security Clearance
GPU Systems Engineer 3 with Security Clearance

Base-2 Solutions • Chevy Chase (MD)

On-site
USD 120,000 - 180,000
Company-paid health premiums
Dental premiums
Vision premiums
+3
GPU Supercomputing Reliability Engineer — Unlimited PTO
GPU Supercomputing Reliability Engineer — Unlimited PTO

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Software Platform Support Engineer - GPU Cloud
Software Platform Support Engineer - GPU Cloud

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 108,000 - 173,000
Software Engineer, Platform
Software Engineer, Platform

fal - Features & Labels • United States

Remote
USD 150,000 - 210,000
Interesting and challenging work
Learning and growth opportunities
Visa sponsorship and relocation to San
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Equity incentives
Technical Support Engineer
Technical Support Engineer

hyperbolic • United States

Remote
USD 110,000 - 150,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Systems Administrator
Systems Administrator

CyberCoders • Houston (TX)

On-site
USD 120,000 - 140,000
Comprehensive benefits
PTO
401k with match
+1