Cloud GPU Support Engineer

Emploive

San Francisco, Northern (CA, KY)

Hybrid

USD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
Tuition Reimbursement
401(k) with employer match

Job summary

CoreWeave is hiring a Technical Support Engineer to support, operate, and maintain our extensive GPU fleet across our U.S. data centers and beyond.

You will work with customers, technicians, and engineering to ensure reliability, performance, and scale of our cloud infrastructure. You will diagnose issues, develop workload knowledge, and create internal docs while collaborating with cross-functional teams to improve hardware and software stability.

Qualifications

  • Experience in data centers, GPU clusters, server deployments, system administration, or hardware troubleshooting.
  • Intermediate knowledge of Linux (Ubuntu, CentOS, or similar), including command-line proficiency.
  • Experience with NVIDIA GPUs, SuperMicro systems, Dell HPC, and large-scale data center environments.
  • Experience in networking fundamentals (TCP/IP, VLANs, DNS, DHCP) and troubleshooting tools.
  • Hands-on experience with firmware updates, BIOS configurations, and driver management.
  • Experience analyzing system logs and debugging issues across firmware, drivers, and hardware layers.
  • Experience working with Jira, Confluence, Notion, or other issue-tracking and documentation platforms.
  • Experience in scripting and automation (Python, Bash, Ansible, or similar).

Responsibilities

  • Provide high-level support for customers utilizing bare-metal GPU fleets on CoreWeave Cloud.
  • Diagnose, triage, and investigate reported customer issues and high-priority incidents, identifying root causes and escalating when necessary.
  • Develop a deep understanding of customer workloads and use cases to provide tailored technical support.
  • Coordinate remote troubleshooting and hardware interventions with Data Center Technicians.
  • Create and maintain internal documentation, including troubleshooting guides, best practices, and knowledge base articles.
  • Participate in an on-call rotation to support production clusters and ensure operational reliability.
  • Collaborate with engineering teams to improve hardware reliability, software stability, and system performance.
  • Implement automation and scripting to streamline support workflows and reduce manual interventions.
  • Perform in-depth log analysis and debugging across multiple layers of the stack (firmware, drivers, hardware).
  • Provide feedback to internal teams on common support issues to drive continuous improvements.
  • Work with networking teams to troubleshoot connectivity issues affecting customer workloads.
  • Support supercomputing infrastructure running GPU workloads at scale.
  • Drive operational excellence by refining internal processes and support methodologies.

Skills

GPU clusters
Linux proficiency
Networking basics
Firmware & BIOS
Log analysis
Scripting & automation

Tools

Jira
Confluence
Notion

Job description

CoreWeave is hiring a Technical Support Engineer to support, operate, and maintain our extensive GPU fleet across our U.S. data centers and beyond.

You will work with customers, technicians, and engineering to ensure reliability, performance, and scale of our cloud infrastructure. You will diagnose issues, develop workload knowledge, and create internal docs while collaborating with cross-functional teams to improve hardware and software stability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Bare‑Metal Support Engineer for AI Cloud
GPU Bare‑Metal Support Engineer for AI Cloud

CoreWeave • San Francisco (CA)

On-site
USD 99,000 - 132,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+2
GPU Bare-Metal Support Engineer
GPU Bare-Metal Support Engineer

AI Chopping Block • California (MO)

Hybrid
USD 99,000 - 132,000
Medical, dental, and vision insurance
401(k) with generous match
Flexible PTO
+2
Infrastructure Support Engineering Manager - GPU/HPC
Infrastructure Support Engineering Manager - GPU/HPC

CoreWeave • Seattle (WA)

On-site
USD 157,000 - 210,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Tuition Reimbursement
+3
GPU Bare-Metal Support Engineer
GPU Bare-Metal Support Engineer

CoreWeave • Sunnyvale (CA)

On-site
USD 99,000 - 132,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+3
GPU Bare-Metal Support Engineer
GPU Bare-Metal Support Engineer

CoreWeave • Bellevue (WA)

On-site
USD 99,000 - 132,000
Medical/Dental/Vision
401(k) Matching
Flexible PTO
+7
Senior Network Services Engineer (GPU Cloud)
Senior Network Services Engineer (GPU Cloud)

CoreWeave • New York (NY)

On-site
USD 182,000 - 242,000
Medical, dental, vision insurance
Equity and 401(k) matching
Flexible PTO
Staff Network Platform Engineer – GPU Cloud Automation
Staff Network Platform Engineer – GPU Cloud Automation

CoreWeave • New York (NY)

On-site
USD 180,000 - 260,000
Medical insurance
Dental insurance
Vision insurance
+13
Senior Network Observability Engineer for GPU Cloud
Senior Network Observability Engineer for GPU Cloud

Socket.dev • Sunnyvale (CA), New York (NY)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+1
Staff Software Engineer - Network Automation
Staff Software Engineer - Network Automation

CoreWeave • Livingston (NJ)

On-site
USD 207,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+1
Technical Support Engineer (Bare Metal)
Technical Support Engineer (Bare Metal)

CoreWeave • Bellevue (WA)

On-site
USD 99,000 - 132,000
Medical/Dental/Vision
401(k) Matching
Flexible PTO
+7