Manager of Technical Support Engineering (Bare Metal)

Coreweave

Sunnyvale (CA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

CoreWeave in Sunnyvale, CA seeks a Manager of Technical Support Engineering to lead the CX Infrastructure Support team, overseeing hardware monitoring and incident escalations while building a scalable Bare Metal support model across client environments.

You will mentor engineers, collaborate with product and engineering, manage escalations, and drive process improvements to ensure reliable, high-performance infrastructure for AI workloads.

Qualifications

  • Experience leading teams responsible for infrastructure support, data center operations, or physical compute environments.
  • Hands-on Linux system administration and command-line tools.
  • Experience with ticket-based workflows (Jira, Zendesk) in high-urgency environments.
  • Understanding of GPU infrastructure and HPC environments.
  • Proven incident and escalation management with ownership of production-impacting issues.
  • Travel up to 30% annually.
  • Strong scheduling and team logistics in 24/7 or hybrid models.
  • Ability to interpret metrics (MTTR, SLOs) to drive improvements.
  • Track record improving infrastructure reliability and accountability.

Responsibilities

  • Lead daily support operations, triage incidents, and drive escalations.
  • Oversee a team of Systems Operations Engineers and build Bare Metal support.
  • Maintain and optimize physical infrastructure across client environments.
  • Develop and lead a dedicated Infrastructure Support team.
  • Oversee incident resolution and collaboration with internal teams.
  • Improve support processes to reduce downtime and meet client expectations.
  • Work with product, infrastructure, and other teams to ensure seamless delivery of resources.
  • Manage client communications during escalations and resolutions.
  • Mentor teammates to grow expertise in critical infrastructure.
  • Ensure scalability of operations as the company grows.

Skills

Leadership
Infrastructure management
Linux administration
Incident escalation
Jira
Zendesk
GPU hardware
Data center operations
24/7 support
Escalation handling

Tools

Jira
Zendesk
Linux

Job description

Job Responsibilities
  • The Customer Experience (CX) Organization at CoreWeave is dedicated to ensuring every client running AI workloads at scale has a seamless, reliable, and high-performance experience. This team supports the infrastructure that powers the AI revolution working across data centers, hardware systems, and customer workloads to maintain the integrity of our cloud platform
  • The CX organization aligns closely with the internal and customer engineering teams, offering valuable insights from the field and having the chance to contribute to the CoreWeave product roadmap and development
  • As a Manager of Technical Support Engineering, you’ll be at the center of ensuring our dedicated infrastructure remains stable, reliable, and performant. You’ll lead daily support operations, triage incidents, drive escalations, and ensure that hardware is monitored, maintained, and delivered effectively for our clients
  • You’ll oversee a team of experienced Systems Operations Engineers and help build a new team focused on our Bare Metal support model. This role balances tactical execution with operational maturity, working cross-functionally with engineering, product, and infrastructure teams to scale processes as we grow
  • Lead a skilled team responsible for maintaining and optimizing physical infrastructure across multiple client environments
  • Build, develop, and lead a dedicated Infrastructure Support team focused on supporting key infrastructure, handling escalations, and ensuring smooth hardware operations
  • Oversee the resolution of infrastructure-related incidents, escalation management, and collaborate with internal teams to deliver effective solutions
  • Improve support processes to enhance efficiency and reduce downtime, ensuring the infrastructure meets client expectations
  • Work closely with product, infrastructure, and other teams to ensure seamless delivery of infrastructure resources
  • Manage client communication during escalations and issue resolution to ensure transparency and client satisfaction
  • Mentor team members, developing their skills to manage and maintain critical infrastructure effectively
Requirements
  • Experience working with high-performance rack-scale hardware, including CPU and GPU-based compute nodes
  • Familiarity with hardware-level diagnostics, troubleshooting, and replacement (servers, power, cabling, etc.)
  • Experience managing ticket-based workflows (Jira, Zendesk, etc.) in a high-urgency technical environment
  • Understanding of GPU infrastructure (e.g., NVIDIA A100/H100s, PCIe/NVLink, liquid cooling) or a demonstrated ability to quickly learn and adapt to HPC environments
  • Proven track record in incident and escalation management, with direct ownership of client or production-impacting issues
  • Travel up to 30% annually
  • 5+ years of experience leading teams responsible for infrastructure support, data center operations, or physical compute environments
  • Hands-on experience with Linux system administration and command-line tools
  • Skilled in managing scheduling, shift coverage, and team logistics in 24/7 or hybrid support models
  • Comfortable interpreting and acting on metrics (MTTR, SLOs, backlog, ticket trends) to drive operational improvements
  • Have a track record of improving infrastructure reliability through clear processes and team accountability
About You

Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams even if you aren’t a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talk

  • Thrive in fast-paced environments where priorities can shift quickly
  • Think critically about how to scale operations without overengineering
  • Comfortable getting close to the work, but know when to step back and lead
  • Communicate clearly, especially under pressure
  • Care about delivering for customers but know when to hold the line to protect the team and long-term goals
  • Experience managing infrastructure support teams in high-growth or rapidly evolving environments
  • Proven ability to develop and implement operational processes that scale with business needs
  • Strong familiarity with server and GPU hardware lifecycle management: deployment, maintenance, thermal/power concerns, RMA coordination, and decommissioning
  • Demonstrated success in coaching and growing technical teams through training, mentorship, and performance development
  • Familiarity with AI/ML workloads, cluster utilization patterns, or the infrastructure needs of GPU-heavy clients is a plus
  • Skilled in both developing and interpreting metrics to drive accountability, continuous improvement, and executive visibility
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Manager, Technical Support Engineering (Bare Metal)
Manager, Technical Support Engineering (Bare Metal)

CoreWeave • San Francisco (CA)

On-site
USD 157,000 - 210,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
Manager, Technical Support Engineering (Bare Metal)
Manager, Technical Support Engineering (Bare Metal)

Socket.dev • San Francisco (CA)

On-site
USD 157,000 - 210,000
Medical, dental, and vision insurance
401(k) with generous employer match
Flexible PTO
+2
Manager, Technical Support Engineering (Bare Metal)
Manager, Technical Support Engineering (Bare Metal)

CoreWeave • Seattle (WA)

On-site
USD 157,000 - 210,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Tuition Reimbursement
+3
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Senior Specialist Field Engineer - Compute Infrastructure
Senior Specialist Field Engineer - Compute Infrastructure

CoreWeave • Bellevue (WA)

On-site
USD 188,000 - 275,000
401(k) with employer match
Medical, dental, and vision insurance
Tuition Reimbursement
+1
Senior Specialist Field Engineer - Compute Infrastructure
Senior Specialist Field Engineer - Compute Infrastructure

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical/Dental/Vision insurance
401(k) with company match
Tuition Reimbursement
+1
Sr HPC Hardware Engineer
Sr HPC Hardware Engineer

Career Techniques • Dallas (TX)

Hybrid
USD 120,000 - 180,000
Senior Specialist Field Engineer - Compute Infrastructure
Senior Specialist Field Engineer - Compute Infrastructure

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical insurance
401(k) match
Paid time off
+3
Senior Manager, Technical Solutions Manager
Senior Manager, Technical Solutions Manager

CoreWeave • Livingston (NJ)

On-site
USD 207,000 - 275,000
Medical Insurance
401k Matching
Flexible PTO
+5
Senior Manager, Technical Solutions Manager
Senior Manager, Technical Solutions Manager

CoreWeave • New York (NY)

On-site
USD 207,000 - 275,000
Medical/Dental/Vision Insurance
401(k) with match
Flexible PTO
+2