Senior Manager, Technical Support Engineering - Cloud

Coreweave

Livingston (NJ)

On-site

USD 198,000 - 264,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CoreWeave is seeking a Senior Manager of Technical Support Engineering - Infrastructure to lead a global, 24/7 infrastructure support team. You will own the function’s strategy, staffing, and performance, guiding engineers and coordinating with Product Engineering and field specialists to resolve complex customer issues end-to-end.

Expect to scale the organization, implement quality programs, and shape the direct-to-expert support model.

Qualifications

  • 5+ years of people-leadership experience within technical support/operations.
  • 8+ years total in technical support/operations, including running a 24/7 function in cloud environments.
  • Strong background in Linux, containerization, and Kubernetes.
  • Experience with hyperscale/cloud infrastructure a plus.
  • Proven ability to lead, coach, and scale teams and processes.

Responsibilities

  • Own and scale the 24/7 infrastructure support function across regions.
  • Manage talent acquisition, onboarding, and performance coaching.
  • Stay hands-on with complex customer issues as escalation point.
  • Develop enablement, training, and progression paths for staff.
  • Implement quality measures and CSAT/SLO targets for the function.
  • Define operating models, staffing, roles, and escalation boundaries.
  • Lead customer comms during critical incidents and resolve conflicts.
  • Track KPIs and plan multi-quarter growth for the function.
  • Coordinate with Product Engineering and domain teams to align feedback to roadmaps.

Skills

People leadership
Linux
Kubernetes
Cloud infrastructure
Incident management
Program management

Tools

Slurm

Job description

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.

About the Team

The Technical Support Engineering - Infrastructure team sits within CoreWeave's Global Field Organization (GFO) and is the front line for every customer running AI and HPC workloads on our platform. Operating 24/7/365, the team supports the Kubernetes-powered infrastructure behind the AI revolution: GPU compute, high-performance networking and storage, Slurm/HPC clusters, and the large-scale, mission-critical training workloads that run on them.

We operate on a Direct-to-Expert model. Instead of routing customers through generic support tiers, we get the right expert on the problem fast while a single owner stays accountable for the issue end-to-end. That means fewer handoffs, clearer ownership, and a high-touch experience that our customers feel on every ticket. The team is the connective tissue between our customers and CoreWeave's engineering organization. We triage and resolve deep technical issues directly, coordinate with Product Engineering, Specialist Field Engineers, and domain specialists when needed, and feed what we learn in the field straight back into the product roadmap.

About the Role

As the Senior Manager of Technical Support Engineering - Infrastructure, you'll own, scale, and continuously improve CoreWeave's infrastructure support function — a large, globally distributed 24/7/365 organization of skilled engineers who resolve our customers' most complex technical challenges with deep expertise, efficiency, and empathy. High-touch, expert-led support is one of CoreWeave's clearest differentiators, which gives this role prominent visibility and influence across the company.

You'll operate at the department level, setting the technical and operational bar for the entire function and designing the operating systems (coverage/staffing model, Direct-to-Expert partnerships, quality program, metrics) that let support scale reliably. You'll lead with empathy and invest in each person's growth, while building the structure and processes that let the team scale with CoreWeave's hyper-growth without compromising the culture we care deeply about.

In this role, you will:
  • Own the strategy, health, and performance of the entire 24/7/365 infrastructure support function, helping scale coverage and capability across regions and domains as CoreWeave grows.
  • Own talent acquisition and retention by hiring, onboarding, and developing engineers through diligent performance management and coaching tailored to each individual's needs.
  • Stay hands-on: dig into complex, customer-impacting issues alongside your team and serve as a senior technical escalation point, ensuring the highest quality of support.
  • Build and facilitate the enablement and career-development frameworks for onboarding, technical training, and progression paths that raise capability across the whole function and create a repeatable path for career growth.
  • Implement quality assurance measures, including ticket reviews and best-practice playbooks, that raise the bar on resolution speed, accuracy, and consistency.
  • Own the support operating model for the function including roles and responsibilities, escalation boundaries, coverage/staffing model, and the response and resolution targets (SLOs) that keep every customer's experience consistently excellent as volume scales.
  • Lead customer communication during critical incidents and resolve conflicts with clarity, composure, and empathy.
  • Track and report on KPIs focused on team performance and customer satisfaction, and own the strategic planning for the team's growth and scalability.
  • Own the cross-functional interface between support and Product Engineering, Specialist Field Engineers, and domain teams, building the alignment mechanisms and feedback loops that make the Direct-to-Expert model work at scale for escalations, live incidents, and customer needs.
  • Champion the voice of the customer, turning recurring support patterns into product, tooling, and process improvements.
  • Help set the multi-quarter vision and operating plan for infrastructure support, and represent the function in company-level planning, headcount, and prioritization discussions.
Who you are:
  • You have 5+ years of people-leadership experience and 8+ years total in technical support/operations, including running a 24/7 support function at scale in a cloud operations environment.
  • You have a strong background in Linux, containerization technologies, and Kubernetes, and you understand virtualization and cloud computing concepts. Experience at a hyperscaler or cloud infrastructure provider is a strong plus.
  • You lead with empathy and aren't afraid to get your hands dirty. You do the work alongside your direct reports and model the standard you set.
  • You're energized by leadership excellence and talent development: diligent performance management, coaching, and growing each engineer according to their individual needs.
  • You've built enablement and quality programs that scaled across a function or multiple teams — and can show the measurable improvement they drove.
  • You've designed the operating model for a support org, including coverage/staffing model, escalation boundaries, SLOs, and evolved it as the org scaled.
  • You're a calm, clear communicator with customers and executives during critical incidents, and you resolve conflicts effectively across teams and organizational boundaries.
  • You think in systems and multi-quarter plans. You've defined KPI/SLO frameworks, reported to senior leadership, and owned capacity and growth planning for a function, not just a single team.
  • You've led a globally distributed team across time zones.
  • Leadership & Communication: proven ability to lead through senior talent and set direction for a function, with executive-level communication skills.
  • Strategic & Operational Planning: you build operating systems, plans, and metrics that scale a function beyond what any one person can hold.
  • Problem-Solving & Adaptability: robust problem-solving skills and adaptability at organizational scale, in a fast-paced, hyper-growth environment.
  • Program Management: experience with program-management tools and methodologies.
Bonus points if you have:
  • Experience supporting AI/ML, HPC, or GPU-accelerated workloads at scale.
  • Hands-on Kubernetes operations experience (CKA certification a plus).
  • Familiarity with Slurm/SUNK, RDMA networking, distributed storage, and observability tooling such as Grafana.
  • You have some experience with infrastructure as it relates to Data Center Operations.

Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk.

  • You're an expert in what it takes to be an excellent leader and foster an environment in which people are excited and inspired to participate.
  • You love to dive into problems, test for solutions, and enjoy engaging with customers.
  • You're excited and curious about AI.

Why CoreWeave?

At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values:

  • Be Curious at Your Core
  • Act Like an Owner
  • Empower Employees
  • Deliver Best-in-Class Client Experiences
  • Achieve More Together

We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!

The base salary range for this role is $198,000 to $264,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager, Technical Support Engineering - Cloud
Senior Manager, Technical Support Engineering - Cloud

CoreWeave • Sunnyvale (CA)

On-site
USD 198,000 - 264,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+3
Senior Manager, Technical Support Engineering - Cloud
Senior Manager, Technical Support Engineering - Cloud

CoreWeave • Bellevue (WA)

On-site
USD 198,000 - 264,000
Medical, dental, and vision insurance
401(k) with match
Flexible PTO
+1
Senior Manager, Technical Support Engineering - Cloud
Senior Manager, Technical Support Engineering - Cloud

CoreWeave • New York (NY)

On-site
USD 198,000 - 264,000
Medical, dental, vision insurance
401(k) with employer match
Flexible PTO
+3
Technical Support Engineer - III (Cloud)
Technical Support Engineer - III (Cloud)

CoreWeave • New York (NY)

On-site
USD 122,000 - 163,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+1
Manager, Technical Support Engineering (Bare Metal)
Manager, Technical Support Engineering (Bare Metal)

Socket.dev • Sunnyvale (CA), Seattle (WA), San Francisco (CA)

On-site
USD 157,000 - 210,000
Medical, dental, and vision insurance
401(k) with generous employer match
Flexible PTO
+2
Manager, Technical Support Engineer
Manager, Technical Support Engineer

CoreWeave • Livingston (NJ)

On-site
USD 198,000 - 264,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+3
Staff Software Engineer
Staff Software Engineer

CoreWeave • San Francisco (CA)

On-site
USD 207,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+2
Staff Software Engineer
Staff Software Engineer

CoreWeave • Bellevue (WA)

On-site
USD 207,000 - 275,000
Medical, dental, vision insurance
Company-paid Life Insurance
401(k) with employer match
+7
Staff Software Engineer
Staff Software Engineer

CoreWeave • Livingston (NJ)

On-site
USD 207,000 - 275,000
Medical, dental, and vision insurance—
Company-paid Life Insurance
Tuition Reimbursement
+4
Principal Engineer, Cloud Infrastructure Services
Principal Engineer, Cloud Infrastructure Services

Weights & Biases • United States

Hybrid
USD 206,000 - 303,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+3