Principal Software Engineer, GPU Compute

OneClick Smart Resume

San Mateo (CA)

On-site

USD 345,000 - 399,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Roblox is seeking a Principal Software Engineer, GPU Compute, in San Mateo, CA to lead GPU compute initiatives and own the machine management layer for production-scale accelerators. You will partner with Kubernetes, networking, and cloud teams to drive end-to-end GPU strategy and reliability.

The role requires deep GPU expertise, strong Go skills, and experience operating GPU/AI workloads in production. Onsite work in our San Mateo office is expected with equity and benefits available.

Qualifications

  • 10+ years of experience building and operating large-scale distributed systems and infrastructure.
  • Deep, hands-on GPU expertise at the machine management layer and above: GPU host provisioning, driver and firmware lifecycle, GPU health and reliability.
  • Proven track record as compute expert with scalable GPU infrastructure.
  • Strong proficiency in Go or other well-structured programming languages.
  • Experience operating GPU and AI workloads in production, including CUDA, GPU scheduling, and high-performance networking (NVLink, InfiniBand, RoCE).
  • Familiarity with Kubernetes for GPU workloads and bare-metal concepts (firmware, BMC/IPMI/Redfish, OS imaging).
  • Anchor expert with leadership to uplift engineers around you.

Responsibilities

  • Serve as the GPU technical leader for the Compute team across Kubernetes, Machine Bootstrap, Networking, and Cloud.
  • Own GPU host lifecycle: driver, firmware, CUDA stack management, health, and remediation of GPU-specific failures.
  • Architect how GPU capacity is exposed to compute platforms and integrated with Kubernetes for GPUs and AI workloads.
  • Drive reliability and performance at fleet scale with automated diagnosis and repair of unhealthy accelerators.
  • Evaluate and onboard new GPU/AI accelerator platforms, networking topologies, and multi-node training/inference patterns.
  • Establish standards, tooling, and APIs for safe, efficient GPU compute usage across teams.

Skills

Go programming
Distributed systems
GPU compute

Tools

Kubernetes
CUDA
NVLink
InfiniBand
RoCE
Bare-metal concepts
BMC/IPMI/Redfish

Job description

Principal Software Engineer, GPU Compute

Roblox San Mateo, CA, United States

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators.

At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there.

A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone.

As a Principal Software Engineer on the Compute team, you will be the technical anchor for Roblox's GPU and AI accelerator capabilities. This is a battle-tested GPU expert role focused on the machine management layer and above: how GPU hosts are made production-ready, kept healthy, and turned into reliable compute for the workloads that depend on them. You will own the hard problems that show up only at scale, from driver and firmware management to GPU health, reliability, and performance across a rapidly growing fleet of accelerators spanning Roblox data centers and cloud environments. You will set the technical direction for GPU compute and up-level the entire organization's GPU expertise.

You Will
  • Serve as the GPU technical leader for the Compute team, partnering across Kubernetes, Machine Bootstrap, Networking, and Cloud to drive GPU strategy end to end.
  • Own the GPU host lifecycle above raw fleet management: driver, firmware, and CUDA stack management, GPU health and telemetry, and remediation of GPU-specific failures (XID errors, ECC, thermal, NVLink and fabric faults).
  • Architect how GPU capacity is exposed to compute platforms, including scheduling, isolation, and integration with Kubernetes for GPU and AI workloads.
  • Drive GPU reliability and performance at fleet scale, defining the detection, diagnosis, and automated repair of unhealthy accelerators before they impact production.
  • Evaluate and onboard new GPU and AI accelerator platforms, networking topologies (NVLink, InfiniBand, RoCE), and multi-node training and inference patterns.
  • Establish the standards, tooling, and APIs that let other engineering teams consume GPU compute safely and efficiently, reducing toil and raising the bar for the org.
You Have
  • 10+ years of experience building and operating large-scale distributed systems and infrastructure.
  • Deep, hands-on GPU expertise at the machine management layer and above: GPU host provisioning, driver and firmware lifecycle, GPU health and reliability, and the realities of running accelerators in production.
  • A track record as an expert for compute, not just fleet management, with the scars to prove you have scaled GPU or accelerator infrastructure that other teams depend on.
  • Strong proficiency in Go or other well-structured programming languages.
  • Experience operating GPU and AI workloads in production, including familiarity with CUDA, GPU scheduling, and high-performance networking (NVLink, InfiniBand, RoCE).
  • Familiarity with Kubernetes for GPU workloads and with bare-metal concepts (firmware, BMC/IPMI/Redfish, OS imaging) is a strong plus.
  • A history of being the anchor expert that an organization relies on for its hardest GPU and compute problems, and the leadership to up-level the engineers around you.

For roles that are based at our headquarters in San Mateo, CA: The starting base pay for this position is as shown below. The actual base pay is dependent upon a variety of job-related factors such as professional background, training, work experience, location, business needs and market demand. Therefore, in some circumstances, the actual salary could fall outside of this expected range. This pay range is subject to change and may be modified in the future. All full-time employees are also eligible for equity compensation and for benefits as described on this page.

Annual Salary Range

$345,040—$399,420 USD

Roles that are based in an office are onsite Tuesday, Wednesday, and Thursday, with optional presence on Monday and Friday (unless otherwise noted).

Roblox provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. Roblox also provides reasonable accommodations to candidates with qualifying disabilities or religious beliefs during the recruiting process.

For US based roles only, please note the Company may not be able to employ candidates for this role who have United States work authorization related to certain U.S. visa categories, or support future H-1B sponsorship at this time.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Software Engineer, GPU Compute San Mateo, CA, United States Engineering
Principal Software Engineer, GPU Compute San Mateo, CA, United States Engineering

Roblox Corporation • San Mateo (CA)

Hybrid
USD 345,000 - 400,000
Principal Software Engineer, GPU Compute
Principal Software Engineer, GPU Compute

Roblox • San Mateo (CA)

Hybrid
USD 345,000 - 400,000
Equity compensation
Comprehensive benefits package
Flexible work schedule
Senior Hardware Engineer – GPU & AI Infrastructure
Senior Hardware Engineer – GPU & AI Infrastructure

Roblox • San Mateo (CA)

On-site
USD 243,000 - 295,000
Senior Product Manager, Compute Platform
Senior Product Manager, Compute Platform

Roblox • San Mateo (CA)

Hybrid
USD 280,000 - 331,000
Equity compensation
Comprehensive benefits package
Senior Hardware Engineer - GPU & AI Infrastructure
Senior Hardware Engineer - GPU & AI Infrastructure

OneClick Smart Resume • San Mateo (CA)

On-site
USD 243,000 - 295,000
Principal Software Engineer – Compute (Kubernetes)
Principal Software Engineer – Compute (Kubernetes)

Roblox • San Mateo (CA)

On-site
USD 345,000 - 399,000
Equity compensation
Full benefits
Senior Product Manager, Compute Platform San Mateo, CA, United States
Senior Product Manager, Compute Platform San Mateo, CA, United States

Roblox Corporation • San Mateo (CA)

On-site
USD 280,000 - 331,000
Principal Software Engineer, Compute Fleet Management
Principal Software Engineer, Compute Fleet Management

Roblox • San Mateo (CA)

On-site
USD 345,000 - 400,000
Equity compensation
Employee benefits
Senior Hardware Engineer – Infrastructure
Senior Hardware Engineer – Infrastructure

Roblox • San Mateo (CA)

On-site
USD 243,000 - 295,000
Equity compensation
Benefits described on page
Principal Software Engineer, Compute Fleet Management San Mateo, CA, United States Engineering
Principal Software Engineer, Compute Fleet Management San Mateo, CA, United States Engineering

Roblox Corporation • San Mateo (CA)

Hybrid
USD 345,000 - 400,000