Software Architect, GPU System

ByteDance

San Jose (CA)

On-site

USD 218,000 - 480,000

Full time

44 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

ByteDance seeks a senior technical leader to drive the GPU system software strategy for data-center compute. You will define architecture and a multi-year roadmap across driver, runtime, firmware, and observability layers, guiding cross-functional teams from silicon to fleet operations.

You will ensure high-quality design reviews, performance analysis, and robust deployment while mentoring engineers and communicating decisions to executives and specialists.

Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
  • 5+ years of hands-on experience in GPU, accelerator, kernel, firmware, runtime, or data-center system software, including technical leadership on complex hardware/software programs.
  • Demonstrated success setting technical direction and leading a major hardware/software program from concept through production.
  • Deep understanding of accelerator architecture, memory hierarchy, interconnects, operating systems, device management, and production reliability.
  • Experience coordinating dependencies and technical decisions across multiple engineering teams and organizational boundaries.
  • Strong coding, design-review, performance-analysis, and system-debugging capability, with evidence of continued hands-on contribution.

Responsibilities

  • Define the architecture and multi-year technical roadmap for GPU system software across driver, runtime, firmware-interface, management, observability, and reliability layers.
  • Set priorities and technical standards for accelerator enablement, balancing near-term product commitments with compatibility, performance, serviceability, and long-term maintenance.
  • Own end-to-end integration from architecture and pre-silicon planning through bring-up, qualification, general availability, fleet monitoring, and sustained operation.
  • Lead cross-functional execution among silicon, firmware, kernel, compiler, library, machine-learning framework, server, network, storage, validation, and production teams.
  • Resolve ambiguous system-level tradeoffs involving APIs, resource management, memory and interconnect topology, telemetry, recovery, security, and workload performance.
  • Remain technically involved through prototypes, critical-path code and design reviews, performance analysis, and leadership during high-severity failure investigations.
  • Define measurable release and reliability criteria, including qualification coverage, regression thresholds, fault containment, automated repair, and fleet-health indicators.
  • Mentor engineers, raise the quality of architecture and debugging practices, and communicate technical decisions to both specialist and executive audiences.

Skills

GPU software
kernel
firmware
runtime
leadership
design-review
performance-analysis
system-debugging

Education

Bachelor’s degree in CS/CE/EE

Tools

GPU driver
runtime libraries
fleet interfaces

Job description

About the Team

We are a systems software team building the foundational software for large-scale compute platforms. We work at the hardware/software boundary across the Linux kernel, accelerators, storage, firmware, and platform validation. We value rigorous engineering, clear interfaces, measurable performance and reliability, and upstream collaboration where appropriate. The team partners closely with hardware, architecture, product, validation, and production engineering groups to move new capabilities from design through dependable deployment.

About the Role

You will lead the technical strategy and cross-functional delivery of the low-level GPU software stack for data-center computing. You will connect accelerator architecture, firmware, kernel drivers, runtimes, communication, telemetry, and fleet reliability into a coherent roadmap. This is a hands-on individual-contributor role: you will make architecture decisions, guide engineers, review critical implementations, and resolve system-level issues without assuming hiring or performance-management responsibilities.

Responsibilities
  • Define the architecture and multi-year technical roadmap for GPU system software across driver, runtime, firmware-interface, management, observability, and reliability layers.
  • Set priorities and technical standards for accelerator enablement, balancing near-term product commitments with compatibility, performance, serviceability, and long-term maintenance.
  • Own end-to-end integration from architecture and pre-silicon planning through bring-up, qualification, general availability, fleet monitoring, and sustained operation.
  • Lead cross-functional execution among silicon, firmware, kernel, compiler, library, machine-learning framework, server, network, storage, validation, and production teams.
  • Resolve ambiguous system-level tradeoffs involving APIs, resource management, memory and interconnect topology, telemetry, recovery, security, and workload performance.
  • Remain technically involved through prototypes, critical-path code and design reviews, performance analysis, and leadership during high-severity failure investigations.
  • Define measurable release and reliability criteria, including qualification coverage, regression thresholds, fault containment, automated repair, and fleet-health indicators.
  • Mentor engineers, raise the quality of architecture and debugging practices, and communicate technical decisions to both specialist and executive audiences.
Qualifications
Minimum Qualifications
  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
  • 5+ years of hands-on experience in GPU, accelerator, kernel, firmware, runtime, or data-center system software, including technical leadership on complex hardware/software programs.
  • Demonstrated success setting technical direction and leading a major hardware/software program from concept through production.
  • Deep understanding of accelerator architecture, memory hierarchy, interconnects, operating systems, device management, and production reliability.
  • Experience coordinating dependencies and technical decisions across multiple engineering teams and organizational boundaries.
  • Strong coding, design-review, performance-analysis, and system-debugging capability, with evidence of continued hands-on contribution.
Preferred Qualifications
  • Experience with a major GPU computing stack and its kernel driver, runtime, libraries, tooling, and fleet-management interfaces.
  • Experience defining platform APIs or hardware/software contracts across multiple accelerator generations.
  • Knowledge of distributed accelerator workloads, collective communication, topology-aware placement, and large-scale serviceability.
  • Track record mentoring senior engineers and building alignment where priorities, ownership, or technical evidence initially conflict.
Job Information

The base salary range for this position in the selected city is $218400 - $480000 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

For Los Angeles County (unincorporated) Candidates

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:

1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;

2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems;

3. Exercising sound judgment.

About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Reasonable Accommodation

ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, GPU Systems
Senior Software Engineer, GPU Systems

ByteDance • San Jose (CA)

On-site
USD 162,000 - 388,000
Medical, dental, vision insurance
401(k) plan with company match
Paid parental leave
+1
Senior Software Engineer, Storage Systems
Senior Software Engineer, Storage Systems

ByteDance • San Jose (CA)

On-site
USD 162,000 - 388,000
System Software Architect, Linux Kernel and Operating System
System Software Architect, Linux Kernel and Operating System

ByteDance • San Jose (CA)

On-site
USD 218,000 - 480,000
Medical insurance
Dental insurance
Vision insurance
+8
Tech Lead Engineer, Physical AI Infrastructure
Tech Lead Engineer, Physical AI Infrastructure

ByteDance • San Jose (CA)

On-site
USD 254,000 - 480,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
Senior Production System Engineer - New York City
Senior Production System Engineer - New York City

Bytedance • New York (NY)

On-site
USD 122,000 - 272,000
Medical, dental and vision insurance
401(k) with company match
Paid parental leave
+1
Backend Development Engineer Graduate (Infrastructure Platform Delivery) - 2027 Start
Backend Development Engineer Graduate (Infrastructure Platform Delivery) - 2027 Start

Bytedance • San Jose (CA)

On-site
USD 128,000 - 256,000
Cloud Site Reliability Engineer - DCS Cloud Seattle Regular
Cloud Site Reliability Engineer - DCS Cloud Seattle Regular

ByteDance • Seattle (WA), Northern (KY)

Hybrid
USD 129,960 - 341,734
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+5
Senior Server Operations & Maintenance Engineer Technology - Infrastructure San Jose Regular
Senior Server Operations & Maintenance Engineer Technology - Infrastructure San Jose Regular

ByteDance • San Jose (CA), Northern (KY)

Hybrid
USD 115,200 - 288,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Research Scientist - DPU & AI Infra Technology - Infrastructure San Jose Regular
Research Scientist - DPU & AI Infra Technology - Infrastructure San Jose Regular

ByteDance • San Jose (CA), Northern (KY)

Hybrid
USD 136,800 - 359,720
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Senior Software Engineer, AI Infrastructure-Compute Efficiency
Senior Software Engineer, AI Infrastructure-Compute Efficiency

ByteDance • Seattle (WA)

On-site
USD 207,000 - 368,000