Senior Deep Learning Systems Engineer, Datacenters

NVIDIA

Santa Clara (CA)

On-site

USD 184,000 - 287,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is looking for an experienced professional to develop software infrastructure and improve datacenter architectures designed for deep learning applications. This role involves collaboration with experts to analyze system characteristics and develop profiling tools in Python, bash, and C++.

The ideal candidate will have a Bachelor’s degree in Electrical Engineering or Computer Science and 8 years of experience in relevant fields. Familiarity with CUDA, PyTorch, and performance analysis is essential.

Qualifications

  • 8 years or more of relevant experience.
  • Hands-on experience in system architecture and performance analysis.
  • Experience with GPU kernels (CUDA) and deep learning frameworks.

Responsibilities

  • Develop software infrastructure for DL application analysis.
  • Create cost-efficient datacenter architectures for large language models.
  • Develop analysis tools for measuring key performance metrics.

Skills

Deep Learning applications
System Software
C/C++ programming
Python programming
Operating Systems (Linux)
Containerization Platforms (Docker)

Education

Bachelor’s degree in Electrical Engineering or Computer Science
Master's or PhD degree

Tools

CUDA
PyTorch
TensorFlow
GPU profiling tools
Perf

Job description

As NVIDIA makes inroads into the Datacenter business, our team plays a central role in getting the most out of our exponentially growing datacenter deployments as well as establishing a data‑driven approach to hardware design and system software development.

Do you want to influence the development of high-performance Datacenters designed for the future of AI? Do you have an interest in system architecture and performance? In this role you will find how CPU, GPU, networking, and IO relate to deep learning (DL) architectures for Natural Language Processing, Computer Vision, Autonomous Driving and other technologies. Come join our team, and bring your interests to help us optimize our next generation systems and Deep Learning Software Stack.

What you’ll be doing:
  • Help develop software infrastructure to characterize and analyze a broad range Deep Learning applications

  • Evolve cost-efficient datacenter architectures tailored to meet the needs of Large Language Models (LLMs).

  • Work with experts to help develop analysis and profiling tools in Python, bash and C++ to measure key performance metrics of DL workloads running on Nvidia systems.

  • Analyze system and software characteristics of DL applications.

  • Develop analysis tools and methodologies to measure key performance metrics and to estimate potential for efficiency improvement.

What we need to see:
  • A Bachelor’s degree in Electrical Engineering or Computer Science or equivalent experience (Masters or PhD degree preferred).

  • 8 years or more of relevant experience.

  • Experience in at least one of the following:

    • System Software: Operating Systems (Linux), Compilers, GPU kernels (CUDA), DL Frameworks (PyTorch, TensorFlow).

    • Silicon Architecture and Performance Modeling/Analysis: CPU, GPU, Memory or Network Architecture

  • Experience programming in C/C++ and Python. Exposure to Containerization Platforms (docker) and Datacenter Workload Managers (slurm) is a plus.

  • A deep understanding of computer system architecture and performance analysis is essential for success in this role. Applicants should have demonstrated hands‑on experience in these domains.

  • Demonstrated ability to work in virtual environments, and a strong drive to own tasks from beginning to end. Prior experience with such environments will make you stand out.

Ways to stand out from the crowd:
  • Background with system software, Operating system intrinsics, GPU kernels (CUDA), or DL Frameworks (PyTorch, TensorFlow).

  • Experience with silicon performance monitoring or profiling tools (e.g. perf, gprof, nvidia-smi, dcgm).

  • In depth performance modeling experience in any one of CPU, GPU, Memory or Network Architecture.

  • Exposure to Containerization Platforms (docker) and Datacenter Workload Managers (slurm).

  • Prior experience with multi‑site teams or multi‑functional teams.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until May 11, 2026.

This posting is for an existing vacancy.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Deep Learning Systems Engineer, Datacenters
Senior Deep Learning Systems Engineer, Datacenters

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Data Center System Architect
Senior Data Center System Architect

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Architect - Data Center Systems
Senior Software Architect - Data Center Systems

NVIDIA • Hillsboro (OR)

On-site
USD 224,000 - 357,000
Equity and benefits
Senior Software Architect - Data Center Systems
Senior Software Architect - Data Center Systems

NVIDIA • Town of Texas (WI)

On-site
USD 224,000 - 357,000
Data Center GPU Performance Engineer – Product
Data Center GPU Performance Engineer – Product

NVIDIA • Santa Clara (CA)

On-site
USD 148,000 - 225,000
Comprehensive benefits package
Equity options
Senior Deep Learning Sofware Infrastructure Engineer
Senior Deep Learning Sofware Infrastructure Engineer

2100 NVIDIA USA • California (MO)

On-site
USD 272,000 - 431,000
Senior Deep Learning Sofware Infrastructure Engineer
Senior Deep Learning Sofware Infrastructure Engineer

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Senior Deep Learning Sofware Infrastructure Engineer
Senior Deep Learning Sofware Infrastructure Engineer

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Senior Software Architect - Data Center Systems
Senior Software Architect - Data Center Systems

NVIDIA • California (MO)

On-site
USD 224,000 - 357,000
Senior Software Engineer, CUDA Deep Learning Systems
Senior Software Engineer, CUDA Deep Learning Systems

NVIDIA • Austin (TX)

On-site
USD 224,000 - 357,000
Equity
Benefits