Senior Distributed Systems Engineer, EDA Infrastructure

Nvidia Corporation

Santa Clara (CA)

On-site

USD 152,000 - 288,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Sr Software Engineer to design and scale the infrastructure supporting our EDA workloads. You will build platforms to automate provisioning, configuration, and lifecycle management for large-scale GPU and CPU compute fleets.

The role emphasizes reliability, observability, and automation, requiring strong Go or Python skills and cross-team collaboration. You will work across Linux, OS, networking, and hardware domains to deliver production-grade services at scale.

Qualifications

  • 5+ years of software or infrastructure engineering experience supporting large-scale production systems.
  • BS in Computer Science, Engineering, Physics, Mathematics, or related field, or equivalent experience.
  • Strong programming in Go or Python with solid data structures and algorithms.
  • Experience designing automation for distributed systems and large fleets of Linux-based compute nodes.
  • Understanding of performance, security, reliability, fault tolerance, and data consistency in complex systems.
  • Experience with infrastructure automation, software deployment, observability, and operational recovery.
  • Strong communication across teams, organizations, and geographic regions.
  • Systematic problem solving with ownership and focus on reducing toil.

Responsibilities

  • Design and build platforms that automate provisioning, configuration, operation, and lifecycle management of large-scale GPU/CPU infrastructure.
  • Develop monitoring, health-management, and remediation systems to improve reliability and utilization.
  • Automate hardware deployment, OS configuration, firmware/software updates, and recovery workflows.
  • Build services and workflows that integrate with schedulers and observability platforms.
  • Use telemetry to identify failures and return unhealthy systems to service.
  • Collaborate with EDA, infra, networking, storage, and hardware teams to deliver scalable solutions.
  • Participate in incident response, root-cause analysis, and capacity planning.

Skills

Go
Python
Distributed systems
Automation
Linux
Cluster management
Data structures
Algorithms
Observability

Education

Bachelor's degree in Computer Science, Engineering, Physics, Mathematics, or related field

Tools

Slurm
LSF
Kubernetes
Bright Cluster Manager

Job description

NVIDIA is seeking a Sr Software Engineer to design and scale the infrastructure supporting our EDA workloads. You will build platforms to automate provisioning, configuration, and lifecycle management for large-scale GPU and CPU compute fleets.

The role emphasizes reliability, observability, and automation, requiring strong Go or Python skills and cross-team collaboration. You will work across Linux, OS, networking, and hardware domains to deliver production-grade services at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Distributed Systems Engineer - EDA Infra
Senior Distributed Systems Engineer - EDA Infra

NVIDIA • United States

On-site
USD 152,000 - 242,000
Senior Distributed Systems Engineer EDA Infra Equity
Senior Distributed Systems Engineer EDA Infra Equity

Socket.dev • North Carolina

Hybrid
USD 152,000 - 288,000
Senior Distributed Systems Engineer – EDA Infrastructure
Senior Distributed Systems Engineer – EDA Infrastructure

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Distributed Systems Engineer, EDA Infra & Automation
Distributed Systems Engineer, EDA Infra & Automation

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior Systems Engineer: GPU Infra & Distributed Automation
Senior Systems Engineer: GPU Infra & Distributed Automation

NVIDIA AI • Durham (CA)

On-site
USD 150,000 - 210,000
Equity
Benefits
Senior Infra Engineer: Scale GPU/CPU Compute & Automation
Senior Infra Engineer: Scale GPU/CPU Compute & Automation

NVIDIA Gruppe • Washington

On-site
USD 180,000 - 288,000
Senior Systems Software Engineer- EDA Infrastructure
Senior Systems Software Engineer- EDA Infrastructure

NVIDIA AI • Durham (CA)

On-site
USD 150,000 - 210,000
Equity
Benefits
Senior Systems Engineer - EDA Infra with Equity Options
Senior Systems Engineer - EDA Infra with Equity Options

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure
Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure

NVIDIA • United States

On-site
USD 152,000 - 242,000
Senior Infrastructure & Systems Engineer
Senior Infrastructure & Systems Engineer

NVIDIA Corporation • Northern (KY)

Hybrid
USD 184,000 - 357,000