Senior Distributed Systems Engineer EDA Infra Equity

Socket.dev

North Carolina

Hybrid

USD 152,000 - 288,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking a Sr Software Engineer - Distributed Systems Engineer, EDA Infrastructure to design and scale automation for vast GPU/CPU compute fleets. You will partner with EDA, networking, storage, and hardware teams to improve reliability and throughput across production workloads.

You will implement provisioning, lifecycle management, and remediation tools, while enhancing observability and integration with schedulers and management systems.

Qualifications

  • 5+ years of software or infrastructure engineering for large-scale production systems.
  • BS in Computer Science, Engineering, Physics, Mathematics, or related field, or equivalent experience.
  • Strong programming in Go or Python with solid data structures, algorithms, testing, and design.
  • Experience designing automation for distributed systems and large fleets of Linux-based compute nodes.
  • Understanding of performance, security, reliability, fault tolerance, state management, and data consistency in complex systems.
  • Experience with infrastructure automation, software deployment, observability, and operational recovery.
  • Strong communication skills and ability to work across teams and geographies.
  • Systematic problem solving with ownership and focus on reducing operational toil.

Responsibilities

  • Design and build platforms that automate provisioning, configuration, operation, and lifecycle management of large-scale GPU and CPU compute infrastructure.
  • Develop monitoring, health-management, and remediation systems to improve reliability and utilization.
  • Automate hardware deployment, OS configuration, firmware/ software updates, and recovery workflows.
  • Build services that integrate with workload schedulers, infrastructure management, and observability platforms.
  • Use telemetry to identify failures and return unhealthy systems to service.
  • Collaborate with EDA, infrastructure, networking, storage, and hardware teams on scalable solutions.
  • Contribute to incident response, root-cause analysis, capacity planning, and ongoing production improvements.

Skills

Go
Python
Distributed systems
Production infrastructure
System design
Communication

Education

Bachelor's degree in Computer Science/Engineering/Physics/Math
Equivalent experience

Tools

Slurm
Kubernetes
Linux
LSF
Bright Cluster Manager

Job description

NVIDIA is seeking a Sr Software Engineer - Distributed Systems Engineer, EDA Infrastructure to design and scale automation for vast GPU/CPU compute fleets. You will partner with EDA, networking, storage, and hardware teams to improve reliability and throughput across production workloads.

You will implement provisioning, lifecycle management, and remediation tools, while enhancing observability and integration with schedulers and management systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Distributed Systems Engineer - EDA Infra
Senior Distributed Systems Engineer - EDA Infra

NVIDIA • United States

On-site
USD 152,000 - 242,000
Senior Distributed Systems Engineer, EDA Infrastructure
Senior Distributed Systems Engineer, EDA Infrastructure

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Distributed Systems Engineer – EDA Infrastructure
Senior Distributed Systems Engineer – EDA Infrastructure

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Distributed Systems Engineer, EDA Infra & Automation
Distributed Systems Engineer, EDA Infra & Automation

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior Systems Engineer: GPU Infra & Distributed Automation
Senior Systems Engineer: GPU Infra & Distributed Automation

NVIDIA AI • Durham (CA)

On-site
USD 150,000 - 210,000
Equity
Benefits
Senior Infra Engineer: Scale GPU/CPU Compute & Automation
Senior Infra Engineer: Scale GPU/CPU Compute & Automation

NVIDIA Gruppe • Washington

On-site
USD 180,000 - 288,000
Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure
Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure

NVIDIA • United States

On-site
USD 152,000 - 242,000
Senior Software Engineer - EDA Infra & System Validation
Senior Software Engineer - EDA Infra & System Validation

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Senior Systems Engineer - EDA Infra with Equity Options
Senior Systems Engineer - EDA Infra with Equity Options

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Systems Engineer - Scale Infra & Reliability Equity
Senior Systems Engineer - Scale Infra & Reliability Equity

NVIDIA • Columbia (SC)

On-site
USD 224,000 - 431,250
Equity
Benefits package