Senior Distributed Systems Engineer, EDA Infra - Equity

NVIDIA

Westford (MA)

On-site

USD 184,000 - 288,000

Full time

13 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is hiring engineers to design and operate the infrastructure that supports our EDA workloads. You will build automated platforms and services that manage huge GPU/CPU compute fleets used by engineering teams across NVIDIA.

The role focuses on distributed systems, automation, monitoring, and reliability, with cross-team collaboration and scalable solutions for chip-design workloads. Equity and benefits are provided.

Qualifications

  • 5+ years of software engineering or infrastructure engineering experience supporting large-scale production systems.
  • BS in Computer Science, Engineering, Physics, Mathematics, or related field, or equivalent experience.
  • Strong programming experience in Go or Python, with data structures, algorithms, testing and design knowledge.
  • Experience designing automation for distributed systems and large fleets of Linux-based compute nodes.
  • Understanding of performance, security, reliability, fault tolerance, state management and data consistency.
  • Experience with infrastructure automation, software deployment, observability, and operational recovery.
  • Strong communication skills and ability to work across teams, organizations, and regions.
  • A systematic, ownership-driven approach to problem solving and reducing toil.

Responsibilities

  • Design and build platforms that automate provisioning, configuration, operation, and lifecycle management of large-scale GPU and CPU compute infrastructure.
  • Develop monitoring, health-management, and remediation systems to improve reliability, availability, and utilization of EDA compute environments.
  • Automate hardware deployment, OS configuration, firmware and software updates, cluster enrollment, and recovery workflows.
  • Build reliable services and workflows that integrate with workload schedulers, infrastructure management systems, and observability platforms.
  • Use diagnostics, OS signals, scheduler data, and network/storage telemetry to identify failures and return unhealthy systems to service.
  • Collaborate with EDA, infrastructure, networking, and hardware teams to deliver scalable chip-design solutions.
  • Participate in incident response, root-cause analysis, capacity planning, and continuous improvement of production services.

Skills

Go
Python
Distributed systems
Automation
Linux
Observability

Education

BS in Computer Science or related field

Tools

Slurm
LSF
Kubernetes
Bright Cluster Manager

Job description

NVIDIA is hiring engineers to design and operate the infrastructure that supports our EDA workloads. You will build automated platforms and services that manage huge GPU/CPU compute fleets used by engineering teams across NVIDIA.

The role focuses on distributed systems, automation, monitoring, and reliability, with cross-team collaboration and scalable solutions for chip-design workloads. Equity and benefits are provided.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Distributed Systems Engineer – EDA Infra (Equity)
Senior Distributed Systems Engineer – EDA Infra (Equity)

NVIDIA • California (MO)

On-site
USD 170,000 - 260,000
Equity
Benefits
Senior Distributed Systems Engineer — EDA Infra (Equity)
Senior Distributed Systems Engineer — EDA Infra (Equity)

NVIDIA • Austin (TX)

On-site
USD 152,000 - 288,000
Senior Distributed Systems Engineer — EDA Infra (Equity)
Senior Distributed Systems Engineer — EDA Infra (Equity)

NVIDIA Corporation • Washington

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Senior Distributed Systems Engineer, EDA Infrastructure
Senior Distributed Systems Engineer, EDA Infrastructure

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer, Distributed Systems for EDA Infra
Senior Software Engineer, Distributed Systems for EDA Infra

NVIDIA • Durham (NC)

On-site
USD 152,000 - 288,000
Equity
Benefits
Distributed Systems Engineer, EDA Infra & Automation
Distributed Systems Engineer, EDA Infra & Automation

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior Systems & Infra Automation Engineer (Equity Eligible)
Senior Systems & Infra Automation Engineer (Equity Eligible)

NVIDIA • Austin (TX)

On-site
USD 184,000 - 357,000
Equity
Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure
Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Infra Engineer: Scale GPU/CPU Compute & Automation
Senior Infra Engineer: Scale GPU/CPU Compute & Automation

NVIDIA Gruppe • Washington

On-site
USD 180,000 - 288,000
Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure
Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure

NVIDIA Gruppe • Washington

On-site
USD 180,000 - 288,000