HPC Network Engineer: Automation, AI-Driven Observability

Lawrence Berkeley National Laboratory

Berkeley (CA)

On-site

USD 120,000 - 180,000

Full time

13 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Lawrence Berkeley National Laboratory's NERSC is seeking a Network Engineer, Platform, Automation & HPC/AI to advance the 1 Tb/s border network and an 800G/400G data center backbone supporting HPC workloads and a wide user base across scientific computing.

This role spans network engineering, automation, and software development, offering opportunities to architect solutions, improve latency, and contribute to modernizing data centers and edge services within a high-performance research

Responsibilities

  • Implement, operate, maintain, and improve network automation and observability solutions.
  • Contribute to Data Center modernization efforts and NERSC's Smart Facility initiative.
  • Support design and delivery of network services for emerging needs (e.g., American Science Cloud).
  • Continuously monitor and optimize network performance, focusing on latency, throughput, and fault tolerance.
  • Create and maintain comprehensive network documentation, including topology diagrams.
  • Collaborate with the Security Group to ensure data integrity and high availability.
  • Share on-call rotation with colleagues and serve as escalation contact.
  • Work on and resolve complex issues requiring in-depth analysis across variables.
  • Exercise judgment in selecting methods and evaluation criteria for obtaining results.
  • Build effective working relationships with technical partners across disciplines.
  • Architect, develop, and establish technical direction for network automation and self-healing capabilities.
  • Lead Data Center modernization efforts in support of Smart Facility and emerging needs.
  • Build and integrate AI/ML-driven observability and predictive analytics.

Job description

Lawrence Berkeley National Laboratory's NERSC is seeking a Network Engineer, Platform, Automation & HPC/AI to advance the 1 Tb/s border network and an 800G/400G data center backbone supporting HPC workloads and a wide user base across scientific computing.

This role spans network engineering, automation, and software development, offering opportunities to architect solutions, improve latency, and contribute to modernizing data centers and edge services within a high-performance research

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Engineer, HPC & AI-Driven Automation
Network Engineer, HPC & AI-Driven Automation

Berkeley Lab • Berkeley (CA)

Hybrid
USD 157,000 - 218,000
Health benefits
Tuition Assistance Program
Hybrid work option
+1
HPC Network Engineer - Platform, Automation & AI
HPC Network Engineer - Platform, Automation & AI

LBL • Berkeley (CA)

Hybrid
USD 157,000 - 218,000
Tuition assistance
Holiday shutdown
Parental leave
+1
Network Engineer, Platform, Automation & HPC/AI
Network Engineer, Platform, Automation & HPC/AI

Lawrence Berkeley National Laboratory • Berkeley (CA)

On-site
USD 120,000 - 180,000
HPC User Success Engineer
HPC User Success Engineer

Lawrence Berkeley National Laboratory • Berkeley (CA)

On-site
USD 156,000 - 192,000
Network Engineer, Platform, Automation & HPC/AI
Network Engineer, Platform, Automation & HPC/AI

LBL • Berkeley (CA)

Hybrid
USD 157,000 - 218,000
Tuition assistance
Holiday shutdown
Parental leave
+1
Network Engineer, Platform, Automation & HPC/AI
Network Engineer, Platform, Automation & HPC/AI

Berkeley Lab • Berkeley (CA)

Hybrid
USD 157,000 - 218,000
Health benefits
Tuition Assistance Program
Hybrid work option
+1
IP Backbone Engineer — High‑Perf Networking & Automation
IP Backbone Engineer — High‑Perf Networking & Automation

Berkeley Lab • Berkeley (CA)

On-site
USD 131,000 - 192,000
Health and retirement benefits
Belonging culture
Vacation and sick time
+2
HPC & AI Performance Engineer for Next‑Gen Systems
HPC & AI Performance Engineer for Next‑Gen Systems

Berkeley Lab • Berkeley (CA)

Hybrid
USD 139,000 - 236,000
Exceptional health and retirement benefits
Tuition Assistance Program
Pet insurance
+2
HPC Scientific Support Engineer | User-Focused & AI-Aware
HPC Scientific Support Engineer | User-Focused & AI-Aware

Berkeley Lab • Berkeley (CA)

Hybrid
USD 156,000 - 219,000
Global HPC Network Engineer for AI Infra
Global HPC Network Engineer for AI Infra

Together • San Francisco (CA)

On-site
USD 190,000 - 280,000
Startup equity
Health insurance
Competitive benefits