Network Engineer, Platform, Automation & HPC/AI

Lawrence Berkeley National Laboratory

Berkeley (CA)

On-site

USD 120,000 - 180,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lawrence Berkeley National Laboratory's NERSC is seeking a Network Engineer, Platform, Automation & HPC/AI to advance the 1 Tb/s border network and an 800G/400G data center backbone supporting HPC workloads and a wide user base across scientific computing.

This role spans network engineering, automation, and software development, offering opportunities to architect solutions, improve latency, and contribute to modernizing data centers and edge services within a high-performance research

Responsibilities

  • Implement, operate, maintain, and improve network automation and observability solutions.
  • Contribute to Data Center modernization efforts and NERSC's Smart Facility initiative.
  • Support design and delivery of network services for emerging needs (e.g., American Science Cloud).
  • Continuously monitor and optimize network performance, focusing on latency, throughput, and fault tolerance.
  • Create and maintain comprehensive network documentation, including topology diagrams.
  • Collaborate with the Security Group to ensure data integrity and high availability.
  • Share on-call rotation with colleagues and serve as escalation contact.
  • Work on and resolve complex issues requiring in-depth analysis across variables.
  • Exercise judgment in selecting methods and evaluation criteria for obtaining results.
  • Build effective working relationships with technical partners across disciplines.
  • Architect, develop, and establish technical direction for network automation and self-healing capabilities.
  • Lead Data Center modernization efforts in support of Smart Facility and emerging needs.
  • Build and integrate AI/ML-driven observability and predictive analytics.

Job description

Join NERSC at Berkeley Lab and help engineer the high-performance network platform powering some of the nation's most advanced scientific computing. As a Network Engineer, Platform, Automation & HPC/AI, you'll work at the intersection of network engineering, automation, and software development while helping advance AI-driven network operations. Our team manages 1 Tb/s of border connectivity to ESnet and an 800G/400G data center network backbone supporting the NERSC-9 and NERSC-10 supercomputers, multi-tier storage, archive, and edge services. Your work will help improve the performance, scalability, automation, and reliability of scientific workflows serving more than 10,000 users.

Our mission is to bring science solutions to the world. We welcome candidates from all backgrounds, including those with non-traditional paths. We value a growth mindset and believe skills and experience are transferable. If you're eager to learn and meet the qualifications below, we encourage you to apply. Join our team where your work can have a high impact for an organization associated with 17 Nobel Prizes... and counting.

This position may be filled at Level 3 or Level 4. Level 3 is intended for experienced engineers who independently solve complex networking and automation challenges. Level 4 is intended for senior technical leaders who architect solutions, lead modernization efforts, and establish new technical approaches for complex HPC and data center environments.

Network Engineer Level 3 will:
  • Implement, operate, maintain, and improve network automation and observability solutions.
  • Contribute to Data Center modernization efforts and NERSC's Smart Facility initiative.
  • Support efforts to design and deliver network services to address emerging needs (e.g., American Science Cloud, new Edge services).
  • Continuously monitor and optimize network performance, focusing on reducing latency, maximizing throughput, and improving fault tolerance.
  • Create and maintain comprehensive network documentation, including physical and logical topology diagrams.
  • Collaborate with the Security Group to implement security measures for data integrity and privacy, ensuring high availability and reliability through redundancy and failover mechanisms.
  • Share on-call rotation with colleagues and serve as an escalation contact for service incidents.
  • Work on and resolve complex issues where analysis of situations or data requires an in-depth evaluation of multiple variables.
  • Exercise judgment in selecting methods, techniques and evaluation criteria for obtaining results.
  • Build effective working relationships with technical partners and stakeholders across disciplines.
In addition to the above, the Senior Network Engineer Level 4 will:
  • Architect, develop, and establish technical direction for network automation, observability, and self-healing capabilities.
  • Lead Data Center modernization efforts in support of NERSC's Smart Facility initiative and emerging needs e.g., American Science Cloud Design, develop, and maintain automation frameworks, infrastructure-as-code, and software solutions to manage, optimize, and self-heal the HPC and data center network.
  • Build and integrate AI/ML-driven observability, predictive analytics
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Network Engineer - Platform, Automation & AI
HPC Network Engineer - Platform, Automation & AI

LBL • Berkeley (CA)

Hybrid
USD 157,000 - 218,000
Tuition assistance
Holiday shutdown
Parental leave
+1
Network Engineer, Platform, Automation & HPC/AI
Network Engineer, Platform, Automation & HPC/AI

Berkeley Lab • Berkeley (CA)

Hybrid
USD 157,000 - 218,000
Health benefits
Tuition Assistance Program
Hybrid work option
+1
Network Engineer, Platform, Automation & HPC/AI
Network Engineer, Platform, Automation & HPC/AI

LBL • Berkeley (CA)

Hybrid
USD 157,000 - 218,000
Tuition assistance
Holiday shutdown
Parental leave
+1
Network Engineer, HPC & AI-Driven Automation
Network Engineer, HPC & AI-Driven Automation

Berkeley Lab • Berkeley (CA)

Hybrid
USD 157,000 - 218,000
Health benefits
Tuition Assistance Program
Hybrid work option
+1
HPC Network Engineer: Automation, AI-Driven Observability
HPC Network Engineer: Automation, AI-Driven Observability

Lawrence Berkeley National Laboratory • Berkeley (CA)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

ADP, Inc. • Berkeley (CA)

On-site
USD 140,000 - 180,000
System Infrastructure / Platform Engineer, HPC Technology Department
System Infrastructure / Platform Engineer, HPC Technology Department

Berkeley Lab • Berkeley (CA)

Hybrid
USD 156,000 - 192,000
Exceptional health benefits
Tuition Assistance Program
Paid vacation and sick time
+2
HPC Scientific Support Engineer
HPC Scientific Support Engineer

Lawrence Berkeley National Laboratory • Berkeley (CA)

On-site
USD 156,000 - 192,000
Site Reliability Engineer
Site Reliability Engineer

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
HPC/AI Performance Specialist
HPC/AI Performance Specialist

Berkeley Lab • Berkeley (CA)

Hybrid
USD 139,000 - 236,000
Exceptional health and retirement benefits
Tuition Assistance Program
Pet insurance
+2