HPC/AI Network Engineer — Equity Eligible

Supermicro

San Jose (CA)

On-site

USD 120,000 - 140,000

Full time

34 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Supermicro in San Jose, CA seeks a Network Engineer to assist roll out and maintain business-critical applications and services. You will work with a senior engineer to resolve escalated issues and implement complex projects.

Candidate will test hardware and software, document procedures, and support HPC/AI workloads, leveraging Linux, Docker, and Kubernetes, with a focus on reliability and performance.

Qualifications

  • BS/MS in Electrical Engineering, Computer Engineering or Computer Science.
  • 1+ years of work-related experience in Deep Learning and Machine Learning.
  • Familiar with Linux networking debugging/testing or relevant experience.
  • Familiar with data center, enterprise, or telecommunication working on routing and switching networking technologies.
  • Knowledge with DevOps or in cloud environments, including Docker/Containers and Kubernetes.
  • Hands-on experience with workload/scheduler Managers (Slurm) for rack/cluster.
  • Familiar with MLPerf Training/Inference benchmark, LLM, HPL-AI or RCCL/NCCL.
  • Programming experience with Windows and Linux shell scripting.
  • Strong sense of teamwork and good communication skills.

Responsibilities

  • Execute system-level rack tests on latest NVidia and AMD GPUs, ARM-based, Intel Xeon, and AMD EPYC processors, encompassing functionality, compatibility, performance, stress, and reliability testing, leveraging proprietary in-house tools.
  • Familiar with HPC/AI applications and benchmarks, address customer support issues, demonstrating innovative problem-solving skills and building robust processes and procedures for HPC/AI solutions.
  • May work on conduct proof of concept design and testing, providing optimized benchmarks for HPC/AI applications in a timely manner. Fine-tune BIOS settings, optimize OS/network configurations, and develop diverse simulation configurations to enhance efficiency across various workloads.
  • Deliver on-site deployment services, ensuring customer acceptance verification and providing post-level 1&2 support. Create and maintain technical documentation, including technical notes, blogs, and diagrams, to facilitate knowledge dissemination.
  • Identify and document hardware and software quality issues and collaborate with Product Management and other Engineering teams to integrate customer feedback into future product enhancements.
  • Proactively engage in HPC roadmap development, planning software and hardware upgrades to sustain exceptional HPC infrastructure performance.
  • Document and analyze test plans, reports, logs, and actively contribute to the development of test utilities and automation scripts to streamline testing processes.

Skills

Deep Learning
Linux Networking
Data Center Networking
DevOps / Cloud
Slurm workload Manager
Scripting
Teamwork/Communication

Education

BS/MS in Electrical/Computer Engineering or CS

Tools

Docker
Kubernetes
Linux
OpenStack/OpenShift
CUDA/oneAPI/ROCm

Job description

Supermicro in San Jose, CA seeks a Network Engineer to assist roll out and maintain business-critical applications and services. You will work with a senior engineer to resolve escalated issues and implement complex projects.

Candidate will test hardware and software, document procedures, and support HPC/AI workloads, leveraging Linux, Docker, and Kubernetes, with a focus on reliability and performance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Network & HPC Systems Engineer
Senior Network & HPC Systems Engineer

Super Micro Computer, Inc. • San Jose (CA)

On-site
USD 120,000 - 140,000
HPC Network Systems Engineer for AI & Data Center
HPC Network Systems Engineer for AI & Data Center

Supermicro • Wayne (CA)

On-site
USD 120,000 - 140,000
HPC & AI Network Engineer – GPU & Cloud
HPC & AI Network Engineer – GPU & Cloud

Support Revolution • San Jose (CA)

On-site
USD 120,000 - 140,000
Network Systems Engineer
Network Systems Engineer

Support Revolution • San Jose (CA)

On-site
USD 120,000 - 140,000
Senior System Engineer, HPC/AI & Cloud Clusters
Senior System Engineer, HPC/AI & Cloud Clusters

Supermicro • San Jose (CA)

On-site
USD 137,000 - 156,000
Network Systems Engineer
Network Systems Engineer

Supermicro • San Jose (CA)

On-site
USD 120,000 - 140,000
Network Systems Engineer
Network Systems Engineer

Supermicro • Wayne (CA)

On-site
USD 120,000 - 140,000
Network Systems Engineer 3
Network Systems Engineer 3

Super Micro Computer, Inc. • San Jose (CA)

On-site
USD 120,000 - 140,000
Engineering Technician - HPC Server Validation & Support
Engineering Technician - HPC Server Validation & Support

Supermicro • San Jose (CA)

On-site
USD 39,000 - 42,000
Senior System Engineer – HPC/AI & Cloud Deployments
Senior System Engineer – HPC/AI & Cloud Deployments

Support Revolution • San Jose (CA), Northern (KY)

Hybrid
USD 137,000 - 156,000