Network Systems Engineer

Supermicro

San Jose (CA)

On-site

USD 120,000 - 140,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Supermicro is seeking a Network Engineer to assist with rollouts and maintenance of business-critical applications and services. You will collaborate with senior engineers to resolve escalated issues, implement complex projects, and deliver on-site deployment services with post-support.

The role emphasizes HPC/AI workloads, system-level testing, and robust documentation. Ideal candidates hold a BS/MS in relevant engineering or CS, have 1+ years in Deep Learning/ML, and are familiar with Linux

Qualifications

  • Requires BS/MS in Electrical Engineering, Computer Engineering or Computer Science.
  • 1+ years of work-related experience in Deep Learning and Machine Learning.
  • Familiar with Linux/networking debugging/testing or relevant experience preferred.
  • Familiar with data center, enterprise, or telecommunications networking.
  • Knowledge of DevOps or cloud environments including Docker/Containers and Kubernetes.
  • Hands-on experience with workload/scheduler Managers (Slurm) for rack/cluster.
  • Familiar with MLPerf Training/Inference benchmarks, LLM, HPL-AI or RCCL/NCCL.
  • Programming experience with Windows and Linux shell scripting.
  • Strong teamwork and communication skills.

Responsibilities

  • Execute system-level rack tests on GPUs and processors, covering functionality and performance.
  • Address HPC/AI application customer issues and build robust processes.
  • Prototype test designs and optimize benchmarks for HPC/AI workloads.
  • Deliver on-site deployment services and provide post-level 1&2 support.
  • Document hardware/software issues and feedback to Product Management.
  • Contribute to HPC roadmap and software/hardware upgrades.
  • Develop test plans, reports, and automation scripts.

Skills

Deep Learning
Machine Learning
Linux debugging
Scripting
DevOps
Cloud environments
Teamwork
CUDA/oneAPI/ROCm
SLURM

Education

BS/MS in Electrical Engineering, Computer Engineering or Computer Science

Tools

Docker
Kubernetes
Slurm

Job description

Job Req ID: 30201
About Supermicro:

Supermicro is a Top Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon Valley Top 50 technology firms. Our unprecedented global expansion has provided us with the opportunity to offer a large number of new positions to the technology community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.

Job Summary:

As a Network Engineer, you will be assisting roll out and maintain business critical applications and services for Supermicro. You will work with senior engineer to resolve escalated service issues, work with other engineers to resolutions, engineering and implementing complex projects.

Essential Duties and Responsibilities:
  • Execute system-level rack tests on latest NVidia and AMD GPUs, ARM-based, Intel Xeon, and AMD EPYC processors, encompassing functionality, compatibility, performance, stress, and reliability testing, leveraging proprietary in-house tools.
  • Familiar with HPC/AI applications and benchmarks, address customer support issues, demonstrating innovative problem-solving skills and building robust processes and procedures for HPC/AI solutions.
  • May work on conduct proof of concept design and testing, providing optimized benchmarks for HPC/AI applications in a timely manner. Fine-tune BIOS settings, optimize OS/network configurations, and develop diverse simulation configurations to enhance efficiency across various workloads.
  • Deliver on-site deployment services, ensuring customer acceptance verification and providing post-level 1&2 support. Create and maintain technical documentation, including technical notes, blogs, and diagrams, to facilitate knowledge dissemination.
  • Identify and document hardware and software quality issues and collaborate with Product Management and other Engineering teams to integrate customer feedback into future product enhancements.
  • Proactively engage in HPC roadmap development, planning software and hardware upgrades to sustain exceptional HPC infrastructure performance.
  • Document and analyze test plans, reports, logs, and actively contribute to the development of test utilities and automation scripts to streamline testing processes.
Qualifications:
  • BS/MS in Electrical Engineering, Computer Engineering or Computer Science
  • 1+ years of work-related experience in Deep Learning and Machine Learning
  • Familiar with Linux/networking debugging/testing or relevant experience preferred
  • Familiar withdata center, enterprise, or telecommunication working on routing and switching networking technologies.
  • Knowledge with DevOps or in cloud environments, including but not limited to Docker/Containers and Kubernetes
  • Hands-on experience with workload/scheduler Managers (Slurm) for rack/cluster
  • Familiar with MLPerf Training/Inference benchmark, LLM, HPL-AI or RCCL/NCCL
  • Programming experience with windows and Linux shell scripting
  • Strong sense of teamwork and good team player, strong communication skills
Desired Skills:
  • Familiar with Intel/AMD/NVIDIA development tool kits such as CUDA, oneAPI, ROCm
  • Relevant certifications such as CCIE, JNCIE, or Arista ACE are highly desirable
  • Experience with server/network hardware debugging and troubleshooting
  • CCNA, OpenStack, OpenShift, Azure or AWS
Salary Range

$120,000 - $140,000

The salary offered will depend on several factors, including your location, level, education, training, specific skills, years of experience, and comparison to other employees already in this role. In addition to a comprehensive benefits package, candidates may be eligible for other forms of compensation, such as participation in bonus and equity award programs.

EEO Statement

Supermicro is an Equal Opportunity Employer and embraces diversity in our employee population. It is the policy of Supermicro to provide equal opportunity to all qualified applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, protected veteran status or special disabled veteran, marital status, pregnancy, genetic information, or any other legally protected status.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Network Systems Engineer
Network Systems Engineer

Supermicro • Wayne (CA)

On-site
USD 120,000 - 140,000
System Engineer
System Engineer

Supermicro • Wayne (CA)

On-site
USD 140,000 - 158,000
Data Center Network Engineer
Data Center Network Engineer

Supermicro • San Jose (CA)

On-site
USD 115,000 - 130,000
Sr. System Engineer
Sr. System Engineer

Supermicro • San Jose (CA)

On-site
USD 137,000 - 156,000
Technical Support Engineer
Technical Support Engineer

Supermicro • San Jose (CA)

On-site
USD 90,000 - 110,000
Data Center Network Engineer (30061)
Data Center Network Engineer (30061)

Supermicro • San Jose (CA)

On-site
USD 115,000 - 130,000
Datacenter Solutions Engineer
Datacenter Solutions Engineer

Super Micro Computer, Inc. • San Jose (CA)

On-site
USD 75,000 - 100,000
Sr. Solutions Architect
Sr. Solutions Architect

Supermicro • San Jose (CA)

On-site
USD 165,000 - 180,000
Bonus program
Equity awards
Benefits package
Network Application Engineer
Network Application Engineer

Supermicro • San Jose (CA)

On-site
USD 120,000 - 140,000
Staff Software Engineer - Network Triage and Test Automation
Staff Software Engineer - Network Triage and Test Automation

Supermicro • San Jose (CA)

On-site
USD 200,000 - 255,000