Job Summary
The Network Operations Engineer is responsible for the daily operation and management of the company's server cluster, ensuring the stable operation of the company's business, as well as the maintenance and security management of the office network.
Key Responsibilities
- Responsible for daily operations and maintenance of the company's data center, office network, and production network environments.
- Manage configuration, monitoring, and troubleshooting of network devices including switches, routers, firewalls, and load balancers.
- Participate in network deployment, expansion, optimization, and change implementation.
- Troubleshoot and resolve network issues such as connectivity, routing, packet loss, and latency problems.
- Ensures table operation of IDC, GPU cluster, and AI/HPC network environments.
- Maintain and optimise network monitoring and alerting systems.
- Develop operational standards, SOPs, incident response procedures, and technical documentation.
- Collaborate with server, storage, system, and cloud platform teams for troubleshooting and integration support.
- Support and maintain inter-DC network solutions including leased lines, BGP, VPN, VXLAN, and EVPN.
- Participate in major incident reviews and continuous improvement initiatives.
Required Qualifications
- Bachelor's degree or above in Computer Science, Telecommunications, Network Engineering, or a related field.
- At least 3 years of experience in network operations and maintenance.
- Strong understanding of TCP/IP networking protocols.
- Familiar with Linux administration and common network troubleshooting tools.
- Strong analytical and troubleshooting skills.
- Good communication, teamwork, and ownership mindset.
- Willing to participate in on-call rotation, maintenance windows, and emergency incident response; willing to accept short-term business trips.
Technical Skills
Networking Fundamentals
Familiar with the following protocols and technologies:
- VLAN / STP / LACP
- OSPF / BGP / ECMP
- VXLAN / EVPN
- NAT / ACL
- DHCP / DNS / NTP
- IPv4 / IPv6
- MPLS (plus)
Data Center Networking
- Familiar with Spine-Leaf architecture.
- Knowledge of AI/HPC networking technologies such as ToR, RoCE, PFC, and ECN.
- Experience with Mellanox / NVIDIA networking products is preferred.
- Experience with 100G / 200G / 400G high-speed network operations is preferred.
- Experience troubleshooting GPU cluster networking issues is a plus.
Operations & Automation
Experience with the following tools is preferred:
- Wireshark
- tcpdump
- iperf3
- mtr
- ethtool
- netperf
Automation experience is a plus:
- Python / Shell
- Ansible
- NetQ / UFM
- Zabbix / Prometheus / Grafana
Cloud & Virtualization (Preferred)
- Familiar with Kubernetes and Docker networking.
- Experience with OpenStack, VMware, or SDN networking is preferred.
- Experience with public cloud networking (AWS / Azure / GCP / Alibaba Cloud) is preferred.
Security Knowledge (Preferred)
- Familiar with network security fundamentals including firewalls, ACLs, IDS/IPS, WAF, and VPN technologies.
- Experience in DDoS mitigation, anti-scanning protection, abnormal traffic analysis, and security incident troubleshooting is preferred.
- Knowledge of Zero Trust architecture, network segmentation, and micro-segmentation is preferred.
- Understanding of common network attack methods and troubleshooting techniques such as ARP spoofing, MAC flooding, TCP SYN flood, and DNS attacks.
- Experience with security auditing, log analysis, and SIEM platforms is preferred.
- Familiar with Linux hardening, SSH security configurations, and access control policies.
Preferred security certifications include:
- CISSP
- CISA
- Security+
- CCNP Security
Preferred Qualities
- Strong incident response capability.
- Good documentation habits and operational discipline.
- Passion for large-scale AI/HPC infrastructure.
- Ability to work effectively under pressure.
- Strong focus on stability and performance optimization.