Senior AI Network & Security Engineer

Techstreet Malaysia

Johor

On-site

MYR 150,000 - 230,000

Full time

11 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Techstreet Malaysia is seeking a senior network engineer to design, implement and operate a high-performance AI/GPU data-center network infrastructure in Johor. You will lead complex incidents, deployments and expansions, mentor engineers and collaborate with Operations, Security and Platform teams while participating in a 24x7 on-call roster.

The role requires 8+ years of experience, strong TCP/IP, routing, switching, BGP and firewall skills, and hands-on experience with NVIDIA InfiniBand,

Qualifications

  • Bachelor's degree in Network Engineering or related field.
  • 8+ years of network engineering experience in data-centre/cloud environments.
  • Hands-on TCP/IP, routing, switching, VLANs, BGP, and high availability.
  • Experience with Juniper, Cisco, Arista switches/routers.
  • Experience with enterprise firewall technologies and security policies.
  • Willingness to work on-site in Johor and participate in 24x7 on-call.

Responsibilities

  • Design, implement, operate and maintain AI/GPU data-centre network infrastructure.
  • Lead network incidents and deployments; mentor engineers.
  • Monitor performance, troubleshoot connectivity and GPU workloads.
  • Develop AI networking expertise: InfiniBand, RoCE, RDMA, QoS.
  • Coordinate with security, platform and operations teams.
  • Maintain network diagrams, SOPs, runbooks and documentation.

Skills

TCP/IP & routing
LAN/WAN/VLAN
Security policies
Troubleshooting
Automation (Python/Ansible)

Education

Bachelor's degree in Networking or related field

Tools

Juniper
Cisco
Arista
NVIDIA InfiniBand
RoCE

Job description

Responsible for the design, implementation, operation and continuous improvement of network and security infrastructure supporting large-scale AI/GPU computing platforms, data centres and cloud services.

The role combines strong data-center networking and security fundamentals with high-performance AI networking technologies, including NVIDIA InfiniBand, Spectrum Ethernet/Spectrum-X and RoCEv2. The candidate must demonstrate strong networking fundamentals, hands-on troubleshooting capability and the ability to rapidly learn and develop expertise in AI/GPU networking.

The role will also provide technical leadership during complex incidents, infrastructure deployments and network expansion projects, while mentoring other engineers and working closely with Operations, Systems, Platform, Security, Data Centre teams and external technology partners. The role requires participation in on-call support as needed.

Responsibilities
  • AI & Data Center Network Infrastructure
  • Design, implement, operate and maintain highly available network infrastructure supporting AI/GPU clusters, data centres and AI Cloud services.
  • Support high-performance GPU network fabrics including NVIDIA InfiniBand and Ethernet/RoCE-based architectures.
  • Design and operate data-center networking technologies including routing, switching, VLANs, BGP, EVPN/VXLAN and high-availability architectures.
  • Configure and support network infrastructure across compute, storage, management, out-of-band (OOB), customer and external connectivity networks.
  • Support LAN, WAN, VPN and private interconnect connectivity between data centres, customers, partners and cloud environments.
  • Participate in network architecture, capacity planning and infrastructure expansion activities for new AI/GPU clusters.
  • Network Operations & Performance
  • Monitor network and AI fabric availability, throughput, latency, utilisation, errors and congestion to ensure infrastructure meets performance and SLA requirements.
  • Troubleshoot complex connectivity and performance issues across switches, ConnectX NICs, DPUs and SuperNICs, servers, host networking and GPU workloads.
  • Work with Systems, Platform and Operations teams to identify network-related issues impacting distributed GPU workloads.
  • Develop expertise in AI networking technologies including InfiniBand, RDMA, RoCEv2, lossless Ethernet, PFC, ECN, QoS and congestion management.
  • Support AI fabric monitoring and management platforms such as NVIDIA UFM, NetQ or equivalent tools.
  • Support network validation, commissioning and performance testing for new GPU clusters and infrastructure deployments.
  • Network Security
  • Design, implement and operate network security infrastructure including firewalls, VPNs, ACLs, segmentation, NAT, IPS and secure connectivity.
  • Configure and manage Fortinet/FortiGate or equivalent enterprise firewall platforms.
  • Implement appropriate network segmentation across compute, management, storage, OOB, customer and external-facing environments.
  • Support security hardening, vulnerability remediation and compliance requirements.
  • Work closely with Security and Risk teams during security incidents, assessments and audits.
  • Network Automation
  • Develop and maintain network automation using technologies such as Python, Ansible, APIs, DCIM and Infrastructure as Code.
  • Automate configuration deployment, backups, compliance validation, provisioning and routine operational activities.
  • Support CI/CD and controlled network change processes to improve consistency, reliability and auditability.
  • Work with platform teams to integrate network and fabric telemetry into monitoring platforms.
  • Operations & Reliability
  • Provide technical leadership for complex network incidents, outages and performance degradation.
  • Perform root-cause analysis and drive corrective and preventive actions.
  • Analyse network performance trends, capacity, hardware utilisation and growth requirements.
  • Plan and test redundancy, failover and recovery mechanisms.
  • Participate in change management, maintenance, upgrades and lifecycle management activities.
  • Participate in the operational standby/on-call roster supporting 24×7 AI Cloud services.
  • Develop and maintain HLD/LLD, network diagrams, SOPs, MOPs, EOPs, troubleshooting runbooks and technical documentation.
Requirements

Bachelor's degree in Network Engineering, Computer Science, Information Technology or related discipline, or equivalent practical experience.

8+ years of relevant network engineering experience, preferably within data-center, cloud, service-provider or large-scale infrastructure environments.

Strong hands-on knowledge of TCP/IP, routing, switching, VLANs, BGP and network redundancy/high availability.

Strong experience designing, implementing and troubleshooting production network infrastructure.

Hands-on experience with enterprise/data-center switching and routing platforms such as Juniper, Cisco, Arista or equivalent.

Experience with enterprise firewall technologies, network segmentation, ACLs, VPNs and security policies.

Solid understanding of advanced networking technologies, particularly those related to AI would be highly advantageous.

Strong communication skills, both written and verbal.

Excellent problem-solving and analytical skills.

Ability to work independently and as part of a team.

Willingness to work site-based in Johor and to participate in a 24×7 escalation roster, if required.

Hands-on experience with NVIDIA InfiniBand, Spectrum Ethernet Platform, and/or RDMA over Converged Ethernet (RoCE) preferred.

Good understanding of Linux and host networking.

Enterprise & AI network delivery (LAN, WAN, WLAN, VPN)

Secure connectivity (site-to-site VPN, MPLS, private interconnects)

Network security (firewalls, VLANs, ACLs, zero-trust)

Operations, monitoring & incident response

Certifications such as NVIDIA-Certified InfiniBand or Networking Professional, JNCIP or JNCIE, CCNP or CCIE, Fortinet NSE 4 and above, CISSP or CISM.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Network & Security Engineer
Senior AI Network & Security Engineer

techstreet • Johor

On-site
MYR 180,000 - 300,000
Health Insurance
Performance Bonus
Dental Coverage
Senior AI Network & Security Engineer (Johor Bahru)
Senior AI Network & Security Engineer (Johor Bahru)

Techstreet • Johor Bahru

On-site
MYR 180,000 - 300,000
AI Network & Security Engineer - Data Centre
AI Network & Security Engineer - Data Centre

Neuron Solutions Sdn. Bhd. • Johor

On-site
MYR 120,000 - 180,000
Senior AI Networking & Security Architect
Senior AI Networking & Security Architect

Techstreet Malaysia • Johor

On-site
MYR 150,000 - 230,000
AI Data Centre Network & Security Engineer
AI Data Centre Network & Security Engineer

Neuron Solutions Sdn. Bhd. • Johor

On-site
MYR 120,000 - 180,000
Technical Manager - GPU Cloud & AI Infrastructure
Technical Manager - GPU Cloud & AI Infrastructure

Risewave Consulting, Inc. • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Network Engineer
Network Engineer

Raydian Cloud Sdn Bhd • Kuala Lumpur

On-site
MYR 3,200 - 6,400
Senior AI Networking & Security Architect
Senior AI Networking & Security Architect

techstreet • Johor

On-site
MYR 180,000 - 300,000
Health Insurance
Performance Bonus
Dental Coverage
Senior Data Centre Operations Engineer
Senior Data Centre Operations Engineer

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 120,000 - 180,000
Senior Network Engineer
Senior Network Engineer

NTT DATA Business Solutions • Cyberjaya

On-site
MYR 120,000 - 180,000
Health Insurance
Optical benefits
Dental benefits
+1