We're building AI/HPC infrastructure that powers the next generation of high-performance computing. As a Network Engineer, you'll responsible to design, build, and operate large-scale network fabrics that support GPU clusters, HPC workloads, and mission-critical AI systems. You'll engage across the full lifecycle—from requirements gathering and architecture design through deployment, optimization, and operational support. This is hands-on infrastructure work.
What you'll be doing:
- Design and deploy EVPN-VXLAN Clos/Leaf-Spine Ethernet fabrics and InfiniBand networks for high-performance computing environments
- Configure and troubleshoot BGP, OSPF, VXLAN, and EVPN protocols across multi-vendor platforms (Arista, Cumulus, SONiC, Juniper, Cisco)
- Deploy lossless fabrics (Ethernet and InfiniBand) which use RDMA technologies for high-performance computing.
- Build and validate network topologies in lab environments using simulation tools (NVIDIA Air, GNS3, EVE-NG)
- Develop and optimize network automation solutions using Python and Ansible for network infrastructure
- Support operational excellence: monitor network health, troubleshoot production issues, and coordinate with vendors for rapid resolution
- Create comprehensive documentation, operational runbooks, and knowledge bases for customer support and knowledge transfer
What we need to see:
Core Networking (10+ years)
- Deep expertise in TCP/IP stack, Layer 2/3 switching, routing, and data center architecture
- Expert-level knowledge of BGP, OSPF, EVPN, and VXLAN (5+ years hands-on experience)
- Strong understanding of Clos/Leaf-Spine fabric design and multi-tier network architectures
- Advanced network troubleshooting using packet sniffers, analysis tools, and diagnostic techniques
- Hands-on with at least one network operating system viz. Arista EOS, Cumulus Linux, Cisco, Nokia SR Linux etc.
- Configuration, testing, validation, and issue resolution on these platforms in production environments
- Ethernet lossless networking concepts (PFC, ECN, DCQCN, RoCEv2)
- InfiniBand networks for HPC: configuration, troubleshooting, and performance optimization in large-scale environments
- Linux networking knowledge
Automation & Infrastructure-as-Code
- Expert-level Python scripting for network automation and tooling
- Extensive experience with Ansible, Salt, or equivalent infrastructure automation frameworks
- Ability to develop CI/CD pipelines for network infrastructure provisioning and testing
- Git version control and Infrastructure-as-Code practices
- Lab environment simulation and topology validation using NVIDIA Air, GNS3, EVE-NG, or similar tools
Ways to stand out from the rest:
- Networking certifications: Cisco, Arista, Nokia, Nvidia or equivalent
- Kubernetes and Docker; container networking
- Cloud networking: AWS, GCP, Azure; hybrid cloud and on-premises integration
- Personal GitHub repository with Python examples and evidence of coding expertise
Minimum Qualifications:
Bachelor’s degree in computer science, Electrical/Computer Engineering, Physics, Mathematics, or related field (or equivalent industry experience) years of networking fundamentals and data center architecture
- 5+ years of hands‑on EVPN/BGP/VXLAN experience
- Proficiency in at least one major platform (Arista, Cumulus, SONiC, Juniper, Cisco)
- Production troubleshooting experience.
Soft Skills:
- Strong problem-solving and debugging abilities
- Ownership mindset with accountability for deployment quality and customer success
- Cross-functional collaboration with other teams
- Proactive approach to continuous learning and staying current with AI infrastructure trends
- Strong documentation and presentation skills, with ability to defend design decisions in peer reviews.