Senior HPC Networking Engineer

Mirantis

United States

On-site

USD 170,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Professional development
Conferences and tech talks

Job summary

Mirantis, a Kubernetes-native AI infrastructure company, seeks a Senior HPC Networking Engineer to design, deploy, and support high-performance networks for AI and HPC workloads. The role emphasizes InfiniBand fabrics, Fortinet security, and scalable infrastructure across data centers.

You will troubleshoot complex networking issues, optimize performance, and collaborate with compute, storage, and platform teams while building documentation and playbooks for operations and future platform growth.

Qualifications

  • 5+ years of experience in network engineering for HPC/data center.
  • Strong hands-on experience with InfiniBand technologies (Mellanox/NVIDIA).
  • Solid knowledge of TCP/IP, routing (BGP/OSPF), VLANs, QoS.
  • Experience deploying Fortinet solutions (FortiGate/ FortiManager) and VPNs.
  • Linux familiarity and scripting for automation (Bash, Python).

Responsibilities

  • Design, deploy, and maintain HPC network infrastructures with InfiniBand fabrics.
  • Troubleshoot network issues across InfiniBand and Ethernet for performance.
  • Manage InfiniBand components: switches, HCAs, subnet managers, fabric configs.
  • Perform capacity planning and performance tuning for HPC networks.
  • Collaborate with compute, storage, and platform teams to support HPC workloads.
  • Develop and maintain network architecture docs and runbooks.
  • Handle on-call rotations and incident escalation.

Tools

Fortinet FortiGate
FortiManager
VPNs
Firewall policies
Bash
Python

Job description

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. https://www.mirantis.com/

Location: US

Employment Type: Full-time

Job Description

We are seeking a highly skilled Senior HPC Networking Engineer to design, deploy, manage, and troubleshoot high-performance networking environments. The ideal candidate will have deep expertise in InfiniBand technologies, strong general networking knowledge, and hands-on experience with Fortinet solutions. You will play a critical role in ensuring the performance, reliability, and scalability of HPC infrastructure.

Key Responsibilities
  • Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics.
  • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
  • Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations.
  • Perform performance tuning, monitoring, and capacity planning for HPC networking systems.
  • Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).
  • Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments.
  • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.
  • Develop and maintain documentation for network architecture, configurations, and operational procedures.
  • Participate in on-call rotations and provide escalation support for critical incidents.
  • Lead or contribute to network upgrades, migrations, and new deployments.
  • You will actively troubleshoot and resolve daily customer incident tickets to keep massive GPU training runs moving.
Build, operate, and scale next-generation GPU infrastructure
  • You will be hands-on on the front lines driving daily triage, incident resolution, and SLA management for the world’s most advanced NVIDIA clusters, InfiniBand/RoCE fabrics, and AI workloads.
Build the playbook, then grow into the platform
  • Designed for engineers energized by standing up new operations from the ground up, this role offers a direct trajectory from operationalizing bare-metal clusters to driving platform and AI capabilities as we scale.
Qualifications
Required
  • 5+ years of experience in network engineering, with a focus on HPC or data center environments.
  • Strong hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).
  • Solid understanding of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.
  • Proven experience deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).
  • Experience with network performance analysis and troubleshooting tools.
  • Familiarity with Linux systems and scripting for automation (e.g., Bash, Python).
  • Strong analytical and problem-solving skills.
Preferred
  • Experience with large-scale HPC clusters or AI/ML infrastructure.
  • Knowledge of RDMA, MPI, and low-latency networking concepts.
  • Certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent.
  • Experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform).
Soft Skills
  • Strong communication and collaboration skills.
  • Ability to work independently and handle complex technical challenges.
  • Detail-oriented with a proactive approach to problem-solving.
What We Offer
  • Opportunity to work on cutting-edge HPC infrastructure.
  • Collaborative and innovative work environment.
  • Competitive salary and benefits package.
Additional Information
What does Mirantis offer you
  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.

We are a Leader for Container Management in G2 (#2 after AWS)!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal HPC Network Engineer (remote in the US)
Principal HPC Network Engineer (remote in the US)

Mirantis • United States

Remote
USD 140,000 - 190,000
Senior HPC Networking Architect - InfiniBand & AI Infra
Senior HPC Networking Architect - InfiniBand & AI Infra

Mirantis • United States

On-site
USD 170,000 - 210,000
Competitive salary
Professional development
Conferences and tech talks
Senior HPC Networking Architect (Remote, InfiniBand Expert)
Senior HPC Networking Architect (Remote, InfiniBand Expert)

Mirantis • United States

Remote
USD 140,000 - 190,000
EU Remote HPC Network Engineer (InfiniBand + Fortinet)
EU Remote HPC Network Engineer (InfiniBand + Fortinet)

Visa Hunt • Town of Poland (NY)

Remote
USD 90,000 - 105,000
Advanced AI infra
NVIDIA GPU tech
Kubernetes platforms
+1
AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

On-site
USD 120,000 - 180,000
Professional development and training
Conferences and working groups
Company outings and social events
+1
Senior Technical Support Engineer, Network
Senior Technical Support Engineer, Network

NVIDIA • Town of Sweden (NY)

On-site
USD 120,000 - 160,000
Senior Technical Support Engineer, Network
Senior Technical Support Engineer, Network

NVIDIA • Germany (OH)

On-site
USD 120,000 - 180,000
Infrastructure Engineer
Infrastructure Engineer

Calance • Illinois

On-site
USD 110,000 - 150,000
Senior Solutions Architect, First Time Deployment Networking InfiniBand - NVIS
Senior Solutions Architect, First Time Deployment Networking InfiniBand - NVIS

NVIDIA • Virginia (MN)

On-site
USD 148,000 - 236,000
Equity
Benefits
Senior HPC Networking Architect | InfiniBand & Fortinet
Senior HPC Networking Architect | InfiniBand & Fortinet

United States Digital Space LLC • United States

Remote
USD 140,000 - 190,000
Competitive pay
Full benefits
Professional development
+1