Principal Network Engineer

nscaleoperationsukltd

Seattle (WA)

On-site

USD 180,000 - 260,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nscale, a GPU cloud engineered for AI, seeks a Principal Network Engineer to lead the design and operation of low-latency, high-bandwidth networks powering large-scale AI training and inference workloads. You will define reference architectures, drive architectural decisions, and mentor a team while collaborating with data center operations, platform engineering, and vendors.

You will own the direction for data center fabrics, automation, and observability, ensuring scalable, reliable networking

Qualifications

  • 10+ years of network engineering experience in HPC/AI scale environments.
  • Expert-level knowledge of data center routing and control planes (BGP, EVPN-VXLAN).
  • Hands-on with InfiniBand/RoCE fabrics and related fabric orchestration.

Responsibilities

  • Define, design, and validate large-scale InfiniBand/RoCE and Ethernet fabric architectures across rack, row, and data center scales.
  • Own technical direction for high-performance fabrics and establish reference architectures across sites.
  • Lead network automation strategy using Python/Ansible, with GitOps workflows and IaC tooling.
  • Mentor engineers and drive standards for architecture, automation, and observability.

Skills

HPC networking
RDMA networking (InfiniBand/RoCE)
Subnet managers (OpenSM/UFM)
BGP EVPN-VXLAN
Network automation (Python/Ansible)
IaC / CI pipelines (Terraform, GitLab)

Tools

OpenSSH/Net automation tooling
Cumulus Linux / Nokia / Arista EOS (multi-vendor)

Job description

Principal Network Engineer, AI Infrastructure & High-Performance Networking
About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud strengthens technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

At Nscale, our Engineering team plays a critical role in deploying and operating the infrastructure and software platforms that power our customers.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.

About the Role

The Network Engineering Team is responsible for the design, validation, and ongoing operation of all networking services that underpin both the internal management platform and the customer-facing cloud infrastructure - including high-performance Ethernet fabrics, InfiniBand, WAN connectivity, and data center networking. The team also acts as a 3rd/4th line escalation point for the support organization.

As a Principal Network Engineer, you will be a senior technical authority for Nscale’s AI-optimized network fabrics. You will set technical direction across low-latency, high-bandwidth InfiniBand and Ethernet networks supporting large-scale training and inference workloads; own critical technical domains end to end; and raise the bar for architecture, automation, operational rigor, and engineering standards across the organization.

You will combine deep hands-on engineering with broad architectural influence. You'll define reference architectures, drive consistency across sites, lead complex technical decisions and escalations, and mentor engineers while partnering closely with deployment, data center operations, platform engineering, and vendors.

What You'll Be Doing
  • Define, design, validate, and evolve large-scale InfiniBand/RoCE and Ethernet fabric architectures at rack, row, and data center scale, with tight integration to bare-metal provisioning and cluster management systems.
  • Own technical direction for high-performance Ethernet fabrics, including BGP, EVPN, VXLAN, LACP, and QoS, and establish reference architectures and standards implemented consistently across sites.
  • Design and engineer perimeter and security infrastructure - firewalls, NAT, VPN, and security policy architecture - across WAN and data center edge environments.
  • Lead network automation strategy in a GitOps model, building and guiding Python/Ansible tooling for provisioning, configuration validation, and compliance, with version-controlled configuration and CI/CD-driven change across multi-vendor environments.
  • Drive operational excellence by leading complex escalations and root-cause analysis for performance and stability issues, and systematically reducing reactive toil through runbooks, automation, and measurable SLOs.
  • Set the direction for network observability, telemetry, monitoring, and alerting to provide clear visibility into fabric health, performance, and traffic patterns.
  • Ensure the accuracy and reliability of source‑of‑truth network inventory and configuration data, with changes flowing through structured engineering and change‑management practices.
  • Partner with deployment, data center operations, platform engineering, systems, storage, and vendors on new site delivery and platform evolution.
  • Act as a technical mentor and force multiplier across the team through architecture reviews, design reviews, incident leadership, documentation, and knowledge sharing.
  • Identify systemic risks and architectural gaps across sites and drive durable solutions that improve scalability, reliability, and operational simplicity.
About You (Skills / Qualifications)
  • 10+ years of network engineering experience, with significant depth in HPC, AI, hyperscale, or large-scale data center environments.
  • Extensive hands-on experience with RDMA-aware networking for AI/HPC workloads, including InfiniBand and/or RoCE, subnet managers such as OpenSM/UFM, and fabric orchestration.
  • Expert-level knowledge of modern data center routing and control planes, including BGP, EVPN-VXLAN, and Clos/spine-leaf architectures, with production experience on platforms such as Cumulus, Nokia, or Arista EOS.
  • Strong network automation expertise using Python and Ansible, Git-based workflows, and modern infrastructure-as-code and pipeline tooling such as Terraform, GitLab CI, or GitHub Actions; you treat the network as code rather than managing devices by hand.
  • Deep design and engineering experien
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Network Engineer
Senior Network Engineer

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Senior Network Engineer
Senior Network Engineer

Nscale • San Francisco (CA)

On-site
USD 150,000 - 210,000
Medical insurance
Retirement plan
Flexible PTO
Principal Network Engineer
Principal Network Engineer

Nscale • Seattle (WA), New York (NY), San Francisco (CA), Houston (TX)

On-site
USD 180,000 - 240,000
Base salary + equity
Equity incentives
Dynamic progression plan
Principal Network Engineer
Principal Network Engineer

Socket.dev • Houston (TX)

On-site
USD 180,000 - 270,000
Equity
Competitive base package
Career growth opportunities
Senior Network Engineer
Senior Network Engineer

Nscale • Seattle (WA)

On-site
USD 150,000 - 210,000
Senior Network Engineer
Senior Network Engineer

Nscale • New York (NY)

On-site
USD 150,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior Back-End Network Engineer - AI Infrastructure Operations
Senior Back-End Network Engineer - AI Infrastructure Operations

Nscale • New York (NY)

On-site
USD 150,000 - 240,000
Equity
Comprehensive benefits
Retirement plan
Senior Back-End Network Engineer - AI Infrastructure Operations
Senior Back-End Network Engineer - AI Infrastructure Operations

Nscale • Houston (TX)

On-site
USD 150,000 - 240,000
Equity
Medical, dental, vision
Flexible paid time off
Senior Back-End Network Engineer - AI Infrastructure Operations
Senior Back-End Network Engineer - AI Infrastructure Operations

Nscale • Seattle (WA)

On-site
USD 150,000 - 240,000
Equity
Medical/dental/vision
Flexible PTO
+2
Senior Back-End Network Engineer - AI Infrastructure Operations
Senior Back-End Network Engineer - AI Infrastructure Operations

Nscale • San Francisco (CA)

On-site
USD 150,000 - 240,000
Base + equity
Medical, dental, vision
Flexible PTO