Network Platform Team Lead

Volta

Palo Alto (CA)

On-site

USD 210,000 - 270,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity in Volta
Comprehensive health benefits
Competitive vacation

Job summary

Volta is seeking a senior leader to head the network platform engineering team, shaping the fabric design and automation for large-scale GPU compute infrastructure. You will set direction for fabric, overlay, and multi-tenant isolation, and drive reliability and observability across sites.

You will own APIs and control-plane integrations that expose network capabilities to tenants, collaborate with bring-up teams, and guide vendor engagements while maintaining security and performance at scale.

Qualifications

  • 5+ years in large scale data center or cloud network engineering
  • Production software development in Python/Go/Rust
  • Experience with Ethernet fabric design and ops
  • RoCE v2 at scale and InfiniBand production experience
  • EVPN/VXLAN in production environments
  • Network automation at scale (NetBox/Nautobot, CI)
  • Kubernetes networking (CNI)
  • Multi-vendor evaluation and procurement
  • Security considerations in multi-tenant networks
  • Strong communication and leadership skills
  • Willing to be on site during cluster bring-up

Responsibilities

  • Lead the network platform engineering team: direction, reviews, delivery
  • Set technical direction for platform network layers: fabric, overlay, isolation
  • Own fabric standards (RoCE v2, InfiniBand) across sites
  • Translate product requirements into technical work and roadmaps
  • Define APIs and IaaS control plane exposure for tenants
  • Partner with bring-up teams to turn pain points into scalable features
  • Evaluate OEM designs and hold vendors to Volta's requirements
  • Support customer-facing technical conversations
  • Collaborate on security boundaries and remediation
  • Coordinate with other platform leads on shared architecture
  • Ensure reliability, observability, versioning standards
  • Stay hands-on with high-risk engineering work
  • Run post-mortems and drive continuous improvements
  • Manage the team: 1:1s, performance reviews, hiring

Skills

Data center networking
People management
Python/Go/Rust
Kubernetes networking (CNI)
RoCE v2 / InfiniBand
Network automation
Multi-vendor credibility
Security in multi-tenant networks
Strong communication
On-site bring-up readiness

Tools

NetBox
Nautobot
CI pipelines
Kubernetes tooling

Job description

About Volta

Volta is the category-defining, fully vertically integrated AI infrastructure platform – from capital to clusters to software, under a founder-led enterprise. Our mission is The Utility of Compute™: AI infrastructure as dependable and available as electricity, for every organization that needs it. Launched with a $10B strategic partnership with one of the leading frontier AI labs, a Series A led by Andreessen Horowitz, and a $5B AI Infrastructure Fund, Volta is building the infrastructure layer of the AI era from the ground up. We are 100+ people across London, Palo Alto, and New York, with rapid growth expectations to hundreds.

About The Role

Volta builds and operates large scale GPU compute infrastructure for AI workloads. The network is not a layer underneath the platform, it is part of it. Fabric design, overlay and multi-tenancy, edge connectivity, and the software that programs and observes all of it sit in one team, deliberately.

You will lead that team: a small group of senior network platform engineers who write production software against the fabric rather than only operating it.

What You Will Be Doing
  • Lead the network platform engineering team: technical direction, design reviews, code review, and day-to-day delivery.

  • Set the technical direction for the network layers of the platform: compute and storage fabric, overlay and multi-tenant isolation, edge connectivity, and the telemetry and automation that make them operable.

  • Own the design and evolution of Volta's fabric standards across sites, including RoCE v2 and InfiniBand deployments, and hold the team to a single reference architecture rather than per site variants.

  • Represent network platform engineering in roadmap planning: translate product requirements into scoped, sequenced technical work, communicate trade-offs, and keep delivery on track.

  • Own how network capability is exposed upward: the APIs, abstractions, and IaaS control plane integration through which tenants get isolated, performant networking.

  • Work closely with the bring-up teams to surface operational pain points from cluster deployment and turn them into scalable platform features and automation.

  • Evaluate and challenge network designs from OEMs and partners, and hold vendors to Volta's requirements rather than accepting reference designs as given.

  • Support customer facing technical conversations: workload requirements, fabric design review, and acceptance criteria.

  • Collaborate with the security engineering team on trust boundaries, tenant isolation, and remediation of network layer findings.

  • Coordinate with the other platform engineering leads on shared architecture, joint initiatives, and cross-team and cross-timezone delivery.

  • Own reliability, observability, and interface and versioning standards for the services and fabric your team ships.

  • Stay hands-on: take on complex or high-risk engineering work directly alongside the team, including during bring-up and incident response.

  • Run post-mortems and close structural gaps after incidents, not just the immediate issue.

  • Manage the team: 1:1s, performance conversations, hiring interviews, and onboarding.

What You Bring
  • 5+ years in large scale data center or cloud network engineering, with at least 2 years leading an engineering team.

  • Production software development, not scripting. Python or Go in a shared repository under normal review and CI standards. Our working languages are Python, Go, and Rust.

  • Ethernet fabric design and operations at depth: leaf spine, BGP including unnumbered BGP, ECMP, and the day-two realities of running it at low latency.

  • RoCE v2 at scale: PFC and ECN tuning, DCQCN, and a clear understanding of how it behaves differently from InfiniBand in a GPU training environment.

  • InfiniBand production experience: fat tree topology, UFM, fabric partitioning, adaptive routing, congestion control, with operational and troubleshooting depth rather than design familiarity alone.

  • EVPN and VXLAN in production, including running an overlay under multi-tenant load.

  • Network automation at scale: configuration as code, a source of truth system such as NetBox or Nautobot, CI validation of network change, and declarative or idempotent tooling.

  • Kubernetes networking: CNI, and how workload networking interacts with the underlying fabric.

  • GPU infrastructure experience: how training and inference traffic patterns shape fabric design, and how capacity is surfaced through a platform or API layer.

  • Multi-vendor credibility: able to design, evaluate, and challenge proposals from any major OEM.

  • Security awareness: trust boundaries, least privilege, and secure defaults in multi-tenant network design.

  • Proven people management: running 1:1s, delivering performance feedback, and taking accountability for a team's delivery and wellbeing.

  • Clear communicator who can translate technical complexity for product, leadership, and customer stakeholders.

  • Willing to be on site during cluster bring-up when it matters.

Nice to Have (But Not Essential)

None of these are required. Several map to specific areas of the team's scope, so strength in one or more helps.

  • Fluency with AI-assisted development, and interest in scaling agent-assisted workflows across the team (agentic CLI tools, MCP, skills, APIs) to amplify delivery.

  • ASN operations: running a public autonomous system, BGP peering with transit providers and IXes, RPKI and IRR hygiene, and DDoS posture.

  • IPv6 at production scale: dual stack DC design, v6 BGP peering, and addressing architecture.

  • NVIDIA Spectrum-X, including NetQ and Cumulus, or SONiC and whitebox platforms.

  • gNMI, OpenConfig, or NETCONF/YANG based telemetry and configuration.

  • OVN/OVS, SR-IOV, DPDK, or BlueField DPU based networking.

  • Familiarity with NVLink and NVSwitch topologies and NCCL behaviour.

  • Experience working distributed across time zones with counterparts in other regions.

  • Open source contributions to networking or infrastructure projects.

  • REQ-76

What We Offer

At Volta, we believe people do their best work when they feel supported, trusted and able to grow. We're building a company where you can make an impact, keep a healthy balance between work and life, and build a career you're proud of.
As a global team, we do our best to provide great benefits wherever you're based. While some benefits vary by country due to local regulations, we believe looking after our people is simply the right thing to do.

  • Competitive salary based on the work you do here, not your previous salary

  • Equity in Volta, giving you the opportunity to share in the company's long-term success

  • Retirement/pension contributions

  • Comprehensive health, wellbeing and insurance benefits

  • Generous number of vacation days each year

Additional Information
Equal Opportunity

Volta is an equal opportunity employer. We are committed to building a diverse and inclusive team and make employment decisions based on skills, qualifications, experience and business needs. We do not discriminate on the basis of race, colour, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status or any other legally protected characteristic.

Accessibility

Volta is committed to providing an accessible recruitment experience for all candidates. If you require accommodations or adjustments at any stage of the application or interview process, please contact us at askpeople@volta.com. We will work with you to identify reasonable accommodations that enable you to participate fully in the hiring process.

Candidate Privacy Notice

By applying, you consent to the processing of your personal data for recruitment purposes in accordance with applicable data protection laws, including the UK GDPR, EU GDPR and relevant US state privacy regulations. Your data will be shared only with those involved in the hiring process and will not be used for unrelated purposes. For details, see our Recruitment & Candidate Privacy Notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Network Engineer (3x Openings)
Senior Network Engineer (3x Openings)

Volta • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Competitive salary
Equity in Volta
Retirement/pension contributions
+2
Platform Engineer (12x Openings)
Platform Engineer (12x Openings)

Volta • Palo Alto (CA)

On-site
USD 140,000 - 210,000
Equity in Volta
Retirement contributions
Health benefits
+2
VP Platform Engineering
VP Platform Engineering

Volta • Palo Alto (CA)

On-site
USD 260,000 - 380,000
Equity
Retirement plan
Health benefits
+1
Network Modeling / Automation Engineer
Network Modeling / Automation Engineer

Volta • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Equity in Volta
Competitive salary
Comprehensive health benefits
+2
Head of Project Engineering
Head of Project Engineering

Volta • Palo Alto (CA)

On-site
USD 240,000 - 320,000
Equity in Volta
Retirement/pension contributions
Comprehensive health benefits
+1
Team Lead, Platform Engineering (2x Openings)
Team Lead, Platform Engineering (2x Openings)

Volta • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Retirement plan
+2
Group Technical Program Manager
Group Technical Program Manager

Volta • Palo Alto (CA)

On-site
USD 160,000 - 230,000
Equity in Volta
Retirement/pension contributions
Comprehensive health, wellbeing and保险
Senior Technical Program Manager (5x Openings)
Senior Technical Program Manager (5x Openings)

Volta • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Competitive salary
Equity in Volta
Retirement contributions
+2
Group Program Manager - Technical Programs
Group Program Manager - Technical Programs

Volta • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Equity in Volta
Competitive salary
Retirement/pension contributions
+2
IT Analyst
IT Analyst

Volta • Palo Alto (CA)

On-site
USD 90,000 - 120,000
Competitive salary
Equity in Volta
Comprehensive benefits
+1