Director, Network Capacity Automation

Socket.dev

Bellevue (WA)

On-site

USD 210,000 - 260,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision coverage
401k with company match (USA)
Wellness and commuter stipends
Generous cash & equity compensation
Flexible paid time off

Job summary

Lambda, The Superintelligence Cloud, seeks a Director of Network Capacity Automation to lead teams building software that enables capacity planning and delivery ahead of demand for our GPU cloud infrastructure. This position reports to the VP of Cloud and AI Networking and sits at the helm of a critical growth area.

You will hire and develop engineers, shape the org culture, and partner with technical leadership to maintain high standards in design, code quality, and operational reliability.

Qualifications

  • 8+ years of engineering management experience, including senior ICs
  • Proven record building and scaling high-performance cloud connectivity teams
  • Background in software engineering with production systems experience
  • Strong automation and data-driven approach
  • Experience designing, building, and operating large-scale networks

Responsibilities

  • Own network capacity end to end — planning, delivery, turn-up and lifecycle — and be accountable for capacity landing ahead of demand
  • Build and lead the organization that delivers it: hire, develop and retain engineers and the leaders who manage them, and design the team structure as the org grows
  • Set the multi-year technical strategy for capacity automation to ensure it continually scales
  • Convert manual process into durable software — the goal is a capacity pipeline that runs without heroics, not a better-organised set of runbooks
  • Own the operational bar for your org: how capacity work is planned, reviewed, shipped, measured and learned from when it goes wrong
  • Manage through leads and senior ICs, developing both — regular 1:1s, clear performance feedback, growth planning, and sponsorship of meaningful work
  • Partner and support our Interconnect team who negotiate, acquire and coordinate on external connectivity
  • Partner with Principal Engineers and technical leadership to maintain a high engineering bar across design, code quality and operational reliability
  • Partner with Infrastructure, Supply Chain, Data Center Operations, Product and Finance to align capacity roadmaps, forecast demand and manage long-lead dependencies
  • Represent your org's work and technical direction to senior leadership, and to enterprise customers and partners where appropriate
  • Contribute to org-wide engineering process improvements — how we plan, how we ship, how we learn from incidents

Skills

Engineering management
Network capacity planning
Software engineering background
Automation mindset
Distributed systems design
Large-scale networking
Team leadership
Communication skills

Job description

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our Bellevue or San Francisco office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

Our vision is bold and is not an incremental exercise. We will continually re-evaluate and reinvent our current ways of working and the automation underneath all of it, while operating the existing network flawlessly for all customer workloads.

Lambda has 10x'd over the last three years, and the network engineering organization is scaling to match. We're looking for a Director of Network Capacity Automation to lead one or more teams building the software that enables us to plan and deliver network capacity ahead of customer demand. This capacity is critical to ensure a flawless customer connectivity experience to our GPU cloud infrastructure. We're looking for a Director to own that outcome end to end — and to build the organization that delivers it. You'll hire and grow a talented team of software and network proficient engineers, shape team culture, and partner with technical leadership to deliver the mechanisms needed to stay ahead of our scale ambitions.

This role will report to the VP of Cloud and AI Networking, working alongside a highly capable and passionate team of engineers.

If you'd like to join a great team of customer obsessed engineers building the world's best AI cloud, come join us.

What You'll Do
  • Own network capacity end to end — planning, delivery, turn-up and lifecycle — and be accountable for capacity landing ahead of demand

  • Build and lead the organization that delivers it: hire, develop and retain engineers and the leaders who manage them, and design the team structure as the org grows

  • Set the multi-year technical strategy for capacity automation to ensure it continually scales

  • Convert manual process into durable software — the goal is a capacity pipeline that runs without heroics, not a better-organised set of runbooks

  • Own the operational bar for your org: how capacity work is planned, reviewed, shipped, measured and learned from when it goes wrong

  • Manage through leads and senior ICs, developing both — regular 1:1s, clear performance feedback, growth planning, and sponsorship of meaningful work

  • Partner and support our Interconnect team who negotiate, acquire and coordinate on external connectivity

  • Partner with Principal Engineers and technical leadership to maintain a high engineering bar across design, code quality and operational reliability

  • Partner with Infrastructure, Supply Chain, Data Center Operations, Product and Finance to align capacity roadmaps, forecast demand and manage long-lead dependencies

  • Represent your org's work and technical direction to senior leadership, and to enterprise customers and partners where appropriate

  • Contribute to org-wide engineering process improvements — how we plan, how we ship, how we learn from incidents

You
  • Have 8+ years of engineering management experience, including directly managing senior individual contributors

  • Proven track record of building and scaling engineering organizations that deliver mission-critical, high-performance cloud connectivity serving millions of users

  • Have a software engineering background — you've shipped production systems and can engage credibly on architecture, technical tradeoffs, and code quality

  • Foster a data driven, automation first, high velocity organization

  • Familiarity designing, building, scaling and operating large-scale networks

  • Familiarity with how to build distributed systems and large-scale networking services

  • Strong automation skills and track record accelerating a business delivery velocity on a large scale

  • Have a track record of building high-performing teams in fast-moving, technically demanding environments

  • Are skilled at translating ambiguous business and product goals into clear team priorities and executable engineering plans

  • Show strong judgment about when to go deep technically, when to delegate, and when to escape

  • Track record of successful project and product delivery in fast-paced, high-pressure environments

  • Excellent communicator both written and verbal. You leverage high quality written artifacts to inform high quality decision making and get approval on problems and solution opportunities

Nice to Have
  • Experience in GPU cloud, HPC, or AI/ML infrastructure environments

  • Deep familiarity with large-scale data center or cloud networking — fabric build-out, capacity modelling, or network supply chain

  • Experience with infrastructure capacity forecasting and long-lead hardware planning

  • Prior experience at a high-growth infrastructure or cloud company

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda
  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Network Engineer
Principal Network Engineer

Socket.dev • Bellevue (WA)

On-site
USD 180,000 - 260,000
Health coverage
Dental coverage
Vision coverage
+3
Principal Network Engineer
Principal Network Engineer

Lambda • San Francisco (CA)

Hybrid
USD 180,000 - 290,000
Health, dental, and vision coverage
Wellness stipends
Commuter stipends
+2
Engineering Manager, Fleet Engineering
Engineering Manager, Fleet Engineering

Socket.dev • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision coverage
Wellness stipend
401k plan with company match (USA)
Senior Software Engineer – Core Cloud Platform
Senior Software Engineer – Core Cloud Platform

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+1
Senior Site Reliability Engineer - Fleet
Senior Site Reliability Engineer - Fleet

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+4
Senior Director, Product Management - Foundations
Senior Director, Product Management - Foundations

Lambda • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision coverage
401k plan with company match
Wellness and commuter stipends
+1
Engineering Manager, Fleet Engineering
Engineering Manager, Fleet Engineering

Lambda • San Francisco (CA)

Hybrid
USD 297,000 - 440,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+2
Director, Product Management - Foundations
Director, Product Management - Foundations

Lambda Inc. • Northern (KY)

Hybrid
USD 150,000 - 190,000
Cash & equity compensation
Health, dental, and vision
Wellness stipend
+2
Senior Site Reliability Engineer - Fleet
Senior Site Reliability Engineer - Fleet

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Technical Success Engineer
Technical Success Engineer

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+1