Applied Researcher – Network Expert

Designworks Talent

Bellevue (KY)

Hybrid

USD 120,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Designworks Talent is seeking an Applied Researcher -- Network Expert for a hybrid role based in Bellevue, WA. You will join an AI infrastructure team focused on large-scale GPU networking.

You will act as horizontal SME, evaluating technologies, architectures, and vendor roadmaps; translate research into engineering decisions; and collaborate with data center, finance, and sales to drive performance and cost efficiency.

Qualifications

  • Deep professional experience in networking at scale.
  • Hands-on experience with high-performance GPU or HPC fabrics.
  • Demonstrated debugging of collective communication performance in production.

Responsibilities

  • Define the future of AI networking strategy.
  • Track vendor roadmaps and industry developments.
  • Formulate defensible fabric product and engineering thesis.
  • Own the fabric position and telemetry strategy.
  • Assist finance with costs per port and optics economics.
  • Support sales with performance and price discussions.
  • Collaborate with Data Center Expert on fabric integration.

Skills

Networking at scale
GPU/HPC fabrics
NCCL/RCCL debugging
NVLink/NVSwitch concepts
InfiniBand RoCEv2
Multi-tenant isolation
Telemetry & observability
Vendor roadmaps

Education

PhD in technical field

Tools

NVIDIA NVLink/NVSwitch
InfiniBand RoCEv2
Ethernet fabrics

Job description

Applied Researcher -- Network Expert

**Location:**Hybrid | Bellevue, WA (downtown)

About the Opportunity

Our client is seeking a Network Expert to join an Applied Research organization focused on the future of large-scale AI infrastructure.

This role is designed for a networking expert with deep experience in GPU-based computing environments and high-performance inter-GPU networking. As AI workloads increasingly rely on thousands of GPUs operating as a single logical system, the network connecting those GPUs becomes a critical component of overall infrastructure performance and scalability.

You will serve as a horizontal subject-matter expert, partnering with engineering and infrastructure leaders to evaluate technologies, architectures, and industry developments. You will help the organization understand where GPU networking is headed and what those developments mean for infrastructure, architecture, and investment decisions.

What You'll Do
  • Hold the company's view of where AI networking is going.
  • Track vendor and hyperscaler roadmaps, research, standards work, and the startup and venture landscape across scale-up, scale-out and scale-across fabrics, topology and optics, collectives and the software above the fabric, operations and telemetry, and tenant-facing capability. Right now that means questions like how fast Ultra Ethernet displaces the RoCEv2 fabrics most clusters actually run, where the scale-up domain should end now that NVLink has open challengers, what co-packaged optics does to power per port, and how to network a cluster that no longer fits in one building. Those specific questions will have changed within a year --- holding the current version of them is the job.
  • Formulate and validate the product and engineering thesis. Turn that view into a defensible position on what we build, buy, or partner for in the fabric, pressure-tested against measured cluster performance, isolation requirements, and cost per port --- and say so plainly when the evidence does not hold up.
  • Own the company's fabric position: where the scale-up and scale-out boundary sits for our workloads, which transport we bet on and when, and what good telemetry and observability look like so fabric problems are diagnosable rather than inferred. A fault that restarts a long training job is a direct cost, not an availability statistic.
  • Help finance to formulate the numbers: cost per port, optics and cabling economics across pluggables and co-packaged options --- reach, power draw, failure rates at scale --- and the performance we can actually substantiate against what vendor benchmarks claim.
  • Own tenant-facing network capability for GPU-as-a-service: multi-tenant isolation and its performance cost, storage traffic alongside GPU traffic, and what we can commit to contractually --- including an honest read of where we are undifferentiated against peers.
  • Make the work land commercially. Support sales and delivery in demanding customer conversations about cluster performance, feed product and go-to-market with what we can offer at what performance and price, and provide technical diligence on network vendors, partners, and prospective tuck-in targets. Work closely with the Data Center Expert where the fabric meets the physical plant.

What We're Looking For

Required Qualifications
  • Deep professional experience in networking at scale,
  • Hands-on experience with high-performance GPU or HPC fabrics at current generations,
  • Demonstrated experience debugging real collective communication performance problems in production

Preferred Qualifications

  • Fluency with the landscape you would be scanning: the switch, NIC, and optics vendors, the standards bodies and consortia, and the startups attacking the fabric layer --- and a view on which of them matter. Expect to be asked what you think is currently overhyped, and why.
  • Demonstrated ability to do research in the applied sense: taking an open question, investigating it from primary sources --- vendor roadmaps, standards drafts, benchmark data, academic literature, your own testing --- and producing a defensible position under genuine uncertainty. A PhD in a relevant technical field is one good route to this and is valued here; sustained industry research, standards-body work, or a body of internal technical assessments that changed real decisions are equally valid. Either way, the role turns on the second half: translating that work for engineering, product, go-to-market, and finance, because it informs all four.
  • Deep professional experience in computer networking and large-scale infrastructure.
  • Significant experience with GPU clusters, AI infrastructure, or high-performance computing environments.
  • Strong understanding of inter-GPU networking.
  • Hands-on experience with InfiniBand at current generations (NDR/XDR), high-performance Ethernet fabrics (RoCEv2, Spectrum-X, or Ultra Ethernet), or comparable HPC interconnects.
  • <
  • Practical understanding of NVLink / NVSwitch scale-up domains and how they interact with the scale-out fabric.
  • Experience debugging real collective communication performance problems --- NCCL/RCCL, congestion, stragglers, topology mismatch --- not just reading about them.
  • Understanding of multi-tenant network isolation and the security and performance tradeoffs involved in serving tenants on shared fabric.
  • Understanding of large-scale AI cluster configurations and the networking requirements associated with thousands of GPUs.
  • Ability to evaluate competing technologies and understand where the networking industry is heading.
  • Familiarity with NVIDIA and AMD GPU infrastructure ecosystems.
  • Strong analytical and technical communication skills.
  • Ability to operate as a horizontal technical expert and influence engineering decisions without necessarily owning implementation.
  • Ability to read technical papers and translate research concepts into practical engineering implications.
  • Ability to operate across engineering, research, and infrastructure organizations.
  • A track record of collaborating with researchers and engineers across groups and levels to shape long‑term research directions and move research into practice.
  • Comfortable operating as an individual contributor with high ownership in a lean, early‑stage team.

Location

  • Hybrid role based in downtown Bellevue, WA.
  • Approximately three days per week in the office.
  • Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.
  • U.S. work authorization is required. Visa sponsorship is not currently available.
  • Export control: this role involves technologies subject to U.S. export control regulations. Candidate eligibility may be subject to export control screening and, where applicable, licensing.
  • Travel: Willingness and ability to travel as needed internationally to data centers and co-locations (up to 25%)

Why Join?

  • High-impact technical role: Directly influence the technology direction of an organization building AI infrastructure at scale.
  • Ground-floor opportunity: Help establish technical strategy, architecture, processes, and culture within a growing organization.
  • High ownership: Operate as a senior individual contributor with substantial autonomy and direct access to senior technical leadership.
  • Cross-disciplinary exposure: Work across AI models, inference, accelerators, software systems, networking, infrastructure, and economics.
  • Cutting-edge technical problems: Work on multi-accelerator inference, intelligent routing, performance optimization, token economics, and compiler, kernel, and runtime technologies.
  • Research with practical impact: Turn emerging research and technology developments into decisions that directly affect engineering, product, commercial strategy, and investment.
  • Lean, senior environment: Work with a small group of highly experienced technical contributors rather than within a large management hierarchy.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied Researcher – Network Expert
Applied Researcher – Network Expert

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Applied Researcher – AI Expert
Applied Researcher – AI Expert

Designworks Talent • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Applied Researcher – AI Expert
Applied Researcher – AI Expert

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 280,000
Applied Researcher – Data Centers
Applied Researcher – Data Centers

Designworks Talent • Bellevue (KY)

Hybrid
USD 180,000 - 230,000
Applied Researcher – Data Centers
Applied Researcher – Data Centers

Designworks Talent • Bellevue (WA)

Hybrid
USD 150,000 - 190,000
Member of Technical Staff - Datacenter Networking
Member of Technical Staff - Datacenter Networking

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Networking
Member of Technical Staff - Datacenter Networking

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Networking at Prime Intellect
Member of Technical Staff - Datacenter Networking at Prime Intellect

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Staff HPC Network Architect
Staff HPC Network Architect

Lambda • United States

Hybrid
USD 180,000 - 280,000
Network Engineer, Supercomputing
Network Engineer, Supercomputing

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision insurance
Unlimited PTO
Paid parental leave
+1