Principal Network Developer

Ll Oefentherapie

Nashville (TN)

On-site

USD 150,000 - 190,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Ll Oefentherapie is seeking an expert Network Architect to lead the design, deployment, and operation of large-scale RDMA/RoCE network fabrics across OCI AI, GPU, and HPC infrastructure in a US-based setting. You will drive scalable designs, validate configurations, and guide cross-functional teams in performance and reliability initiatives.

You will mentor engineers, develop automation pipelines, and own test strategies, change reviews, and deployment readiness for high-performance networks,

Qualifications

  • Experience designing, operating, and troubleshooting large-scale RDMA/RoCE, cloud, data center, or high-performance networks.
  • Knowledge of RDMA, RoCE, Ethernet fabrics, congestion control, QoS, and AI/GPU networking.
  • Expertise in routing and switching technologies, including BGP, OSPF, EVPN-VXLAN, and data center networking.
  • Experience with network automation using Python, Ansible, APIs, or similar technologies.
  • Experience with network telemetry, observability, monitoring, performance analysis, and incident management.
  • Experience supporting hyperscale cloud, AI/GPU, HPC, or large-scale distributed infrastructure.
  • Ability to lead complex technical initiatives and collaborate across engineering, operations, customers, and vendors.

Responsibilities

  • Lead the design, deployment, validation, and lifecycle management of large-scale RDMA/RoCE network fabrics supporting OCI AI, GPU, and HPC infrastructure.
  • Translate network architectures into scalable designs and deployment plans, ensuring performance, reliability, and operational readiness across OCI’s global cloud environment.
  • Serve as technical lead for complex network initiatives spanning RDMA/RoCE fabrics, data center networking, automation, testing, deployment, and operations.
  • Develop automation frameworks, tools, scripts, and infrastructure pipelines to improve network deployment, testing, reliability, and operational efficiency.
  • Design test strategies and lead pre-production validation, network change reviews, and deployment readiness for high-performance network fabrics.
  • Build and enhance telemetry, monitoring, dashboards, and alerting to identify network health, congestion, performance, and reliability issues.
  • Lead incident response, complex troubleshooting, root cause analysis, and corrective actions for network issues impacting OCI AI and GPU workloads.
  • Analyze network performance and capacity, including latency, throughput, packet loss, and congestion, to drive scalable improvements.
  • Partner across OCI Network Engineering, SRE, AI Infrastructure, Data Center Operations, product teams, and vendors to deliver reliable network solutions.
  • Mentor engineers and contribute to network architecture, engineering standards, operational tooling, and continuous improvement.

Job description

Key Responsibilities
  • Lead the design, deployment, validation, and lifecycle management of large-scale RDMA/RoCE network fabrics supporting OCI AI, GPU, and HPC infrastructure.
  • Translate network architectures into scalable designs and deployment plans, ensuring performance, reliability, and operational readiness across OCI’s global cloud environment.
  • Serve as technical lead for complex network initiatives spanning RDMA/RoCE fabrics, data center networking, automation, testing, deployment, and operations.
  • Develop automation frameworks, tools, scripts, and infrastructure pipelines to improve network deployment, testing, reliability, and operational efficiency.
  • Design test strategies and lead pre-production validation, network change reviews, and deployment readiness for high-performance network fabrics.
  • Build and enhance telemetry, monitoring, dashboards, and alerting to identify network health, congestion, performance, and reliability issues.
  • Lead incident response, complex troubleshooting, root cause analysis, and corrective actions for network issues impacting OCI AI and GPU workloads.
  • Analyze network performance and capacity, including latency, throughput, packet loss, and congestion, to drive scalable improvements.
  • Partner across OCI Network Engineering, SRE, AI Infrastructure, Data Center Operations, product teams, and vendors to deliver reliable network solutions.
  • Mentor engineers and contribute to network architecture, engineering standards, operational tooling, and continuous improvement.
  • Preferred Skills & Experience
  • Experience designing, operating, and troubleshooting large-scale RDMA/RoCE, cloud, data center, or high-performance networks.
  • Knowledge of RDMA, RoCE, Ethernet fabrics, congestion control, QoS, and AI/GPU networking.
  • Expertise in routing and switching technologies, including BGP, OSPF, EVPN-VXLAN, and data center networking.
  • Experience with network automation using Python, Ansible, APIs, or similar technologies.
  • Experience with network telemetry, observability, monitoring, performance analysis, and incident management.
  • Experience supporting hyperscale cloud, AI/GPU, HPC, or large-scale distributed infrastructure.
  • Ability to lead complex technical initiatives and collaborate across engineering, operations, customers, and vendors.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Manager, Core Infrastructure Engineering
Senior Manager, Core Infrastructure Engineering

Oracle Corporation • Nashville (TN)

On-site
USD 180,000 - 240,000
Network Developer 4
Network Developer 4

Oracle • Seattle (WA)

On-site
USD 102,300 - 209,500
Medical, dental & vision insurance
401(k) with company match
Paid time off & flexible vacation
Network Design Engineer
Network Design Engineer

Infotree Global Solutions • Austin (TX)

On-site
USD 140,000 - 190,000
Principal Network Engineer - AI Infra (OCI)
Principal Network Engineer - AI Infra (OCI)

Netrolynx AI • United States

On-site
USD 180,000 - 240,000
Medical, Dental, Vision insurance
401(k) with company match
Paid time off and holidays
+1
Senior Network Architect – RDMA & Cloud Scale
Senior Network Architect – RDMA & Cloud Scale

Emerson • Nashville (TN)

On-site
USD 102,000 - 210,000
Health insurance
Dental insurance
Vision insurance
+7
Lead RDMA Network Architect – AI Cloud
Lead RDMA Network Architect – AI Cloud

Ll Oefentherapie • Nashville (TN)

On-site
USD 150,000 - 190,000
Senior Principal Engineer - AI Networking
Senior Principal Engineer - AI Networking

Ll Oefentherapie • Seattle (WA)

On-site
USD 120,000 - 160,000
Senior Manager, RDMA Fabric Design & Engineering
Senior Manager, RDMA Fabric Design & Engineering

Oracle • United States

On-site
USD 133,100 - 306,400
Medical insurance
401(k) Matching
Stock Purchase Plan
+1
Principal Software Engineer
Principal Software Engineer

Lambda • San Francisco (CA)

On-site
USD 220,000 - 300,000
Health coverage
Dental coverage
Vision coverage
+6
Cloud Network Engineer - RDMA, AI/HPC & Automation
Cloud Network Engineer - RDMA, AI/HPC & Automation

Ll Oefentherapie • Austin (TX)

Hybrid
USD 102,000 - 210,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off