Network Engineer - Data Center

Xai

Memphis (TN)

On-site

USD 110,000 - 170,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

SpaceXAI seeks a seasoned networks engineer to support AI data centers and high-density training clusters. You'll design highly available networks with spine-leaf fabrics, manage data center and campus networks, and contribute to automation and documentation.

Strong urgency and operational excellence are expected in this role. The ideal candidate has extensive Layer 2/3 expertise in latency-sensitive environments, experience with multiple vendors, and proficiency in scripting and IaC.

Qualifications

  • Bachelor’s degree in computer science, computer engineering, or STEM; 3+ years in network engineering, or 5+ years without a degree.
  • Hands-on experience designing, deploying, and troubleshooting Layer 2/3 networks in latency-sensitive environments.
  • Experience with multiple network vendors in production or lab environments.
  • Experience with GitOps and Infrastructure as Code frameworks.

Responsibilities

  • Design and implement highly available, high-bandwidth networks for AI data centers with spine-leaf/Clos fabrics.
  • Maintain data center and campus networks per company standards; collaborate with cross-functional teams.
  • Evaluate, procure, and deploy network hardware including high-speed switches, optics, and firewalls.
  • Advance network automation tooling; add configuration analysis and scalability to deployment framework.
  • Plan maintenance windows for software updates and hardware refreshes in live production.
  • Troubleshoot network issues; publish RCA documentation and conduct retrospective reviews.
  • Provide direct networking support during cluster bring-up and high-load operations; participate in on-call.
  • Tailor monitoring to detect congestion and loss; prevent impact on workloads.
  • Create and update network documentation including diagrams and procedures.
  • Identify and resolve systemic failure modes and insufficient redundancy with cross-functional teams.
  • Gather requirements and create implementation plans for new halls and expansions.
  • Ensure networks comply with cybersecurity standards and segment compute, management, and OT networks.

Skills

Layer 2/3 networks
GitOps
Infrastructure as Code
Scripting
Linux administration
CCNA/CCNP

Education

Bachelor’s degree in CS/engineering
5+ years experience in lieu of a degree

Tools

Cisco switches
Arista switches
Juniper switches
NVIDIA Spectrum

Job description

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

SpaceXAI is looking for an exceptional network engineer with experience in mission-critical and large-scale production environments to support high-density AI training and inference clusters as well as high-reliability data center networks. As a member of the Data Center Network Engineering team, you will provide operational and design services for networks used by compute clusters, automation & controls engineering, and facilities teams. The ideal candidate thrives in intense, high-flux environments, brings a strong sense of urgency balanced with operational excellence, communicates effectively, and demonstrates high levels of technical acumen.

RESPONSIBILITIES
  • Design and implement highly available, high-bandwidth networks for AI data centers, carefully balancing routing, overlay, and redundancy technologies (including spine-leaf / Clos fabrics) to meet unique requirements of GPU training clusters, inference, storage, and management planes.
  • Design and maintain data center and campus networks in accordance with company network standards. Collaborate with adjacent compute, infrastructure, facilities, and enterprise teams.
  • Evaluate, procure, and deploy network hardware including high-speed switches, optics, firewalls, multiplexing, and related appliances.
  • Contribute to ever-maturing network automation tooling; implementing configuration analysis, linting, and scalability into the deployment framework.
  • Plan and coordinate network maintenance windows with stakeholders to perform software updates, hardware refreshes, and general network work (sometimes on weekends and evenings) in live production environments.
  • Troubleshoot and resolve network-related issues, publishing root cause analysis (RCA) documentation and hosting retrospective reviews.
  • Provide direct networking support during cluster bring-up, capacity expansions, and high-load operations; participate in on-call or serve as networking responsible engineer during critical events.
  • Proactively tailor network monitoring and telemetry to detect congestion, packet loss, and other issues before they impact training or inference workloads.
  • Continuously create and update network documentation, including architecture overviews, design drawings, and operational procedures.
  • Collaborate with cross-functional teams to proactively identify and resolve potential technical issues with network designs, especially systemic and cascading failure modes and insufficient redundancy.
  • Perform job walks with customers, vendors, and contractors to gather network and connectivity requirements and create implementation plans for new halls and expansions.
  • Ensure networks are configured and maintained in compliance with industry and cybersecurity standards, with particular attention to segmentation between compute fabrics, management, and facilities/OT networks.
BASIC QUALIFICATIONS
  • Bachelor’s degree in computer science, computer engineering, or other STEM discipline and 3+ years of professional network engineering experience;
    • OR 5+ years of professional network engineering experience in lieu of a degree.
  • Extensive hands-on experience designing, deploying, supporting, and troubleshooting Layer 2 and Layer 3 networks in latency-sensitive and/or industrial / high-reliability environments.
  • Functional experience with multiple network vendors either in production or lab environments.
  • Experience with GitOps and Infrastructure as Code frameworks, both as user and contributor.
PREFERRED SKILLS AND EXPERIENCE
  • Strong understanding of the OSI model and network standards.
  • Hands-on experience with data center class switches (Cisco, Arista, Juniper, NVIDIA Spectrum or similar).
  • Experience with high-speed Ethernet fabrics for AI/HPC (RoCEv2, PFC/ECN, lossless or near-lossless designs), EVPN/VXLAN, and related overlay technologies.
  • Experience with Wave Division Multiplexing (WDM) and fiber optics infrastructure.
  • Extensive knowledge of common enterprise, data center, and media protocols.
  • Experience with switch port security features, network segmentation, QoS, multicast, and redundancy protocols.
  • Familiarity with network monitoring tools and Layer 1 testing tools.
  • Proficiency in scripting (Bash/PowerShell/Python) and automation frameworks (Terraform, Ansible, etc.).
  • Linux and Windows system administration experience professionally or from labs.
  • Industry standard certifications such as CCNA or CCNP.
  • Experience supporting real-time systems, industrial control networks, or operational technology (OT) environments in data center, energy, or similar high-reliability industries.
  • Excellent communication skills with the ability to communicate with internal/external customers, vendors, and management in both formal and informal situations.
ADDITIONAL REQUIREMENTS
  • Ability to work in data center environments (hot/cold aisles, raised floor or slab, cable trays).
  • Ability to work in tight quarters; physical dexterity is necessary to perform job functions.
  • Availability for extended hours and/or weekends as the schedule varies depending on site operational needs; flexibility is required.
  • Ability to provide 24x7 on-call support in emergency situations and participate in an after-hours on-call rotation.
  • Willingness to travel (up to 20–30%, including other SpaceXAI data center sites) as needed.
  • Ability to lift 30 lbs.
  • Ability to work at heights (e.g., cable trays, ladders).
  • Ability to drive (active valid driver’s license).

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Engineer (Supercomputer Infrastructure)
Network Engineer (Supercomputer Infrastructure)

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 160,000
Network Engineer (Supercomputer Infrastructure) - Memphis
Network Engineer (Supercomputer Infrastructure) - Memphis

Pantera Capital • Southaven (MS)

On-site
USD 120,000 - 180,000
Network Engineer (Supercomputer Infrastructure) - Memphis
Network Engineer (Supercomputer Infrastructure) - Memphis

SpaceXAI • Memphis (TN)

On-site
USD 120,000 - 190,000
Network Engineer (Supercomputer Infrastructure) - Memphis
Network Engineer (Supercomputer Infrastructure) - Memphis

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 180,000
Network Engineer (Supercomputer Infrastructure) - Memphis
Network Engineer (Supercomputer Infrastructure) - Memphis

xAI • Memphis (TN)

Hybrid
USD 120,000 - 190,000
Network Engineer
Network Engineer

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 250,000
Equity
Comprehensive medical coverage
Vision coverage
+5
Network Engineer (Supercomputer Infrastructure)
Network Engineer (Supercomputer Infrastructure)

Spacex • Memphis (TN), Northern (KY)

Hybrid
USD 120,000 - 160,000
Network Connectivity Engineer - Memphis
Network Connectivity Engineer - Memphis

Xai • Southaven (MS)

On-site
USD 85,000 - 130,000
Network Connectivity Engineer - Memphis
Network Connectivity Engineer - Memphis

Pantera Capital • Southaven (MS)

On-site
USD 120,000 - 160,000
Network Engineer
Network Engineer

SpaceXAI • Palo Alto (CA)

On-site
USD 150,000 - 250,000
Equity
Medical insurance
Vision
+5