Senior Systems Engineer- Network Infrastructure

Nscale

United States

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading tech startup in the United States is seeking a Senior Deployment Engineer to lead the deployment of network clusters in its data center. The ideal candidate will have 5-8 years of experience in infrastructure engineering and a hands-on background in high-speed networking and server deployment. This role requires technical expertise in ensuring clusters are production-ready, with responsibilities including integration, validation, and automation. Join a mission-driven team dedicated to building cutting-edge AI infrastructure.

Qualifications

  • 5–8+ years in infrastructure engineering or data center operations.
  • Hands-on experience deploying network servers.
  • Experience with InfiniBand and Ethernet fabrics.

Responsibilities

  • Execute end-to-end bringup of network nodes and racks.
  • Validate BIOS/BMC/firmware configurations.
  • Perform rack-level integration and airflow validation.

Skills

Infrastructure engineering
Hardware deployment
High-speed networking
Troubleshooting distributed systems
Automation (Python, Ansible, Terraform, Bash)

Job description

About Us

We are building next-generation AI infrastructure from the ground up. Our mission is to deliver highly performant, reliable, and scalable network clusters purpose-built for large-scale AI training and inference.

As a startup, we operate with urgency, ownership, and a bias toward action. We are assembling the foundational infrastructure that will power frontier AI workloads—and we’re looking for engineers who want to build it from zero to scale.

The Role

We are hiring a Senior Deployment Engineer to lead hands‑on bringup of network clusters across our data center environments. You will own the execution of node, rack, and network deployment, ensuring clusters are validated, performant, and production‑ready.

This role is deeply technical and execution‑focused. You will be in the details—cabling racks, validating firmware, tuning fabrics, debugging performance—and helping us build repeatable processes as we scale.

What You’ll Do
  • Execute end‑to‑end bringup of network nodes and racks from installation to production readiness.
  • Validate BIOS/BMC/firmware configurations and network health.
  • Perform rack‑level integration including power, cabling, and airflow validation.
  • Bring up and validate high‑speed network fabrics (InfiniBand, RoCE, 100–400G Ethernet).
Network & Performance Validation
  • Configure and validate leaf/spine network connectivity.
  • Run cluster‑wide burn‑in and stress testing.
  • Validate node‑to‑node performance (NCCL, RDMA, GPUDirect).
  • Troubleshoot hardware, firmware, and fabric‑level issues.
Automation & Process
  • Contribute to automation for provisioning and cluster validation.
  • Improve deployment playbooks and documentation.
  • Identify reliability issues early and drive corrective actions.
  • Help turn ad‑hoc deployments into repeatable systems.
Cross-Functional Collaboration
  • Work closely with networking, systems software, and data center teams.
  • Coordinate with hardware vendors to resolve bringup issues.
  • Support rapid capacity expansion as we scale.
What We’re Looking For
Required
  • 5–8+ years in infrastructure engineering, hardware deployment, or data center operations.
  • Hands‑on experience deploying network servers (HGX/DGX or similar platforms).
  • Experience with high‑speed networking (InfiniBand, RoCE, Ethernet fabrics).
  • Experience troubleshooting distributed systems performance issues.
  • Comfortable working onsite in data center environments as needed.
Strongly Preferred
  • Experience in AI/ML infrastructure or HPC environments.
  • Familiarity with NCCL, CUDA, RDMA.
  • Automation experience (Python, Ansible, Terraform, Bash).
  • Experience in high‑density power and cooling environments.
What Success Looks Like
  • Clusters are brought online quickly and correctly.
  • Performance baselines meet or exceed expectations.
  • Deployment processes become faster and more reliable over time.
  • You help build the foundation for scaled infrastructure growth.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Network Architect – AI Infrastructure
Senior Network Architect – AI Infrastructure

Nscale • Seattle (WA)

On-site
USD 120,000 - 160,000
Senior Network Architect – AI Infrastructure
Senior Network Architect – AI Infrastructure

Nscale • United States

On-site
USD 120,000 - 150,000
Senior Network Engineer
Senior Network Engineer

Nscale • Seattle (WA)

On-site
USD 150,000 - 210,000
Principal Network Engineer
Principal Network Engineer

Nscale • Seattle (WA), New York (NY), San Francisco (CA), Houston (TX)

On-site
USD 180,000 - 240,000
Base salary + equity
Equity incentives
Dynamic progression plan
Senior Network Engineer
Senior Network Engineer

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Senior Network Engineer
Senior Network Engineer

Nscale • San Francisco (CA)

On-site
USD 150,000 - 210,000
Medical insurance
Retirement plan
Flexible PTO
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities
Senior Network Engineer
Senior Network Engineer

Nscale • New York (NY)

On-site
USD 150,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+3
Principal Front-End Network Engineer
Principal Front-End Network Engineer

Nscale • San Francisco (CA)

On-site
USD 200,000 - 260,000
Medical, dental, vision coverage
Flexible paid time off
Parental leave
+1
Senior Network & Site Reliability Engineer
Senior Network & Site Reliability Engineer

Alembic Technologies • San Francisco (CA)

On-site
USD 210,000 - 240,000