Network Operations Center Specialist

Socket.dev

Southaven (MS)

On-site

USD 60,000 - 80,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

SpaceXAI is seeking a Network Operations Center (NOC) Specialist to monitor campus health signals around the clock and coordinate rapid incident response. You will be the eyes and voice of the campus, ensuring timely updates, proper escalation, and durable incident timelines.

The role focuses on communications, judgment, and collaboration with SRE, SiteOps, and Facilities; it does not involve hands-on engineering work.

Qualifications

  • Experience in a 24/7 operations environment (NOC, SOC, dispatch, or equivalent).
  • Ability to acknowledge, classify, and elevate incidents under SLA in a high-signal environment.
  • Experience opening and running incident bridges with fixed cadence updates.
  • Excellent written and verbal communication during incidents.
  • Pattern recognition across compute, network, storage, or facilities signals.
  • Experience maintaining runbooks, escalation matrices, and handoffs.
  • Willingness to work rotating shifts including nights and weekends.

Responsibilities

  • Staff the console per shift schedule and monitor cluster health, node availability, network health, facility trends, storage alarms and thresholds.
  • Acknowledge, classify and log disposition; feed noise patterns back to SRE for signal quality.
  • Detect, verify and elevate within time budgets; follow escalation matrix and page correctly the first time.
  • Open and run incident bridges; manage updates on a fixed cadence and live timeline hygiene.
  • Produce first-pass RCA framing and hand to SRE/Hardware Failure Analysis; root cause not published by NOC.
  • Perform structured shift handoffs and maintain cross-site awareness.
  • Write major-incident reports; open corrective Linear projects and drive to closure.
  • Maintain and improve NOC runbooks, escalation matrices, and communications templates; participate in SRE game days.

Skills

24/7 operations
Incident management
SLA adherence
Communication skills
Runbooks & processes
Shift work
Cross-domain pattern recognition
Linear tooling

Tools

Linear

Job description

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As a Network Operations Center (NOC) Specialist, you are the eyes and the voice of the campus — never the hands. You watch campus health signals around the clock, detect and verify site-impacting events, assemble the right responders fast, and run incident communications leadership can trust. You make sure no major incident closes without a timeline, a report, and a tracked corrective project. This is a communications-and-judgment role at the center of site operations, not a junior-engineering holding pen. You do not do wrench work, plant operation, deep root-cause analysis, monitoring design, technical SEV command, or tool building — those belong to SiteOps, Facilities, Hardware Failure Analysis, Site SRE, and Software Platforms.

RESPONSIBILITIES:
  • Staff the console per shift schedule and watch the designated signal surface: cluster health, node availability, network health, facility trend panels, storage alarms, and threshold breaches.
  • Acknowledge every page within SLA; classify (actionable / known / noise) and log disposition; feed noise patterns back to SRE so signal quality keeps improving.
  • Detect, verify, and elevate within time budgets; operate the escalation matrix (NOC - on-call SRE - domain owners) and page correctly the first time.
  • Open and run incident bridges; own stakeholder communications (first update within SLA, then fixed cadence); maintain the incident timeline in real time; call out ownership stalls.
  • Produce first-pass RCA framing (what happened, when, what's impacted, who's engaged) and hand it to SRE / Hardware Failure Analysis for depth - the NOC does not publish root cause.
  • Run structured shift handoffs and durable shift logs; maintain cross-site awareness.
  • Write major-incident reports; open corrective projects in Linear and chase them to closure - the NOC is the nag of record.
  • Maintain and continuously improve NOC runbooks, escalation matrices, and communications templates; participate in SRE-run game days.
BASIC QUALIFICATIONS:
  • Experience in a 24/7 operations environment (NOC, SOC, dispatch, mission control, or equivalent).
  • Proven ability to acknowledge, classify, and elevate incidents under SLA in a high-signal environment.
  • Experience opening and running incident bridges, including stakeholder updates on a fixed cadence and live timeline hygiene.
  • Excellent written and verbal communication skills; able to write clear updates while an incident is in progress.
  • Demonstrated pattern recognition across multiple domains (compute, network, storage, and/or facilities signals) and curiosity about how those systems interact.
  • Experience following, maintaining, and improving operational process (runbooks, escalation matrices, handoffs, or similar).
  • Willingness and ability to work a rotating shift schedule, including nights and weekends, as part of continuous campus coverage.
PREFERRED SKILLS AND EXPERIENCE:
  • Prior NOC, data center operations, or campus reliability experience in a high-performance computing, AI/ML infrastructure, or large-scale production environment.
  • Experience writing major-incident reports and driving corrective follow-ups to closed (e.g. tickets, projects, or Linear).
  • Familiarity with Linear or similar work-tracking tools for corrective action programs.
  • Experience partnering with SRE, SiteOps, and Facilities on escalations and post-incident follow-through.
  • Participation in game days, tabletop exercises, or runbook improvement programs.
  • Prior work in a fast-paced startup or tech company like SpaceXAI.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Operations Center Specialist - Memphis
Network Operations Center Specialist - Memphis

Pantera Capital • Southaven (MS)

On-site
USD 70,000 - 100,000
Network Operations Center Specialist
Network Operations Center Specialist

SpaceXAI • Southaven (MS)

On-site
USD 90,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

Socket.dev • Southaven (MS)

On-site
USD 120,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Pantera Capital • Southaven (MS)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Spacex • Memphis (TN), Northern (KY)

Hybrid
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 170,000
Site Reliability Engineer - Memphis
Site Reliability Engineer - Memphis

Pantera Capital • Southaven (MS)

On-site
USD 130,000 - 180,000
Software Engineer - Datacenter
Software Engineer - Datacenter

SpaceXAI • Memphis (TN)

On-site
USD 120,000 - 160,000
Network Engineer (Supercomputer Infrastructure)
Network Engineer (Supercomputer Infrastructure)

SpaceXAI • Southaven (MS)

On-site
USD 120,000 - 160,000
Network Engineer (Supercomputer Infrastructure)
Network Engineer (Supercomputer Infrastructure)

Socket.dev • Southaven (MS)

On-site
USD 140,000 - 200,000