NOC Engineer

Axe Compute

Miami (FL)

On-site

USD 65,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Axe Compute is seeking NOC Engineers who own the customer experience for live GPU cluster deployments, using automation, AI tools, and continuous process improvement to drive SLA performance rather than merely monitor it.

You will build tools and runbooks, write and improve SOPs, and treat every shift as an opportunity to make the NOC better than you found it, coordinating with data center operators and vendors to resolve client-impacting issues.

Qualifications

  • Experience building automation for monitoring and incident response.
  • Interest or experience in automation or tooling beyond following procedures.
  • Willingness to work 24/7 shifts and on-call rotations.

Responsibilities

  • Monitor live GPU clusters, health, power, cooling, and network status.
  • Triage and resolve incidents within SLA, escalating to Tier 3 when needed.
  • Create and improve runbooks, SOPs, and automation to reduce toil.

Skills

Automation
Scripting
Datadog
Grafana
PagerDuty
Incident response

Tools

Datadog
Grafana
PagerDuty

Job description

Axe Compute is seeking NOC Engineers who do more than watch a screen. This role owns customer experience for live GPU cluster deployments, using automation, AI tools, and continuous process improvement to drive SLA performance, not just monitor it. We're looking for people who build tools, write and improve their own runbooks and SOPs, and treat every shift as an opportunity to make the NOC better than they found it.

ROLE AT A GLANCE
  • Mandate: Monitor live GPU clusters, drive incident response, and continuously improve the tools, processes, and SOPs the NOC runs on.
  • Scope: 24/7 shift-based monitoring and incident response, tool-building and automation, runbook/SOP creation and maintenance, third-party coordination, SLA metric improvement.
  • Key Outcomes: SLA compliance, a measurably improving NOC (fewer repeat incidents, faster resolution, better tooling), strong customer experience outcomes.
WHAT YOU WILL OWN
  • Monitor cluster health, power, cooling, and network status across live deployments, triaging and resolving incidents within SLA.
  • Escalate to Tier 3 (OEM or data center operator) only when an issue genuinely requires vendor-level support.
Tooling, Automation & AI Workflows
  • Build and improve internal tools and automation to reduce manual, repetitive work, using AI-enabled workflows wherever they make the team faster or more accurate.
  • Identify opportunities to automate recurring monitoring, triage, or reporting tasks rather than performing them manually shift after shift.
Runbooks, SOPs & Process Improvement
  • Create and maintain runbooks and standard operating procedures, updating them based on real incidents rather than leaving them static.
  • Contribute to after-action reviews and turn recurring issues into permanent process or tooling fixes.
Customer Experience Ownership
  • Own the customer experience of every incident, communicating clearly and proactively rather than treating tickets as a checklist.
  • Work directly with third parties (data center operators, OEMs, network providers) as needed to resolve client-impacting issues.
  • Track and actively work to improve SLA metrics, using data and software tooling to identify where performance is slipping before it becomes a client-facing problem.
REQUIRED QUALIFICATIONS
  • 1+ years in a NOC, network operations, or infrastructure monitoring role. This role can be junior, but not passive.
  • Demonstrated interest or experience in automation, scripting, or tool-building, not just following existing procedures.
  • Comfort working rotating shifts, including nights/weekends, as part of a 24/7 coverage model.
  • Familiarity with monitoring tools (e.g., Datadog, Grafana, PagerDuty) and a genuine curiosity about AI-enabled operations tools.
PREFERRED QUALIFICATIONS
  • Experience monitoring GPU or high-performance computing infrastructure.
  • Scripting ability (Python, Bash, or similar) to build or modify automation.
  • Additional languages beyond English are a plus for coordinating with clients and global vendors.
  • Multiple shift options available to accommodate different timezones and flexible working hours.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

NOC Engineer - AI-Driven GPU Incident Response
NOC Engineer - AI-Driven GPU Incident Response

Axe Compute • Miami (FL)

On-site
USD 65,000 - 90,000
NOC Technician (Data Center and Site Ops)
NOC Technician (Data Center and Site Ops)

Socket.dev • Memphis (TN)

On-site
USD 42,000 - 66,000
NOC Technician (Data Center and Site Ops)
NOC Technician (Data Center and Site Ops)

Pantera Capital • Memphis (TN)

On-site
USD 48,000 - 64,000
Network Operations Center Technician II
Network Operations Center Technician II

Cirrascale Corporation • Austin (TX)

On-site
USD 55,000 - 90,000
Senior Director, Customer Success
Senior Director, Customer Success

Axe Compute • Miami (FL)

On-site
USD 180,000 - 240,000
NOC Technician (Data Center and Site Ops)
NOC Technician (Data Center and Site Ops)

SpaceXAI • Memphis (TN)

On-site
USD 52,000 - 72,000
Remote NOC Engineer - GPU Clusters & Incident Response
Remote NOC Engineer - GPU Clusters & Incident Response

REALM • United States

On-site
USD 70,000 - 110,000
Network Operations Center Technician II
Network Operations Center Technician II

Cirrascale Cloud Services • Austin (TX)

On-site
USD 80,000 - 110,000
NOC Engineer
NOC Engineer

TechDigital Group • United States

On-site
USD 70,000 - 90,000
NOC Technician (Data Center and Site Ops)
NOC Technician (Data Center and Site Ops)

AI Need That • Memphis (TN), Northern (KY)

Hybrid
USD 48,000 - 72,000