Remote NOC Engineer - GPU Clusters & Incident Response

REALM

United States

On-site

USD 70,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

REALM is seeking a NOC Engineer to support live GPU clusters in a remote setup across US or Southeast Asia. You’ll monitor health, drive incident response, and build automation to reduce manual work.

You’ll own runbooks, communicate proactively during incidents, and work across time zones with a 24/7 coverage model. Prior NOC or infrastructure monitoring experience is required, with a passion for automation and reliability.

Qualifications

  • 1+ years in a NOC, network operations, or infrastructure monitoring
  • Interest or experience in automation and scripting
  • Comfortable with rotating shifts including nights and weekends for 24/7 coverage

Responsibilities

  • Monitor live GPU clusters and triage incidents within SLA
  • Escalate to Tier 3 when vendor-level issues arise
  • Build and improve internal tools and automation to reduce manual work
  • Write and maintain runbooks and SOPs based on real incidents
  • Own the customer experience of every incident with proactive communication

Skills

Automation and scripting
Observability tooling
GPU/HPC monitoring awareness

Tools

Datadog
Grafana
PagerDuty

Job description

REALM is seeking a NOC Engineer to support live GPU clusters in a remote setup across US or Southeast Asia. You’ll monitor health, drive incident response, and build automation to reduce manual work.

You’ll own runbooks, communicate proactively during incidents, and work across time zones with a 24/7 coverage model. Prior NOC or infrastructure monitoring experience is required, with a passion for automation and reliability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote NOC Engineer — Automate Incidents & AI Ops
Remote NOC Engineer — Automate Incidents & AI Ops

REALM • United States

On-site
USD 75,000 - 140,000
NOC Engineer - AI-Driven GPU Incident Response
NOC Engineer - AI-Driven GPU Incident Response

Axe Compute • Miami (FL)

On-site
USD 65,000 - 90,000
NOC Engineer
NOC Engineer

Axe Compute • Miami (FL)

On-site
USD 65,000 - 90,000
GPU Cloud Ops SRE (L1) - Incident Response & Automation
GPU Cloud Ops SRE (L1) - Incident Response & Automation

Bitdeer • San Jose (CA)

On-site
USD 65,000 - 95,000
SRE L1: GPU Cloud Platform Ops & Incident Response
SRE L1: GPU Cloud Platform Ops & Incident Response

Bitdeer Technologies Group • United States

On-site
USD 60,000 - 90,000
GPU Cloud Platform Ops (L1) - Monitoring & Automation
GPU Cloud Platform Ops (L1) - Monitoring & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 60,000 - 90,000
Network Operations Center Technician II
Network Operations Center Technician II

Cirrascale Corporation • Austin (TX)

On-site
USD 55,000 - 90,000
Remote Senior GPU Network Engineer – Fabric & Cluster Builds
Remote Senior GPU Network Engineer – Fabric & Cluster Builds

REALM • United States

On-site
USD 170,000 - 230,000
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)

Rune • Mountain View (CA)

Hybrid
USD 175,000 - 260,000
AI GPU Cloud Ops Engineer - L1 Support
AI GPU Cloud Ops Engineer - L1 Support

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 70,000 - 100,000