SRE L1: GPU Cloud Platform Ops & Incident Response

Bitdeer Technologies Group

United States

On-site

USD 60,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bitdeer Technologies Group seeks an L1 NOC engineer to monitor US GPU data centers and act as the escalation point for incidents. You will execute SOPs, triage hardware, and feed the platform with ground truth to improve automation.

Work primarily on-site, handling incident tickets, performing physical tasks, and collaborating with the APAC team. Strong communication and Linux skills are essential for day-to-day operations.

Qualifications

  • 2+ years in NOC, data center operations, or IT support.
  • Basic Linux system administration and log analysis.
  • Familiarity with monitoring tools and ticketing systems.
  • Ability to perform physical data center tasks and handoffs.
  • Strong communication and structured data handling skills.

Responsibilities

  • Monitor GPU cluster health, networks, storage, and environment.
  • Respond to alerts and execute runbooks for incidents.
  • Identify failed hardware during triage and perform remediation.
  • Manage incident tickets from creation to resolution.
  • Update runbooks and contribute to automation data.

Skills

NOC experience
Linux basics
Monitoring tools
Ticketing systems
Physical data center tasks
Strong communication
Shift work (8AM-8PM PST)
Automation curiosity
Structured data handling

Tools

Prometheus
Grafana
Nagios
ServiceNow
Jira Service Management

Job description

Bitdeer Technologies Group seeks an L1 NOC engineer to monitor US GPU data centers and act as the escalation point for incidents. You will execute SOPs, triage hardware, and feed the platform with ground truth to improve automation.

Work primarily on-site, handling incident tickets, performing physical tasks, and collaborating with the APAC team. Strong communication and Linux skills are essential for day-to-day operations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Cloud Ops SRE (L1) - Incident Response & Automation
GPU Cloud Ops SRE (L1) - Incident Response & Automation

Bitdeer • San Jose (CA)

On-site
USD 65,000 - 95,000
GPU Cloud Platform Ops (L1) - Monitoring & Automation
GPU Cloud Platform Ops (L1) - Monitoring & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 60,000 - 90,000
AI GPU Cloud Ops Engineer - L1 Support
AI GPU Cloud Ops Engineer - L1 Support

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 70,000 - 100,000
SRE L1 Support/Cloud Platform Ops Engineer
SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer • San Jose (CA)

On-site
USD 65,000 - 95,000
SRE L1 SupportCloud Platform Ops Engineer
SRE L1 SupportCloud Platform Ops Engineer

Bitdeer Technologies Group • United States

On-site
USD 60,000 - 90,000
SRE L1 Support/Cloud Platform Ops Engineer
SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 70,000 - 100,000
SRE L1 Support/Cloud Platform Ops Engineer
SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 60,000 - 90,000
Senior SRE & Automation Engineer — GPU Cloud Reliability
Senior SRE & Automation Engineer — GPU Cloud Reliability

Bitdeer Group • San Jose (CA), Austin (TX)

Hybrid
USD 180,000 - 240,000
Remote NOC Engineer - GPU Clusters & Incident Response
Remote NOC Engineer - GPU Clusters & Incident Response

REALM • United States

On-site
USD 70,000 - 110,000
Junior SRE: Monitoring Platform Engineer (GPU Cloud)
Junior SRE: Monitoring Platform Engineer (GPU Cloud)

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 90,000 - 130,000