SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer Technologies Group

Austin (TX)

On-site

USD 60,000 - 90,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bitdeer Technologies Group in Austin is hiring for an L1 NOC operator focused on front-line monitoring and incident response for NeoCloud's US GPU data centers during the 8AM–8PM PST shift. You will execute SOPs, escalate as needed, perform hardware triage, and feed the AIOps platform with ground truth to train automations.

This role emphasizes hands-on operational tasks and structured, repeatable processes.

Qualifications

  • 2+ years in NOC, data center ops, or IT support.
  • Basic Linux administration (CLI, logs, services).
  • Familiar with monitoring tools: Prometheus, Grafana, Nagios.
  • Experience with ticketing systems: ServiceNow, Jira.
  • Ability to perform physical DC tasks: rack/stack, cabling.
  • Strong communication for shift handoffs and escalation.
  • Willingness to work 12-hour 8AM-8PM PST shifts with rotation.
  • Curiosity about automation and process improvement.
  • Appreciation for structured data and proper ticketing.

Responsibilities

  • Monitor GPU cluster health, network, storage, and environmental sensors.
  • Respond to alerts and runbooks for incidents.
  • Perform hardware triage and identify failed components.
  • Execute remediation steps: resets, drains, reboots.
  • Collect diagnostic data for escalation.
  • Manage incident tickets from creation to resolution.
  • Perform on-site DC tasks: cable installation, hardware swaps.
  • Handoff notes during shifts.
  • Update operational runbooks based on recurring issues.
  • Assist with hardware deployment and inventory under SME guidance.

Skills

NOC operations
Incident response
Communication skills
Automation curiosity
Structured data handling

Tools

Prometheus
Grafana
Nagios
ServiceNow
Jira Service Management

Job description

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit https://ir.bitdeer.com/

Position Overview

You are the first human in the loop — the escalation target when the AIOps system needs a decision, and the source of ground truth that turns novel incidents into new automations.

NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer humans — it means humans focus on judgment calls the platform can't yet make, and every judgment call trains the platform to do it next time. In this L1 role you cover front-line monitoring and incident response for NeoCloud's US GPU DCs during the 8AM–8PM PST shift. You execute SOPs, upscale the hard cases, and feed the AIOps substrate the ground truth it needs to learn from novel incidents.

What you'll own
  • Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards.
  • Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts.
  • Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection.
  • Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery.
  • Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports.
  • Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira).
  • Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles).
  • Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team.
  • Maintain and update operational runbooks based on recurring issues.
  • Assist with hardware deployment, firmware updates, and inventory management under SME guidance.
Feed the AIOps substrate
  • Every novel incident you resolve is data the platform team needs — you tag it, describe it, and hand it back so it becomes an automation.
  • Every runbook you touch should get closer to being executable by the platform, not by you.
  • Your handoff notes are structured signal, not free-form email.
Why this role is different from a NOC job
  • You are not the last line of defense — the platform is. You are the training signal.
  • Growth path is real: strong L1s here move into SME roles, or into the platform team as automation authors.
Job Requirement:
  • 2+ years in NOC, data center operations, or IT support role
  • Basic Linux system administration (command line, log analysis, service management)
  • Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent)
  • Experience with ticketing systems (ServiceNow, Jira Service Management)
  • Ability to perform physical data center tasks: rack and stack, cabling, hardware replacement
  • Strong communication skills for shift handoffs, incident documentation, and escalation
  • Ability to work 8AM-8PM PST shift schedule (12-hour shifts with rotation)
  • Curiosity about automation — you don't just execute the runbook, you notice when it's the third time this month and ask what should change.
  • Comfort with structured data — you understand that how you file a ticket matters, because it may train a model that decides how the next one is filed.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE L1 Support/Cloud Platform Ops Engineer
SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 70,000 - 100,000
SRE L1 SupportCloud Platform Ops Engineer
SRE L1 SupportCloud Platform Ops Engineer

Bitdeer Technologies Group • United States

On-site
USD 60,000 - 90,000
SRE L1 Support/Cloud Platform Ops Engineer
SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 60,000 - 90,000
SRE L1 Support/Cloud Platform Ops Engineer
SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer Group • San Jose (CA), Austin (TX)

Hybrid
USD 60,000 - 90,000
SRE L1 Support/Cloud Platform Ops Engineer
SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer • San Jose (CA)

On-site
USD 65,000 - 95,000
Sr. SRE Platform Software Engineer
Sr. SRE Platform Software Engineer

Bitdeer • San Jose (CA)

On-site
USD 130,000 - 160,000
Sr. SRE Platform Software Engineer
Sr. SRE Platform Software Engineer

Bitdeer Group • San Jose (CA)

On-site
USD 120,000 - 160,000
Cloud Service Security Platform DevOps & Maintenance
Cloud Service Security Platform DevOps & Maintenance

Bitdeer • San Jose (CA)

On-site
USD 140,000 - 210,000
Cloud Service Security Platform DevOps & Maintenance
Cloud Service Security Platform DevOps & Maintenance

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 190,000
K8 Site Reliability SME
K8 Site Reliability SME

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 180,000 - 240,000