NOC Engineer: AI GPU Cloud & HPC Operations

Cirrascale Cloud Services

Austin (TX)

On-site

USD 80,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cirrascale Cloud Services, a provider of high-performance cloud infrastructure for AI workloads, seeks a skilled support engineer to triage alerts, troubleshoot GPU nodes, and coordinate incident responses across global data centers. You will engage with customers, document configurations, and help optimize NOC procedures.

We value 2-4 years in HPC/AI infrastructure, scripting ability, and strong communication.

Qualifications

  • 2-4 years of experience in HPC, AI infrastructure, cloud systems or related fields.
  • Scripting (Python, Bash) and GPU/resource monitoring.
  • Solid understanding of HPC datacenter networking and troubleshooting.
  • Strong analytical and problem-solving skills; ability to work independently.
  • Excellent communication and customer service skills.
  • Certifications: Advanced Linux or AI/ML certifications are a plus.
  • Experience in Datacenter Network Operations.
  • Familiarity with Jira ticketing system; RMAs/logistics is a plus.
  • Proficiency with Microsoft 365 tools.

Responsibilities

  • Respond to alerts and incidents for HPC infrastructure and GPU nodes.
  • Understand deployment, networking, and clustering in datacenters.
  • Triage tickets and perform basic troubleshooting via Jira.
  • Monitor and troubleshoot nodes and network equipment.
  • Remote troubleshooting of servers and GPUs across data centers.
  • Lead major incident response and coordination when needed.
  • Review and optimize NOC procedures and capacity planning.
  • Document configurations, updates, and asset inventory.

Skills

HPC infrastructure
Python/Bash scripting
Datacenter networking
Analytical problem-solving
Customer service
Communication skills

Education

Advanced Linux certification
AI/ML certification

Tools

Jira
Microsoft 365

Job description

Cirrascale Cloud Services, a provider of high-performance cloud infrastructure for AI workloads, seeks a skilled support engineer to triage alerts, troubleshoot GPU nodes, and coordinate incident responses across global data centers. You will engage with customers, document configurations, and help optimize NOC procedures.

We value 2-4 years in HPC/AI infrastructure, scripting ability, and strong communication.

Get your free, confidential resume review.
or drag and drop your file here.