Network Operations Center Technician II

Cirrascale Cloud Services

Austin (TX)

On-site

USD 80,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cirrascale Cloud Services, a provider of high-performance cloud infrastructure for AI workloads, seeks a skilled support engineer to triage alerts, troubleshoot GPU nodes, and coordinate incident responses across global data centers. You will engage with customers, document configurations, and help optimize NOC procedures.

We value 2-4 years in HPC/AI infrastructure, scripting ability, and strong communication.

Qualifications

  • 2-4 years of experience in HPC, AI infrastructure, cloud systems or related fields.
  • Scripting (Python, Bash) and GPU/resource monitoring.
  • Solid understanding of HPC datacenter networking and troubleshooting.
  • Strong analytical and problem-solving skills; ability to work independently.
  • Excellent communication and customer service skills.
  • Certifications: Advanced Linux or AI/ML certifications are a plus.
  • Experience in Datacenter Network Operations.
  • Familiarity with Jira ticketing system; RMAs/logistics is a plus.
  • Proficiency with Microsoft 365 tools.

Responsibilities

  • Respond to alerts and incidents for HPC infrastructure and GPU nodes.
  • Understand deployment, networking, and clustering in datacenters.
  • Triage tickets and perform basic troubleshooting via Jira.
  • Monitor and troubleshoot nodes and network equipment.
  • Remote troubleshooting of servers and GPUs across data centers.
  • Lead major incident response and coordination when needed.
  • Review and optimize NOC procedures and capacity planning.
  • Document configurations, updates, and asset inventory.

Skills

HPC infrastructure
Python/Bash scripting
Datacenter networking
Analytical problem-solving
Customer service
Communication skills

Education

Advanced Linux certification
AI/ML certification

Tools

Jira
Microsoft 365

Job description

Cirrascale Cloud Services provides high-performance cloud infrastructure purpose-built for deep learning, generative AI, and large-scale AI inference workloads. We specialize in dedicated GPU cloud solutions tailored to the unique needs of startups, research labs, and enterprise AI teams. Our mission is to accelerate AI innovation by combining powerful hardware with white-glove service and flexible, custom-built environments.

Key Responsibilities
  • First-line whit glove response to alerts and incidents to systems and job failures
  • Demonstrate understanding of GPU nodes and how they are deployed, networked, and clustered within a datacenter
  • Assist customers with ticket triage and basic troubleshooting using the Jira (Atlassian) ticketing system
  • Perform high-level monitoring and troubleshooting on all nodes and network equipment
  • Remotely Troubleshoot Installed Servers & GPUs at various global datacenter locations
  • Resolve complex and critical incidents within our datacenters
  • Lead major incident response and coordination
  • Review and optimize existing NOC procedures
  • Work on capacity planning and performance monitoring
  • Knowledge and experience working with Dell, SuperMicro & Lenovo type Servers is highly recommended
  • Perform deep troubleshooting of GPU node failures, job preemption conflicts, and cluster imbalance
  • Analyze alerts for GPU utilization inefficiencies, failed Machine Learning pipelines, or I/O bottlenecks
  • Provide on-site and remote support to resolve urgent technical issues
  • Document system configurations, updates, and inventory, maintaining accurate records of data center assets.
  • Stay current with industry trends, emerging technologies, and best practices in HPC network operations center trends
Qualifications
  • 2-4 years of experience in HPC, AI infrastructure, cloud systems, or related
  • Understanding of scripting (Python, Bash, etc.), GPU resource monitoring preferred
  • Solid understanding of HPC datacenter networking principles and experience with network troubleshooting
  • Strong analytical and problem-solving skills, with the ability to work independently and manage multiple tasks
  • Excellent communication skills and the ability to collaborate effectively with the customer and the team. Customer Service is a must
  • Certifications: Advanced Linux or any other AI/ML certifications are a huge plus
  • Experience in Datacenter Network Operations
  • Experience with RMAs, logistics, shipping, and receiving a plus
  • It is a plus with experience working in Jira (Atlassian) ticketing system.
  • Proficient in Microsoft 365(Outlook, Word, Excel)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Operations Center Technician II
Network Operations Center Technician II

Cirrascale Corporation • Austin (TX)

On-site
USD 55,000 - 90,000
Data Center Technician II
Data Center Technician II

Cirrascale Cloud Services • Austin (TX)

On-site
USD 55,000 - 85,000
NOC Engineer: AI GPU Cloud & HPC Operations
NOC Engineer: AI GPU Cloud & HPC Operations

Cirrascale Cloud Services • Austin (TX)

On-site
USD 80,000 - 110,000
GPU HPC NOC Technician II: Data Center Ops
GPU HPC NOC Technician II: Data Center Ops

Cirrascale Corporation • Austin (TX)

On-site
USD 55,000 - 90,000
Data Center Technician II Cleveland
Data Center Technician II Cleveland

Cirrascale Corporation • Cleveland (OH)

On-site
USD 50,000 - 70,000
Network Operating Technician - Multiple Levels
Network Operating Technician - Multiple Levels

Cirrascale Cloud Services, LLC • Austin (TX)

On-site
USD 70,000 - 90,000
401(k)
Health insurance
Paid time off
+2
Data Center Technician II AUS (Afternoon Shift)
Data Center Technician II AUS (Afternoon Shift)

Cirrascale Corporation • Austin (TX)

On-site
USD 60,000 - 80,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Linux Admin
Linux Admin

TechDigital Group • United States

On-site
USD 100,000 - 130,000