Network Operations Center Technician II

Cirrascale Corporation

Austin (TX)

On-site

USD 55,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cirrascale, a provider of high-performance GPU cloud infrastructure, seeks a Network Operations Technician II to join our 24/7 NOC. You will monitor alerts, triage incidents, and coordinate resolutions across global data centers, using Jira for ticketing and collaborating with customers.

The ideal candidate has Linux expertise, networking knowledge, and experience with GPU nodes. You will mentor teammates, help optimize procedures, and ensure SOC 2 compliant operations while delivering strong

Qualifications

  • 2-4 years of experience in HPC, AI infrastructure, cloud systems, or related fields.
  • Understanding of scripting (Python, Bash) and GPU resource monitoring preferred.
  • Solid understanding of HPC datacenter networking principles and troubleshooting.
  • Strong analytical and problem-solving skills; able to work independently and manage multiple tasks.
  • Excellent communication skills and customer service is a must.
  • Certifications: Advanced Linux or AI/ML certifications are a plus.
  • Experience in Datacenter Network Operations and with RMAs/logistics a plus.
  • Experience with Jira (Atlassian) ticketing system is a plus.

Responsibilities

  • First-line response to alerts and incidents affecting systems and jobs.
  • Demonstrate understanding of GPU nodes, deployment, networking, and clustering in a datacenter.
  • Assist customers with ticket triage and basic troubleshooting using Jira.
  • Perform high-level monitoring and troubleshooting on all nodes and network equipment.
  • Remotely troubleshoot servers and GPUs at global datacenters.
  • Lead major incident response and coordinate resolutions.
  • Review and optimize NOC procedures and workflows.
  • Support capacity planning and performance monitoring.
  • Document system configurations, updates, and inventory; maintain asset records.

Skills

HPC infra
Scripting: Python
Linux admin
Network troubleshooting
Customer service
Communication
Datacenter networking
GPU orchestration

Tools

Jira (Atlassian)
Microsoft 365

Job description

Cirrascale Cloud Services provides high-performance cloud infrastructure purpose-built for deep learning, generative AI, and large-scale AI inference workloads. We specialize in dedicated GPU cloud solutions tailored to the unique needs of startups, research labs, and enterprise AI teams. Our mission is to accelerate AI innovation by combining powerful hardware with white-glove service and flexible, custom-built environments.

Position Overview

As a Network Operations Technician II at Cirrascale, you will be an integral part of our Operations team, responsible for maintaining the integrity and functionality of our data centers. The ideal candidate will have a strong background working in a 365-days 24/7 Network Operations Center environment. In this role, the NOC Technician will be responsible for monitoring the network via an alarm reporting tool. NOC Technician will create an outage ticket in Jira (Atlassian). Customers will be notified of the outage and that we are working on a trouble ticket to resolve it. You are the face of our company to our customers, ensuring that every interaction is a best-in-class customer service experience, so being able to demonstrate quality written and oral communication with customers and other stakeholders is of paramount importance. The primary responsibilities associated with the position include providing technical support to customers who are experiencing problems with their network and mentoring less experienced colleagues. The ideal candidate will have a strong background in SOC 2 compliance, Linux operating systems, a thorough understanding of networking principles, and experience interfacing GPUs remotely and Internet outages.

Key Responsibilities

  • First-line response to alerts and incidents to systems and job failures
  • Demonstrate understanding of GPU nodes and how they are deployed, networked, and clustered within a datacenter
  • Assist customers with ticket triage and basic troubleshooting using the Jira (Atlassian) ticketing system
  • Perform high-level monitoring and troubleshooting on all nodes and network equipmen t
  • Remotely Troubleshoot Installed Servers & GPUs at various global datacenter locations
  • Resolve complex and critical incidents within our datacenter
  • Lead major incident response and coordinatio n
  • Review and optimize existing NOC procedures
  • Work on capacity planning and performance monitorin g
  • Knowledge and experience working with Dell, SuperMicro & Lenovo type Servers is highly recommended
  • Perform deep troubleshooting of GPU node failures, job preemption conflicts, and cluster imbalance
  • Analyze alerts for GPU utilization inefficiencies, failed Machine Learning pipelines, or I/O bottlenecks
  • Provide on-site and remote support to resolve urgent technical issues
  • Document system configurations, updates, and inventory, maintaining accurate records of data center assets.
  • Stay current with industry trends, emerging technologies, and best practices in HPC network operations center trends
  • Analyze alerts for GPU utilization inefficiencies, failed Machine Learning pipelines, or I/O bottlenecks
  • Provide on-site and remote support to resolve urgent technical issues
  • Document system configurations, updates, and inventory, maintaining accurate records of data center assets.
  • Stay current with industry trends, emerging technologies, and best practices in HPC network operations center trends

Qualifications

  • 2-4 years of experience in HPC, AI infrastructure, cloud systems, or related
  • Understanding of scripting (Python, Bash, etc.), GPU resource monitoring preferred
  • Solid understanding of HPC datacenter networking principles and experience with network troubleshooting
  • Strong analytical and problem-solving skills, with the ability to work independently and manage multiple tasks
  • Excellent communication skills and the ability to collaborate effectively with the customer and the team. Customer Service is a must
  • Certifications: Advanced Linux or any other AI/ML certifications are a huge plus
  • Experience in Datacenter Network Operations
  • Experience with RMAs, logistics, shipping, and receiving a plus
  • It is a plus with experience working in Jira (Atlassian) ticketing system.
  • Proficient in Microsoft 365(Outlook, Word, Excel)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Operations Center Technician II
Network Operations Center Technician II

Cirrascale Cloud Services • Austin (TX)

On-site
USD 80,000 - 110,000
Network Operating Technician - Multiple Levels
Network Operating Technician - Multiple Levels

Cirrascale Cloud Services, LLC • Austin (TX)

On-site
USD 70,000 - 90,000
401(k)
Health insurance
Paid time off
+2
GPU HPC NOC Technician II: Data Center Ops
GPU HPC NOC Technician II: Data Center Ops

Cirrascale Corporation • Austin (TX)

On-site
USD 55,000 - 90,000
Data Center Technician II
Data Center Technician II

Cirrascale Cloud Services • Austin (TX)

On-site
USD 55,000 - 85,000
Data Center Technician II AUS (Afternoon Shift)
Data Center Technician II AUS (Afternoon Shift)

Cirrascale Corporation • Austin (TX)

On-site
USD 60,000 - 80,000
Data Center Technician II Cleveland
Data Center Technician II Cleveland

Cirrascale Corporation • Cleveland (OH)

On-site
USD 50,000 - 70,000
NOC Engineer: AI GPU Cloud & HPC Operations
NOC Engineer: AI GPU Cloud & HPC Operations

Cirrascale Cloud Services • Austin (TX)

On-site
USD 80,000 - 110,000
Deployment Technician
Deployment Technician

Cirrascale Corporation • Austin (TX)

Hybrid
Health insurance
Dental insurance
Vision insurance
+2
NOC Engineer
NOC Engineer

Axe Compute • Miami (FL)

On-site
USD 65,000 - 90,000
Network Operations Center (NOC) Technician II - Morning Shift
Network Operations Center (NOC) Technician II - Morning Shift

Dynascale Technologies • Los Angeles (CA)

On-site
USD 65,000 - 90,000