IOC Systems Analyst

Optomi

Fort Worth (TX)

On-site

USD 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Opportunities to grow within cloud and AI technologies
Collaborate with engineering teams
Sustainability-focused organization

Job summary

Optomi, in partnership with a leading AI Cloud Service Provider, is seeking an IOC Systems Specialist to join a fast-paced operations team. The role focuses on providing Tier 2 operational support for high-performance computing cloud infrastructure while maintaining system stability and performance.

The ideal candidate will have 2–5 years of experience with HPC clusters, Kubernetes, and relevant cloud platforms. Opportunities for growth within AI and HPC technologies are available.

Qualifications

  • 2–5 years of experience supporting HPC clusters in a production IOC/NOC environment.
  • Hands-on experience with Kubernetes and Slurm workload manager.
  • Experience with storage technologies such as WEKA and VAST.

Responsibilities

  • Provide Tier 2 operational support for HPC cloud environments.
  • Monitor, troubleshoot, and resolve incidents related to Kubernetes and Slurm.
  • Act as escalation point for Tier 1 support teams.
  • Perform root cause analysis and contribute to improvements.

Skills

HPC cluster support
Kubernetes
Slurm workload manager
Incident response
Cloud platforms (AWS, Azure, GCP)
HPC networking

Education

Post-secondary education in Computer Science, Engineering, or related field

Tools

WEKA
VAST

Job description

Optomi, in partnership with a leading AI Cloud Service Provider, is seeking an IOC Systems Specialist to join a fast-paced operations team supporting large-scale HPC and GPU cloud environments.

Position Summary

The IOC Systems Specialist is responsible for providing Tier 2 operational support for high-performance computing (HPC) cloud infrastructure in a 24x7 IOC/NOC environment. This role focuses on monitoring, troubleshooting, and resolving complex incidents across Kubernetes clusters, Slurm-managed workloads, cloud services, and large-scale storage environments. The specialist will help ensure system stability, performance, and uptime while supporting mission-critical AI and GPU computing operations.

What the right candidate will enjoy
  • Working with cutting-edge AI and HPC infrastructure technologies
  • Supporting large-scale GPU cloud environments in a highly technical operations setting
  • Collaborating with engineering and infrastructure teams to solve complex production issues
  • Opportunities to grow within cloud, Kubernetes, HPC, and observability technologies
  • Being part of a sustainability-focused organization powered by renewable energy
What type of experience the right candidate has
  • 2–5 years of experience supporting or operating HPC clusters in a production IOC/NOC environment
  • Hands‑on experience with Kubernetes and Slurm workload manager
  • Experience supporting storage technologies such as WEKA and VAST
  • Background in incident response, troubleshooting, and root cause analysis within complex systems
  • Familiarity with cloud platforms such as AWS, Azure, or GCP
  • Understanding of HPC networking and storage infrastructure, including InfiniBand, Ethernet fabrics, and high‑throughput storage environments
  • Post‑secondary education in Computer Science, Engineering, or related technical discipline, or equivalent hands‑on experience
What the responsibilities are of the right candidate
  • Provide Tier 2 operational support for HPC cloud environments while maintaining system stability and SLA adherence
  • Monitor, troubleshoot, and resolve incidents related to Kubernetes, Slurm, storage systems, and associated cloud infrastructure
  • Act as an escalation point for Tier 1 support teams and coordinate with engineering teams for permanent resolution of issues
  • Perform root cause analysis and contribute to continuous operational improvements
  • Execute operational changes, maintenance activities, patching, and upgrades following change management procedures
  • Support and maintain monitoring, alerting, and observability tools for proactive issue detection
  • Maintain runbooks, operational documentation, incident reports, and knowledge base articles
  • Support operational readiness for new HPC technologies and infrastructure deployments
  • Provide guidance and mentorship to Tier 1 operations staff and data center technicians
  • Participate in a 24x7 rotating shift schedule and major incident response activities
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Cloud Operations Specialist
HPC Cloud Operations Specialist

Optomi • Fort Worth (TX)

On-site
USD 70,000 - 90,000
Opportunities to grow within cloud and AI technologies
Collaborate with engineering teams
Sustainability-focused organization
IOC Network Operations Specialist
IOC Network Operations Specialist

Optomi • Fort Worth (TX)

On-site
USD 80,000 - 100,000
Competitive total rewards package
100% company-paid medical, dental, and vision coverage
401(k) with company match
Staff Engineer, Senior Manager
Staff Engineer, Senior Manager

Jobtailor • Connecticut

On-site
USD 140,000 - 190,000
Virtualization & Orchestration Engineer
Virtualization & Orchestration Engineer

Jobtailor • Bellevue (WA)

On-site
USD 150,000 - 210,000
Network Operations Center Technician II
Network Operations Center Technician II

Cirrascale Corporation • Austin (TX)

On-site
USD 55,000 - 90,000
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Principal Software Developer, AI Infrastructure
Principal Software Developer, AI Infrastructure

Ll Oefentherapie • Austin (TX)

On-site
USD 120,000 - 160,000
Member of Technical Staff – AI Cloud Infrastructure
Member of Technical Staff – AI Cloud Infrastructure

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
Operations Specialist
Operations Specialist

Cirrascale Corporation • San Diego (CA)

On-site
USD 70,000 - 110,000
Software Developer 5
Software Developer 5

Ll Oefentherapie • Seattle (WA)

On-site
USD 130,000 - 160,000