HPC Cloud Operations Specialist

Optomi

Fort Worth (TX)

On-site

USD 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Opportunities to grow within cloud and AI technologies
Collaborate with engineering teams
Sustainability-focused organization

Job summary

Optomi, in partnership with a leading AI Cloud Service Provider, is seeking an IOC Systems Specialist to join a fast-paced operations team. The role focuses on providing Tier 2 operational support for high-performance computing cloud infrastructure while maintaining system stability and performance.

The ideal candidate will have 2–5 years of experience with HPC clusters, Kubernetes, and relevant cloud platforms. Opportunities for growth within AI and HPC technologies are available.

Qualifications

  • 2–5 years of experience supporting HPC clusters in a production IOC/NOC environment.
  • Hands-on experience with Kubernetes and Slurm workload manager.
  • Experience with storage technologies such as WEKA and VAST.

Responsibilities

  • Provide Tier 2 operational support for HPC cloud environments.
  • Monitor, troubleshoot, and resolve incidents related to Kubernetes and Slurm.
  • Act as escalation point for Tier 1 support teams.
  • Perform root cause analysis and contribute to improvements.

Skills

HPC cluster support
Kubernetes
Slurm workload manager
Incident response
Cloud platforms (AWS, Azure, GCP)
HPC networking

Education

Post-secondary education in Computer Science, Engineering, or related field

Tools

WEKA
VAST

Job description

Optomi, in partnership with a leading AI Cloud Service Provider, is seeking an IOC Systems Specialist to join a fast-paced operations team. The role focuses on providing Tier 2 operational support for high-performance computing cloud infrastructure while maintaining system stability and performance.

The ideal candidate will have 2–5 years of experience with HPC clusters, Kubernetes, and relevant cloud platforms. Opportunities for growth within AI and HPC technologies are available.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

IOC Systems Analyst
IOC Systems Analyst

Optomi • Fort Worth (TX)

On-site
USD 70,000 - 90,000
Opportunities to grow within cloud and AI technologies
Collaborate with engineering teams
Sustainability-focused organization
Hybrid Cloud Network Operations Tier-2 Specialist
Hybrid Cloud Network Operations Tier-2 Specialist

Optomi • Fort Worth (TX)

On-site
USD 80,000 - 100,000
Competitive total rewards package
100% company-paid medical, dental, and vision coverage
401(k) with company match
IOC Network Operations Specialist
IOC Network Operations Specialist

Optomi • Fort Worth (TX)

On-site
USD 80,000 - 100,000
Competitive total rewards package
100% company-paid medical, dental, and vision coverage
401(k) with company match
OCI AI & GPU HPC Infrastructure Architect
OCI AI & GPU HPC Infrastructure Architect

Ll Oefentherapie • United States

On-site
USD 180,000 - 240,000
AI/HPC Systems Engineer — Hybrid Cloud & GPUs
AI/HPC Systems Engineer — Hybrid Cloud & GPUs

Protingent • San Jose (CA)

On-site
USD 110,000 - 124,000
Insurance plan options (HDHP/POS)
Pre-tax commuter benefits
401k plan
+1
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
HPC & AI Cloud Solutions Engineer
HPC & AI Cloud Solutions Engineer

VC5 Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
AI/ML Infrastructure Lead — GPU Cluster & Cloud Ops
AI/ML Infrastructure Lead — GPU Cluster & Cloud Ops

Oracle • United States

On-site
USD 121,000 - 307,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off
Senior Platform Reliability Engineer (Kubernetes & CI/CD)
Senior Platform Reliability Engineer (Kubernetes & CI/CD)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
GPU / HPC Consultant
GPU / HPC Consultant

Arke • United States

On-site
USD 120,000 - 180,000