Cloud Engineer

Cirrascale Corporation

San Diego (CA)

On-site

USD 121,500 - 178,750

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive medical, dental, vision coverage
401(k) with company match
Generous paid time-off
Professional development support

Job summary

Cirrascale Corporation is seeking a Cloud Engineer to architect and manage scalable storage solutions tailored for AI workloads. The ideal candidate will have over 7 years of experience managing enterprise storage, specifically with Ceph and WEKA, and deep knowledge of various storage architectures.

This role involves collaborating with engineering teams to deliver high-performance infrastructure, optimizing system performance, and ensuring data integrity. Join Cirrascale to be part of a forward-thinking team driving AI innovation.

Qualifications

  • 7+ years of experience managing enterprise storage systems.
  • 3+ years of hands-on experience with Ceph and/or WEKA in production.
  • Deep knowledge of object, block, and parallel file systems.

Responsibilities

  • Architect and manage scalable storage solutions.
  • Optimize IOPS across distributed systems for AI workloads.
  • Monitor performance and ensure low latency and data integrity.

Skills

Distributed storage systems
Ceph
WEKA
Linux (Ubuntu/CentOS)
Bash scripting
Python scripting
AI/ML performance tuning
Docker
Kubernetes

Tools

Prometheus
Grafana
Ansible
Terraform

Job description

Cirrascale Cloud Services provides high-performance cloud infrastructure purpose-built for deep learning, generative AI, and large-scale AI inference workloads. We specialize in dedicated GPU cloud solutions tailored to the unique needs of startups, research labs, and enterprise AI teams. Our mission is to accelerate AI innovation by combining powerful hardware with white-glove service and flexible, custom-built environments.

We are seeking a Cloud Engineer with deep expertise in distributed storage systems, specifically Cephand WEKA, to architect, deploy, and maintain scalable storage infrastructures supporting AI and HPC workloads. This role is critical to ensure performance, resiliency, and data integrity across our customer environments.

You will play a key role in supporting large-scale GPU infrastructure deployments, collaborating with engineering and operations teams to deliver best-in-class storage solutions tailored to the demanding requirements of AI and ML workloads.

Key Responsibilities
  • Architect, deploy, and manage high-performance, scalable storage solutions based on Cephand WEKA for AI and deep learning workloads.
  • Optimize IOPS and throughput across distributed systems supporting hundreds of GPUs per cluster.
  • Develop and maintain infrastructure-as-code templates for automated storage deployments.
  • Monitor system performance and implement improvements to ensure low latency, high bandwidth, and data integrity.
  • Lead incident response and root cause analysis for storage-related issues across production environments.
  • Collaborate with system engineers, network teams, and customer success to tailor storage performance to specific workload needs.
  • Evaluate and integrate new storage technologies and NVMe architectures into the AI stack.
  • Write and maintain detailed technical documentation and runbooks.
  • Contribute to strategic infrastructure planning and scaling initiatives.
Required Qualifications
  • 7+ years of experience managing and scaling enterprise storage systems.
  • 3+ years of hands‑on experience with Ceph and/or WEKA in production environments.
  • Deep knowledge of storage architectures: object, block, and parallel file systems.
  • Strong understanding of RDMA, InfiniBand, NVMe‑oF, and distributed metadata systems.
  • Proficiency in Linux (preferably Ubuntu/CentOS) and scripting (Bash, Python).
  • Experience with performance tuning for AI/ML workloads using storage-intensive frameworks like TensorFlow, PyTorch, etc.
  • Familiarity with containerized and virtualized environments: Docker, Kubernetes, KVM, etc.
  • Strong troubleshooting and diagnostic skills in large-scale, multi‑tenant environments.
Preferred Qualifications
  • Experience with AI model training workflows, particularly storage IO patterns in multi-node GPU clusters.
  • Familiarity with storage solutions from NetApp, DDN, VAST Data, and cloud‑native offerings.
  • Experience with monitoring tools (Prometheus, Grafana), and configuration management (Ansible, Terraform).
  • Knowledge of data governance and compliance standards in AI environments.
Compensation and Benefits

We are committed to supporting the people who help us grow. We offer a competitive compensation package that includes base salary, performance bonuses and stock option opportunities.

The base salary range for the Cloud Engineer is $121,500 to $178,750. This pay range reflects the broad, minimum to maximum, pay range for this job for the location for which it has been posted. Compensation decisions are dependent on several factors including, but not limited to, an individual’s qualifications, location where the role is to be performed, internal equity, and alignment with market data.

Our benefits package includes comprehensive medical, dental, vision coverage, 401(k) with company match, generous paid time-off and professional development support.

Why Join Cirrascale?

Join a growing team that’s pushing the boundaries of AI infrastructure. At Cirrascale, you’ll contribute to projects powering next-generation AI applications while working with top‑tier hardware in a collaborative and innovative environment. From custom deployments to hands‑on customer support, every role here plays a part in enabling breakthroughs in AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Storage Engineer
Sr. Storage Engineer

Cirrascale Cloud Services, LLC • San Diego (CA)

On-site
USD 121,000 - 179,000
Health insurance
Dental insurance
Vision insurance
+3
Cloud Engineer
Cloud Engineer

Cirrascale Cloud Services, LLC • Austin (TX)

Hybrid
USD 101,000 - 194,000
Health insurance
Dental and vision insurance
Retirement plans
+2
Senior AI Storage Engineer: Ceph/WEKA for GPU Clusters
Senior AI Storage Engineer: Ceph/WEKA for GPU Clusters

Cirrascale Corporation • San Diego (CA)

Hybrid
USD 121,000 - 179,000
Comprehensive medical, dental, vision coverage
401(k) with company match
Generous paid time-off
+1
Senior Software Engineer, Backend Systems
Senior Software Engineer, Backend Systems

Cirrascale Corporation • San Diego (CA)

On-site
USD 175,000 - 200,000
Health insurance
Dental insurance
Vision insurance
+3
Deployment Technician
Deployment Technician

Cirrascale Corporation • Austin (TX)

Hybrid
Health insurance
Dental insurance
Vision insurance
+2
Software Engineer, Database Systems
Software Engineer, Database Systems

Cirrascale Corporation • Town of Texas (WI)

On-site
USD 145,000 - 175,000
Health insurance
Dental insurance
Vision insurance
+3
Infrastructure Engineer, Patching & Compliance
Infrastructure Engineer, Patching & Compliance

Cirrascale Corporation • Town of Texas (WI)

On-site
USD 100,000 - 140,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Engineer – AI & HPC Observability
Senior Engineer – AI & HPC Observability

Cirrascale Corporation • San Diego (CA)

On-site
USD 155,000 - 206,000
Health, dental, and vision insurance
Retirement plans
Paid time off
+1
Solutions Architect
Solutions Architect

Cirrascale Corporation • New York (NY)

On-site
USD 150,000 - 195,000
401(k) with company match
Health, dental, and vision insurance
Paid time off
Logistics Manager
Logistics Manager

Cirrascale • Austin (TX)

On-site
USD 110,000 - 120,000