Hyperbolic Labs - Senior Site Reliability Engineer

deCircle

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

deCircle is seeking a Site Reliability Engineer based in San Francisco to ensure operational excellence for our GPU marketplace and AI infrastructure. The role involves defining service level objectives, managing capacity for a distributed system, and ensuring security protocols are adhered to. Candidates should have a strong background in reliability engineering, capacity planning, and incident response, with an emphasis on developing resilient infrastructures. Join us to contribute to our mission of making AI accessible and affordable globally.

Qualifications

  • Expert in site reliability engineering with proven experience defining, monitoring, and maintaining SLOs.
  • Strong background in capacity planning and management for distributed systems.
  • Experienced in incident response and post‑mortem processes.
  • Knowledge of deployment systems including progressive rollouts and automated rollback.
  • Proficient in observability tools and practices such as metrics, logging, and tracing.

Responsibilities

  • Ensure reliability, performance, and security of GPU marketplace and AI infrastructure.
  • Define and maintain service level objectives for job success rates.
  • Build robust incident response systems and manage capacity across distributed GPU network.
  • Implement security and compliance frameworks to protect the infrastructure.

Skills

Site Reliability Engineering
Capacity Planning
Incident Response
Deployment Systems Knowledge
Observability Tools
Infrastructure Security
Secrets Management
Problem-Solving
Automation Mindset

Tools

Prometheus
Grafana
ELK Stack

Job description

Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By aggregating computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open‑source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We're looking for forward‑thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality.

As we prepare for growth after our Series A, our team — led by co‑founders with PhDs in AI, Math, and Computer Science — is poised to redefine computing.

About the Role

We're seeking a Site Reliability Engineer to ensure Hyperbolic's GPU marketplace and AI infrastructure operate with exceptional reliability, performance, and security. As an aggregator of compute resources from hundreds of global suppliers, our SLOs, trust, and economic efficiency are product‑critical. You'll be responsible for defining and maintaining service level objectives for job success rates, building robust incident response systems, managing capacity across our distributed GPU network, and implementing secure rollout and rollback mechanisms that keep our platform running smoothly 24/7.

In this role, you'll establish the reliability standards that define customer trust in our platform, design monitoring and alerting systems that provide deep visibility into our infrastructure, build automation for capacity management and resource allocation, lead incident response and post‑mortem processes, and work closely with engineering teams to improve system resilience. You'll also focus on security and infrastructure hardening, ensuring strong isolation between tenants and suppliers, implementing key management systems, and building compliance frameworks. This is a high‑impact position where your work directly influences our ability to deliver on our promise of affordable, accessible AI compute at scale.


  • Expert in site reliability engineering with proven experience defining, monitoring, and maintaining SLOs and SLAs for production systems

  • Strong background in capacity planning and management, including forecasting, resource allocation, and cost optimization for distributed systems

  • Experienced in incident response, on‑call rotations, and post‑mortem processes with a track record of reducing MTTR and improving system resilience

  • Deep knowledge of deployment systems including progressive rollouts, canary deployments, feature flags, and automated rollback mechanisms

  • Proficient in observability tools and practices including metrics, logging, tracing, and alerting systems (Prometheus, Grafana, ELK stack, or similar)

  • Strong understanding of infrastructure security including tenant isolation, workload isolation, network segmentation, and security hardening

  • Experience with secrets management, key management systems (KMS), certificate management, and secure credential rotation

  • Knowledge of compliance frameworks and security best practices for cloud platforms (SOC 2, ISO 27001, or similar)

  • Excellent problem‑solving skills with ability to debug complex distributed systems issues under pressure

  • Strong automation mindset with experience using infrastructure‑as‑code, configuration management, and CI/CD pipelines

Preferred Qualifications
  • Experience operating GPU infrastructure, AI/ML platforms, or compute marketplaces at scale

  • Background in distributed systems, peer‑to‑peer networks, or decentralized infrastructure

  • Knowledge of multi‑tenancy security patterns, container security, and runtime security tools

  • Experience with chaos engineering, fault injection, and resilience testing

  • Familiarity with cost optimization strategies for cloud infrastructure and GPU resources

  • Experience building and operating systems with demanding uptime requirements (99.9%+ SLAs)

  • Background at companies like AWS, Google Cloud, Azure, or fast‑growing infrastructure startups

  • Contributions to open‑source reliability, observability, or security tools

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior GPU Infrastructure Engineer
Senior GPU Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Staff Software Engineer (Platform/Infrastructure)
Staff Software Engineer (Platform/Infrastructure)

Hyperbolic • San Francisco (CA)

On-site
USD 130,000 - 170,000
Hyperbolic Labs -Member of Technical Staff - Full Stack
Hyperbolic Labs -Member of Technical Staff - Full Stack

deCircle • San Francisco (CA)

On-site
USD 130,000 - 170,000
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior Data Analyst
Senior Data Analyst

Unchain Data • San Francisco (CA)

On-site
USD 80,000 - 120,000
Senior Data Analytics Engineer
Senior Data Analytics Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 90,000 - 120,000
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000