Senior AI Infra SRE — GPU Cloud Reliability Leader

deCircle

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

deCircle is seeking a Site Reliability Engineer based in San Francisco to ensure operational excellence for our GPU marketplace and AI infrastructure. The role involves defining service level objectives, managing capacity for a distributed system, and ensuring security protocols are adhered to. Candidates should have a strong background in reliability engineering, capacity planning, and incident response, with an emphasis on developing resilient infrastructures. Join us to contribute to our mission of making AI accessible and affordable globally.

Qualifications

  • Expert in site reliability engineering with proven experience defining, monitoring, and maintaining SLOs.
  • Strong background in capacity planning and management for distributed systems.
  • Experienced in incident response and post‑mortem processes.
  • Knowledge of deployment systems including progressive rollouts and automated rollback.
  • Proficient in observability tools and practices such as metrics, logging, and tracing.

Responsibilities

  • Ensure reliability, performance, and security of GPU marketplace and AI infrastructure.
  • Define and maintain service level objectives for job success rates.
  • Build robust incident response systems and manage capacity across distributed GPU network.
  • Implement security and compliance frameworks to protect the infrastructure.

Skills

Site Reliability Engineering
Capacity Planning
Incident Response
Deployment Systems Knowledge
Observability Tools
Infrastructure Security
Secrets Management
Problem-Solving
Automation Mindset

Tools

Prometheus
Grafana
ELK Stack

Job description

deCircle is seeking a Site Reliability Engineer based in San Francisco to ensure operational excellence for our GPU marketplace and AI infrastructure. The role involves defining service level objectives, managing capacity for a distributed system, and ensuring security protocols are adhered to. Candidates should have a strong background in reliability engineering, capacity planning, and incident response, with an emphasis on developing resilient infrastructures. Join us to contribute to our mission of making AI accessible and affordable globally.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI GPU Infra SRE - Scale, Automation & Equity
Senior AI GPU Infra SRE - Scale, Automation & Equity

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior SRE: AI Infra on-Site in SF, GPU & Cloud
Senior SRE: AI Infra on-Site in SF, GPU & Cloud

The Recruiting Guy • Arlington (VA)

On-site
USD 175,000 - 250,000
Senior AI Infra SRE: GPU Clusters & High-Perf Networking
Senior AI Infra SRE: GPU Clusters & High-Perf Networking

Andromeda • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Senior AI-Driven SRE for Cloud Reliability
Senior AI-Driven SRE for Cloud Reliability

Cerebras • Mountain View (CA)

Hybrid
USD 100,000 - 150,000
Competitive salary and benefits package
Opportunities for professional growth
Collaborative work environment
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity

FLUIX • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Attractive compensation package including equity options
Comprehensive health, dental, and vision insurance
Opportunities for professional growth
Remote Senior SRE - AI-Driven Platform & Cloud
Remote Senior SRE - AI-Driven Platform & Cloud

Circle Internet Management Services LLC • California (MO)

On-site
USD 153,000 - 205,000
Senior SRE: GPU-Driven, Global Scale & Causal AI
Senior SRE: GPU-Driven, Global Scale & Causal AI

Crossing Hurdles • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior SRE — AI-Driven Cloud Reliability
Senior SRE — AI-Driven Cloud Reliability

BetterUp • New York (NY)

Hybrid
USD 164,000 - 205,000
Senior SRE: Blockchain & AI Infra Lead
Senior SRE: Blockchain & AI Infra Lead

ZetaChain • San Francisco (CA)

On-site
USD 140,000 - 190,000
Competitive compensation
Remote work with quarterly meetups
Strong open-source culture
Cloud SRE Architect — AI-Driven CI/CD & Scale
Cloud SRE Architect — AI-Driven CI/CD & Scale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits