Senior AI Infra SRE — GPU Cloud Reliability Leader
deCircle
San Francisco (CA)
On-site
USD 120,000 - 150,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
deCircle is seeking a Site Reliability Engineer based in San Francisco to ensure operational excellence for our GPU marketplace and AI infrastructure. The role involves defining service level objectives, managing capacity for a distributed system, and ensuring security protocols are adhered to. Candidates should have a strong background in reliability engineering, capacity planning, and incident response, with an emphasis on developing resilient infrastructures. Join us to contribute to our mission of making AI accessible and affordable globally.
Qualifications
Expert in site reliability engineering with proven experience defining, monitoring, and maintaining SLOs.
Strong background in capacity planning and management for distributed systems.
Experienced in incident response and post‑mortem processes.
Knowledge of deployment systems including progressive rollouts and automated rollback.
Proficient in observability tools and practices such as metrics, logging, and tracing.
Responsibilities
Ensure reliability, performance, and security of GPU marketplace and AI infrastructure.
Define and maintain service level objectives for job success rates.
Build robust incident response systems and manage capacity across distributed GPU network.
Implement security and compliance frameworks to protect the infrastructure.
Skills
Site Reliability Engineering
Capacity Planning
Incident Response
Deployment Systems Knowledge
Observability Tools
Infrastructure Security
Secrets Management
Problem-Solving
Automation Mindset
Tools
Prometheus
Grafana
ELK Stack
Job description
deCircle is seeking a Site Reliability Engineer based in San Francisco to ensure operational excellence for our GPU marketplace and AI infrastructure. The role involves defining service level objectives, managing capacity for a distributed system, and ensuring security protocols are adhered to. Candidates should have a strong background in reliability engineering, capacity planning, and incident response, with an emphasis on developing resilient infrastructures. Join us to contribute to our mission of making AI accessible and affordable globally.