Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops

Arcoro Holdings Corp

Phoenix, Northern (AZ, KY)

Hybrid

USD 200,000 - 220,000

Full time

2 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote Work
401(k) with Company match
Flexible PTO and Company-paid holidays

Job summary

Arcoro is seeking a Site Reliability Engineering Manager to lead the SRE team, ensuring availability, performance, and operational excellence for production systems. The role blends leadership with hands‑on coding, automation, on‑call duties, and collaboration with product teams.

You will drive SLOs/SLIs, mentor engineers, and champion reliable, AI‑driven practices while enabling agentic AI development across the platform. Remote work options available.

Qualifications

  • Proven experience leading SRE, operations, or reliability-focused engineering teams in a production software environment.
  • Willingness and ability to operate as a hands‑on individual contributor in addition to managing the team, including writing code, building automation, and participating in on‑call
  • Strong understanding of SRE principles, including SLOs/SLIs, error budgets, and blameless postmortems
  • Hands‑on background in incident response, on‑call management, and production troubleshooting
  • Experience with modern observability practices, including metrics, logging, tracing, and alerting
  • Demonstrated experience applying AI and automation to reliability work, including using AI‑assisted tooling, building automated remediation, and leading the adoption of AI‑driven practices on a team
  • Strong leadership, coaching, and team development skills
  • Excellent communication skills, including the ability to lead through high‑pressure incidents and communicate clearly with technical and non‑technical stakeholders

Responsibilities

  • Lead and manage a team of Site Reliability Engineers responsible for the reliability, performance, and operational health of production systems
  • Serve as a hands‑on technical contributor by writing code and automation, building reliability tooling, participating in on‑call, and working directly in production systems alongside the team
  • Support career growth and development of team members through coaching, mentoring, and performance management
  • Define, measure, and drive Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets in partnership with engineering and product teams
  • Own incident response, including on‑call rotations, escalation processes, severity management, and blameless postmortems
  • Drive continuous improvement in monitoring, observability, alerting, and on‑call practices to reduce toil and mean‑time‑to‑recovery
  • Lead the adoption of AI and automation across SRE practices, including AI‑assisted incident response, intelligent alerting, automated remediation, and the use of AI tooling to reduce toil and accelerate operational workflows
  • Partner with Engineering to refine our products to better support agentic AI development, including improving APIs, telemetry, environments, and platform capabilities that enable AI agents to safely build on and operate against our systems
  • Drive cloud cost optimization and FinOps practices in partnership with Engineering, including vendor management, cost allocation, rightsizing, and engineering best practices that reduce cloud spend
  • Partner with Engineering on operational readiness reviews, production change management, and release safety
  • Champion reliability best practices and ensure they are embedded across the engineering organization
  • Track and report on key reliability metrics, incident trends, and team health to leadership
  • Stay current with emerging SRE practices, tooling, and industry standards

Skills

SRE leadership
Hands-on coding
Incident management
Observability
AI automation
Cloud infrastructure
Leadership coaching
Communication

Education

Bachelor's degree in Computer Science

Tools

Kubernetes
Helm
Argo
Datadog
Grafana
OpenTelemetry
Azure Monitor
Terraform
CloudFormation
Azure DevOps
GitHub Actions
Bicep
PagerDuty

Job description

Arcoro is seeking a Site Reliability Engineering Manager to lead the SRE team, ensuring availability, performance, and operational excellence for production systems. The role blends leadership with hands‑on coding, automation, on‑call duties, and collaboration with product teams.

You will drive SLOs/SLIs, mentor engineers, and champion reliable, AI‑driven practices while enabling agentic AI development across the platform. Remote work options available.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE – AI-Driven Infra, Remote & Async
Senior SRE – AI-Driven Infra, Remote & Async

Remote • United States

Remote
USD 53,300 - 119,850
Budget for local social events
Flexible life-work balance
Senior SRE - AI-Driven Cloud Infra (Remote)
Senior SRE - AI-Driven Cloud Infra (Remote)

ioet • United States

Remote
USD 140,000 - 190,000
Remote work
USD compensation
Paid holidays
+1
Remote SRE Lead — AI Reliability & Observability
Remote SRE Lead — AI Reliability & Observability

United States Digital Space LLC • United States

Remote
USD 198,000 - 303,000
Remote SRE: AI Platform Reliability & Automation
Remote SRE: AI Platform Reliability & Automation

Runpod • United States

On-site
USD 150,000 - 200,000
Remote work first
Competitive base salary
Stock options equity
+2
Senior SRE: AI-Driven Reliability & Automation (Hybrid)
Senior SRE: AI-Driven Reliability & Automation (Hybrid)

Namely • United States

Hybrid
USD 120,000 - 150,000
Principal SRE: AI-Driven Reliability Leader (Remote)
Principal SRE: AI-Driven Reliability Leader (Remote)

Papa John'S International, Inc. • Kalispell (MT)

On-site
USD 150,000 - 190,000
Comprehensive benefits package
Equity stock purchase
401(k) contribution
+1
Remote SRE Lead - Incident, Reliability & Observability
Remote SRE Lead - Incident, Reliability & Observability

NightDragon Acquisition Corp. • United States

On-site
USD 260,000 - 280,000
Hybrid & Remote Work
Competitive Compensation
Equity package
Remote SRE Manager: Lead Reliability & Automation
Remote SRE Manager: Lead Reliability & Automation

NationsBenefits, LLC • Plantation (FL)

Remote
USD 140,000 - 180,000
Fully remote
Unlimited PTO
Competitive compensation
+1
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
SRE Manager: Reliability Leader for Scalable Cloud
SRE Manager: Reliability Leader for Scalable Cloud

Litera • Denver (CO)

Hybrid
USD 120,000 - 160,000