Site Reliability Engineer

asobbi

California (MO)

On-site

USD 170,000 - 220,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fully remote (US timezone)

Job summary

asobbi is building a UK sovereign AI cloud powered by renewable energy and expanding a US-based operations team. We are seeking Senior Site Reliability Engineers to build that capability from the ground up, with a strong emphasis on automation and software-engineering rather than AI model work.

The role is fully remote in the US timezone, offering high-visibility, groundbreaking infrastructure for next-generation AI and an automation-first culture.

Qualifications

  • Experience in SRE or platform engineering roles.
  • Hands-on with observability and monitoring tools.
  • Exposure to incident management/on-call and automation of runbooks.
  • Strong Python for automation and integrations.

Responsibilities

  • Build Python-based automation for incident triage and runbook execution.
  • Integrate observability, ITSM, and infrastructure APIs to enrich alerts.
  • Improve monitoring signal quality through correlation and deduplication.
  • Develop internal tools, dashboards, and CLI utilities for self-service.
  • Maintain runbook-as-code and automation libraries.
  • Turn post-incident learnings into better tooling and standards.

Skills

SRE / Platform Engineering
Python automation
Incident management
On-call experience

Tools

Prometheus
Grafana
ServiceNow
Halo
Jira Service Management
OpenTelemetry
ChatOps bots
DCIM
IPAM
Hypervisor integration

Job description

SENIOR SITE RELIABILITY ENGINEER | FULLY REMOTE (US)

Location: Fully remote - US-based, working a US timezone

Package: $170,000 - $220,000

Overvie w

We're supporting a specialist AI infrastructure company - a UK sovereign AI cloud powered by renewable energy - that builds and operates large-scale compute on regenerated industrial and energy sites.

Having recently secured a major US customer and taken on their entire cluster, the business is standing up a US-based operations team and is looking for Senior Site Reliability Engineers to build that capability from the groundup.

This is a Platform/SRE role with a strong automation and software-engineering bias — not an AI model-building role. You'll turn runbooks, alerts and operational workflows into safe, auditable automation that improves reliability across the platform.

  • Get in at the ground floor of a brand-new US operations function and shape how it runs at scale
  • Work on critical infrastructure powering the next generation of AI
  • Automation-first culture - reduce toil and build tooling, rather than fire fight
  • Real autonomy and high visibility with leadership
  • Fully remote, on a US timezone (West-coast preferred)
What you'll be doing
  • Building Python-based automation for incident triage, runbook execution and routine operational tasks
  • Integrating observability, ITSM and infrastructure APIs to enrich alerts and automate workflows
  • Improving monitoring signal quality through correlation, enrichment, suppression and deduplication
  • Building internal tools and self-service capabilities - CLI utilities, ChatOps integrations and dashboards
  • Maintaining version-controlled runbook-as-code and automation libraries
  • Turning post-incident learnings into better tooling, automation and operational standards
We're keen to speak with candidates who have
  • Experience in SRE, Platform Engineering or production infrastructure operations
  • Hands-on experience with observability/monitoring tooling (Prometheus, Grafana or similar)
  • Exposure to incident management / on-call, and converting manual runbooks into automation
  • Strong Python for automation, APIs and integrations
Nice to have:
  • GPU, datacentre or colocation infrastructure experience
  • ITSM integrations (ServiceNow, Halo, Jira Service Management or similar)
  • ChatOps tooling (Slack or Microsoft Teams bots)
  • OpenTelemetry, logging or distributed tracing experience
  • DCIM, IPAM or hypervisor-control-plane integrations
  • Experience with LLM-assisted or agent-based operational automation
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Hampshire

On-site
USD 110,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

BridgeSource Utilities Solutions • United States

Hybrid
USD 140,000 - 190,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Oaks (PA)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare coverage
401(k) matching
Tuition reimbursement
+1