Senior SRE: AI-Driven Reliability & Observability

Accelerant

United States

Remote

USD 140,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Accelerant is seeking an experienced reliability and observability leader to own end-to-end reliability initiatives across our financial data platform. You will define SLOs, harden deployments for enterprise-grade resiliency, and broaden observability across Velocity, MuleSoft, Snowflake, Fabric, and AWS.

You will prototype AI-powered operational tooling using Cursor and drive a scalable incident lifecycle, including blameless postmortems.

Qualifications

  • Proven experience designing, operating, and scaling reliable production systems.
  • Hands-on with modern observability tooling (Datadog, Prometheus/Grafana, OpenTelemetry).
  • Experience defining SLIs, SLOs, and error budgets and translating to business KPIs, not just infra metrics.
  • Experience operating data platforms (Snowflake, Fabric) and enterprise integration layers (MuleSoft) with enterprise SaaS such as D365.
  • Incident management experience with Incident.io and ServiceNow; on-call and postmortem practices.
  • Hands-on experience building with LLMs and AI coding assistants — Cursor in particular; building agents is a plus.
  • Ability to define reliability strategy and defend it to engineering leadership and the business.
  • Strong communication skills and autonomy in ambiguous, fast-changing environments.

Responsibilities

  • Own the reliability roadmap end to end and set SLOs and error budgets.
  • Harden the platform for availability, performance, and recoverability.
  • Extend instrumentation across Velocity, Red Panda, MuleSoft, Snowflake, Fabric, and AWS; cover service health and business KPIs.
  • Implement scalable incident and blameless postmortem processes with defined escalation paths.
  • Scale automation, data lineage, and auto-remediation; build actionable dashboards.
  • Build specialized SRE agents using Cursor AI for incident triage and root-cause analysis.
  • Host SRE agents on the AI fabric with governance and reuse across teams.

Skills

Datadog
OpenTelemetry
Prometheus/Grafana
SLIs and SLOs
Incident management
Cursor AI
AI agents
System architecture
Autonomy

Tools

Snowflake
Fabric
MuleSoft
D365
Cursor

Job description

Accelerant is seeking an experienced reliability and observability leader to own end-to-end reliability initiatives across our financial data platform. You will define SLOs, harden deployments for enterprise-grade resiliency, and broaden observability across Velocity, MuleSoft, Snowflake, Fabric, and AWS.

You will prototype AI-powered operational tooling using Cursor and drive a scalable incident lifecycle, including blameless postmortems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI-Driven Reliability & Observability
Senior SRE: AI-Driven Reliability & Observability

Accelerant • United States

Remote
USD 120,000 - 160,000
Senior AI Reliability Engineer - Observability & Incidents
Senior AI Reliability Engineer - Observability & Incidents

Cisco • Milwaukee (WI)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+16
Remote Senior Reliability Engineer: AI-Driven Observability
Remote Senior Reliability Engineer: AI-Driven Observability

Hims, Inc. • Northern (KY)

Remote
USD 140,000 - 190,000
Equity compensation
Unlimited PTO
Health benefits including medical, etc
+3
Senior SRE: AI-Driven Reliability & Incident Leadership
Senior SRE: AI-Driven Reliability & Incident Leadership

Salesforce • San Francisco (CA)

On-site
USD 149,000 - 246,000
Senior SRE
Senior SRE

Accelerant • United States

Remote
USD 140,000 - 210,000
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

1611 Ally Bank • Charlotte (NC)

On-site
USD 110,000 - 180,000
Annual incentive plan
Relocation assistance
Senior Reliability Engineer — AI-Driven Observability
Senior Reliability Engineer — AI-Driven Observability

hims & hers • United States

Remote
USD 140,000 - 210,000
Competitive salary & equity
Unlimited PTO
Comprehensive health benefits
+3
Staff SRE: AI-Driven Reliability & Platform Architect
Staff SRE: AI-Driven Reliability & Platform Architect

Devopsroles • Northern (KY)

Remote
USD 150,000 - 225,000
Equity
Benefits program
Staff SRE: AI-Driven Reliability Architect (Hybrid)
Staff SRE: AI-Driven Reliability Architect (Hybrid)

EarnIn Bfwf • Mountain View (CA)

Hybrid
USD 252,000 - 308,000
Senior SRE & Observability Consultant — Client-Facing
Senior SRE & Observability Consultant — Client-Facing

Quality Ai • Northern (KY)

Hybrid
USD 150,000 - 170,000
Global mobility opportunities
Technical training academy