Senior SRE

Accelerant

United States

Remote

USD 140,000 - 210,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Accelerant is seeking an experienced reliability and observability leader to own end-to-end reliability initiatives across our financial data platform. You will define SLOs, harden deployments for enterprise-grade resiliency, and broaden observability across Velocity, MuleSoft, Snowflake, Fabric, and AWS.

You will prototype AI-powered operational tooling using Cursor and drive a scalable incident lifecycle, including blameless postmortems.

Qualifications

  • Proven experience designing, operating, and scaling reliable production systems.
  • Hands-on with modern observability tooling (Datadog, Prometheus/Grafana, OpenTelemetry).
  • Experience defining SLIs, SLOs, and error budgets and translating to business KPIs, not just infra metrics.
  • Experience operating data platforms (Snowflake, Fabric) and enterprise integration layers (MuleSoft) with enterprise SaaS such as D365.
  • Incident management experience with Incident.io and ServiceNow; on-call and postmortem practices.
  • Hands-on experience building with LLMs and AI coding assistants — Cursor in particular; building agents is a plus.
  • Ability to define reliability strategy and defend it to engineering leadership and the business.
  • Strong communication skills and autonomy in ambiguous, fast-changing environments.

Responsibilities

  • Own the reliability roadmap end to end and set SLOs and error budgets.
  • Harden the platform for availability, performance, and recoverability.
  • Extend instrumentation across Velocity, Red Panda, MuleSoft, Snowflake, Fabric, and AWS; cover service health and business KPIs.
  • Implement scalable incident and blameless postmortem processes with defined escalation paths.
  • Scale automation, data lineage, and auto-remediation; build actionable dashboards.
  • Build specialized SRE agents using Cursor AI for incident triage and root-cause analysis.
  • Host SRE agents on the AI fabric with governance and reuse across teams.

Skills

Datadog
OpenTelemetry
Prometheus/Grafana
SLIs and SLOs
Incident management
Cursor AI
AI agents
System architecture
Autonomy

Tools

Snowflake
Fabric
MuleSoft
D365
Cursor

Job description

The Role

We’re building the financial data platform at Accelerant — the premium, claims, and paid data products that underpin financial processing, reserving analysis, and the monthly close — and it needs to stay fast, resilient, and observable as we scale. You’ll drive the reliability and observability strategy across the platform and the enterprise systems it depends on: Velocity, MuleSoft, D365, Snowflake, Fabric, and the streaming and integration layers that move data through it. You are a key decider about what gets measured, how we define reliability, and where engineering needs to invest to keep production healthy.

We need someone who can prove a repeatable, define-to-alert observability pipeline, harden it, and scale it into systems that have never had real SLOs — and build modern, AI-assisted operational tooling that lets a small team punch far above its weight.

This Is a High-Autonomy, High-Impact Role for Someone Who:
  • Sees a recurring alert or a fragile deploy path and cannot leave it alone. Excels at shipping the right fix and the right automation, not the perfect one.
  • Has run real production systems at scale — not just written runbooks for them.
  • Has built with Datadog, OpenTelemetry, incident tooling, and AI coding assistants long enough to have strong opinions about what fits our needs.
  • Can prototype an operational agent in Cursor and iterate as they go.
  • Operates with autonomy, and can carry a technical discussion on system architecture, failure modes, and tradeoffs.
  • Is genuinely curious about applying emerging AI to reliability and operations.
What You’ll Do
Drive the reliability and observability initiative

Own the reliability roadmap end to end. Prove a repeatable define → emit → ingest → dashboard → alert metric pipeline, set SLOs and error budgets, prioritize the work, and drive execution. You’ll partner with engineering on what we monitor, how, and when — indexing on user impact over low-level infrastructure.

Harden the foundational platform

Take the financial data platform from functional to enterprise-grade, with a focus on availability, performance, and recoverability. Strengthen deployment paths, straight-through processing, and failover so the monthly close runs faster and cleaner as legacy hops are retired.

Expand observability breadth and depth

Extend instrumentation across the six target systems — Velocity, Red Panda, MuleSoft, Snowflake, Fabric, and AWS (with D365 ledger to follow) — proving both push (OpenTelemetry) and pull (agent) ingestion. Cover service health (latency, error rates, throughput) and business KPIs (match rate, reconciliation completeness, settlement correctness and latency).

Implement a scalable incident and review process

Build the on-call, alerting, and blameless postmortem process that keeps reliability high as systems and the team grow. Route alerts Datadog → Incident.io with ServiceNow as the system of record, and set severity standards, escalation norms, and follow-up tracking that actually closes the loop.

Scale automation, auditability, and reduce toil

Build the tooling that automates routine operations, self-heals common failures, and surfaces signal over noise. Establish data lineage and retention, and validate reliability at scale — 5,000+ transactions before go-live — through auto-remediation, capacity planning, and actionable dashboards.

Build specialized SRE agents using Cursor AI

Design and ship AI agents for incident triage, log analysis, and root-cause investigation (to name a few). Use Cursor as your build environment. Treat the agents as products solving specific problems.

Host SRE agents on the AI fabric

Partner with the AI platform team to deploy your agents on the org’s AI fabric. Make them discoverable, governed, and reusable across functions.

What You’ll Bring
Must-Haves
  • Proven experience designing, operating, and scaling reliable production systems.
  • Deep hands‑on expertise with modern observability tooling — Datadog, Prometheus/Grafana, and OpenTelemetry — including both push and pull ingestion patterns.
  • Strong background defining SLIs, SLOs, and error budgets — and translating them into business‑level KPIs, not just infrastructure metrics.
  • Experience operating data platforms (Snowflake, Fabric) and enterprise integration layers (MuleSoft) alongside enterprise SaaS such as D365 (F&O and/or Power Apps).
  • Hands‑on incident management experience with tools like Incident.io and ServiceNow, and a track record of running effective on‑call and postmortem practices.
  • Hands‑on experience building with LLMs and AI coding assistants — Cursor in particular. Bonus if you’ve built and deployed agents.
  • Ability to define reliability strategy, reliability targets, and operational metrics — and defend them to engineering leadership and the business.
  • Strong communication skills — you can explain a root cause to a junior engineer and a reliability risk to a product lead.
  • Demonstrated bias for action and ability to operate autonomously in ambiguous, fast‑changing environments.
Nice-to-Haves
  • Experience in insurance, fintech, or other regulated financial services industries.
  • Familiarity with insurance and finance concepts (premium, claims, settlement, reserving, monthly close) or willingness to learn them deeply.
  • Experience with streaming and event pipelines (Red Panda / Kafka) and data lineage, retention, and auditability requirements.
  • Strong working knowledge of chaos engineering, performance and load testing, and capacity planning.
  • Experience deploying AI agents on an internal AI platform or fabric (governance, eval harnesses, prompt/version management).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7
Technical Operations Lead
Technical Operations Lead

First Citizens • Raleigh (NC)

On-site
USD 140,000 - 190,000
Job Posting Title AI/ DevOps Engineer
Job Posting Title AI/ DevOps Engineer

Adobe • San Jose (CA)

On-site
USD 180,000 - 240,000
Data & AI Reliability Engineering Consultant/Architect
Data & AI Reliability Engineering Consultant/Architect

EPAM Systems • United States

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hard Rock Digital • United States

On-site
USD 150,000 - 210,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Senior Site Reliability Engineer, AI-DNA, $100k/year USD
Senior Site Reliability Engineer, AI-DNA, $100k/year USD

IgniteTech • United States

On-site
USD 120,000 - 160,000
Technical Operations Lead
Technical Operations Lead

First Citizens Bank • Phoenix (AZ)

On-site
USD 140,000 - 190,000
Benefits program
Technical Operations Lead
Technical Operations Lead

First Citizens Bank • Dallas (TX)

On-site
USD 140,000 - 180,000