Systems Reliability Engineer

CirrusLabs, LLC

Jacksonville (FL)

On-site

USD 120,000 - 170,000

Full time

18 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

CirrusLabs, LLC in Jacksonville, FL seeks a Systems Reliability Engineer to own the reliability, resiliency and security of a large-scale embedded finance platform. You will design monitoring, implement observability, and drive incident response end-to-end to improve MTTR and platform health.

You will work with engineering, product and risk teams to embed reliability best practices, define SLOs/SLIs, and automate toil-reducing tooling across the stack in a fast-paced, customer-focused

Qualifications

  • Hands-on experience with monitoring, observability and alerting tools.
  • Proven experience operating and supporting a large-scale enterprise platform.
  • Experience with incident response and leading RCA processes.
  • Strong understanding of reliability engineering principles and best practices.
  • Excellent communication skills to drive remediation across teams.

Responsibilities

  • Own the reliability, resiliency and availability of the Embedded Finance platform.
  • Design, implement and maintain monitoring and alerting frameworks for end-to-end observability.
  • Define and track SLOs/SLIs and error budgets to measure platform health.
  • Lead incident response and coordinate with engineering, product and risk teams.
  • Own RCA processes, document events and remediation actions to prevent recurrence.
  • Ensure incident tickets (e.g., ServiceNow) are created and tracked.
  • Build automation and tooling to reduce toil and improve MTTD/MTTR.
  • Collaborate with stakeholders to embed reliability best practices across the lifecycle.

Skills

Observability
Incident response
Root cause analysis
SRE principles
Collaboration

Tools

Splunk
Dynatrace
Grafana
Datadog
ServiceNow

Job description

Job Title: Systems Reliability Engineer,
Job Description:
What does a successful Systems Reliability Engineer (SRE) in Embedded Finance do on this project?

We are seeking a Systems Reliability Engineer to join the Technical Operations team supporting our Embedded Finance (EmFi) platform. In this role, you will own the reliability and resiliency of a large-scale enterprise platform, ensuring that our services remain highly available, performant and secure. You will design and implement monitoring and alerting frameworks, lead incident response and drive the root cause analysis (RCA) process to continuously improve platform stability. This is an opportunity to be a Partner in Possibility - helping our clients deliver financial services experiences that are essential to everyday life.

What you will do:
  • Own the reliability, resiliency and availability of the Embedded Finance platform, proactively identifying and mitigating risks to service continuity.
  • Design, implement and maintain comprehensive monitoring and alerting frameworks leveraging Splunk, Dynatrace, Grafana and Datadog to provide end-to-end observability across the platform.
  • Define and track service level objectives (SLOs), service level indicators (SLIs) and error budgets to measure and improve platform health.
  • Lead and participate in incident response, serving as a technical driver during remediation calls and coordinating with impacted and impacting technical and product teams.
  • Own and advance the root cause analysis (RCA) process — investigating incidents, documenting the sequence of events and remediating actions, and clearly identifying underlying root causes to prevent recurrence.
  • Ensure timely creation and management of incident tickets (e.g., ServiceNow) and accurate incident tracking, aging and reporting.
  • Build automation and tooling to reduce toil, improve mean time to detection (MTTD) and mean time to resolution (MTTR), and increase operational efficiency.
  • Collaborate with engineering, product and risk stakeholders to embed reliability best practices into the platform lifecycle.
What you will need to have:
  • Hands-on experience with monitoring, observability and alerting tools, specifically Splunk, Dynatrace, Grafana and Datadog.
  • Proven experience operating and supporting a large-scale enterprise platform environment.
  • Demonstrated experience with incident response and leading or contributing to root cause analysis (RCA) processes.
  • Strong understanding of reliability engineering principles, including availability, resiliency, monitoring and alerting best practices.
  • Experience with ticketing and incident management workflows (e.g., ServiceNow).
  • Excellent communication skills, with the ability to drive remediation efforts and collaborate across technical, product and risk teams.
What would be great to have:
  • Experience in financial services, payments or embedded finance environments.
  • Proficiency with scripting or programming languages for automation (e.g., Python, Go, Bash).
  • Familiarity with cloud platforms, containerization and CI/CD pipelines.
  • Experience defining and managing SLOs, SLIs and error budgets
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE - Embedded Finance Platform Reliability
SRE - Embedded Finance Platform Reliability

Ascendion • Jacksonville (FL), Northern (KY)

Hybrid
USD 125,000 - 136,000
Medical insurance
Dental insurance
Vision insurance
+7
Platform Reliability Engineer - Embedded Finance
Platform Reliability Engineer - Embedded Finance

Ascendion • Atlanta (GA), Northern (KY)

Hybrid
USD 76,000 - 90,000
Medical insurance
Dental insurance
Vision insurance
+3
Embedded Finance SRE: Reliability & Observability Lead
Embedded Finance SRE: Reliability & Observability Lead

CirrusLabs, LLC • Jacksonville (FL)

On-site
USD 120,000 - 170,000
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

Goldman Sachs • Dallas (TX)

On-site
USD 120,000 - 160,000
None
SRE
SRE

Ascendion • Jacksonville (FL), Northern (KY)

Hybrid
USD 125,000 - 136,000
Medical insurance
Dental insurance
Vision insurance
+7
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

On-site
USD 140,000 - 190,000
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

The Goldman Sachs Group • Dallas (TX)

On-site
USD 110,000 - 140,000
System Reliability Engineer
System Reliability Engineer

Ascendion • Atlanta (GA), Northern (KY)

Hybrid
USD 76,000 - 90,000
Medical insurance
Dental insurance
Vision insurance
+3
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

State of Wisconsin Investment Board • Madison (WI)

On-site
USD 140,000 - 180,000
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)

PulseRise Technologies • New York (NY)

Hybrid
USD 130,000 - 160,000