Site Reliability Engineer (Hybrid)

Unlimited Systems

Cincinnati (OH)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Unlimited Systems is seeking a Site Reliability Engineer to own the observability strategy for our AKS-based production systems in Cincinnati, Ohio. You will define SLOs/SLIs, partner with DevOps and software engineering, and drive reliability as a first‑class concern to keep our services healthy and scalable.

You will build tooling and dashboards, refine incident response processes, and contribute to security reviews as we grow our cloud-native platform.

Qualifications

  • 5+ years in SRE/DevOps/Platform Engineering in production cloud.
  • Hands-on AKS: provisioning, upgrades, CNI, lifecycle management.
  • Splunk SPL authoring, dashboards, alerting, data onboarding.
  • Azure Monitor, Log Analytics, Application Insights proficiency.
  • IaC with Terraform; Helm and Kustomize familiarity.
  • Scripting in Python/Go/Bash.
  • Drive incident resolution and post-mortems.
  • Translate complex topics for non-technical stakeholders.

Responsibilities

  • Own the observability strategy for AKS-based production systems.
  • Define and drive meaningful SLOs and SLIs.
  • Collaborate with DevOps and software engineering to improve reliability.
  • Build internal tooling to reduce toil and accelerate incident response.
  • Participate in security reviews and threat modeling for platform capabilities.

Skills

AKS experience
Splunk SPL
Instana APM
Azure Monitor & Log Analytics
Terraform
Helm & Kustomize
Python/Go/Bash
Incident response
Clear communication

Tools

Terraform
Helm
Kustomize

Job description

The Company

Unlimited Systems authors the category‑leading Unlimited Financials practice management system focused on the unique revenue cycle requirements of specialty healthcare providers. Unlimited Systems customers enjoy streamlined business office workflows, reduced claim denial rates, and accelerated and amplified revenue streams. Unlimited Systems is committed to ensuring that specialty healthcare providers thrive in a dynamic reimbursement environment.

Unlimited Systems is a portfolio company of Francisco Partners, a leading technology investment firm with deep sector focus and a track record of delivering outstanding returns. Through private equity and credit funds, they provide flexible capital and partnership to growth‑aspiring technology companies.

The Challenge

Unlimited Financials operates across a sophisticated cloud‑native environment — Azure Kubernetes Service, event‑sourced microservices, healthcare integrations, and a growing network of downstream data pipelines — all of which requires dedicated technical support to remain stable, observable, and performant for our customers. We are building the reliability practice this platform deserves. As we scale into new specialty healthcare markets and take on greater complexity, we need someone who thinks proactively — not just responding to incidents but engineering the systems and signals that prevent them. This is not a run‑and‑maintain role. It is a “build‑and‑shape‑the‑future” role.

Our Site Reliability Engineer will own the observability strategy across our observability applications, define and drive meaningful SLOs and SLIs for our critical services, and work shoulder‑to‑shoulder with DevOps and Software Engineering to make reliability a first‑class concern — not an afterthought. You will bring structure to chaos, clarity to ambiguity, and measurable improvement to customer experience.

The Position

A full‑time, hybrid role reporting to our Technical Support Manager, our Site Reliability Engineer will work frequently from our headquarters in Cincinnati, Ohio. Residency in the state of Ohio is preferred.

As Site Reliability Engineer, you will own the reliability, scalability, and performance of our production systems hosted on Azure Kubernetes Service (AKS), working at the intersection of software engineering and operations to build the tooling, processes, and culture that keep our services running at scale. You will be a key contributor to our observability practice — using Splunk for log analytics and alerting, Instana for APM and distributed tracing, and native Azure tools including Azure Monitor, Log Analytics, and Application Insights to provide a comprehensive, real‑time view of system health.

Responsibilities

Reliability & Incident Management

  • Defining, tracking, and reporting on SLIs, SLOs, and error budgets for all critical services.
  • Designing and maintaining runbooks, escalation paths, and on‑call rotation schedules.
  • Designing chaos engineering practices to proactively surface reliability weaknesses before they impact customers.

Observability — Splunk, Instana & Azure

  • Building and maintaining Splunk searches, dashboards, and alert policies covering application and infrastructure logs.
  • Developing KPIs and unified service‑health views for engineering and leadership.
  • Configuring and extending instrumentation across microservices for distributed tracing and real‑time baselining.
  • Creating smart alerts integrated with on‑call and ticketing systems for automated incident routing.
  • Maintaining Azure Monitor alert rules, action groups, and workbooks across Azure subscriptions.
  • Utilizing Log Analytics workspaces including data‑retention policies and ingestion cost governance using KQL.
  • Leveraging Application Insights for APM, availability testing.
  • Driving convergence of Splunk, Instana, and Azure signals into a unified observability strategy.
  • Building internal tooling in Python, Go, or Bash to eliminate toil and accelerate incident response.
  • Participating in security reviews and threat‑modeling sessions for new platform capabilities.
The Person

The opportunity to cast a vision for success, to blend art and science with proactive strategy and tactical execution, will be a key to success in this role.

  • You’re a self‑starter. In many ways, you’ll help define success metrics for this position. You intuitively understand what will be required to excel at this work and won’t expect others to produce the blueprint for you. You’ll hit the ground running as the architect for this function, helping to ideate and inform where, and how, with what, and by whom site reliability will be achieved.
  • You enjoy wearing multiple hats. At this stage in Unlimited’s growth, our thought leaders have an opportunity to speak into and influence our technology, operations, and processes. Having a green field to work in will mean as you build your role, you’ll touch on others’ work and play in others’ courts. If we agree on successful outcomes, how we get there will be a collective strategy and effort.
  • You’re comfortable with shifting priorities, and some might say you’re an expert at context shifting. Not everyone can sustain the exercise of planning strategically but building reactively. Not every day will present the requirement to shift gears, but some days will present the need to be tactically agile. When that moment comes, you’ll not only be prepared, but you’ll also be ready.
Qualifications

Required

  • 5+ years in an SRE, DevOps, or Platform Engineering role in a production cloud environment.
  • Hands‑on AKS experience: cluster provisioning, upgrades, CNI networking, and workload lifecycle management.
  • Proficiency with Splunk: SPL authoring, dashboard creation, alert configuration, and data onboarding.
  • Working knowledge of Instana APM: agent deployment, custom tracing, alerting, and performance analysis.
  • Solid command of Azure Monitor, Log Analytics (KQL), and Application Insights.
  • Infrastructure‑as‑Code experience with Terraform; familiarity with Helm and Kustomize.
  • Scripting or development proficiency in at least one of: Python, Go, or Bash.
  • Demonstrated ability to drive incident resolution and lead blameless post‑mortems.
  • Clear communicator — able to translate complex technical topics for non‑technical stakeholders.

Preferred

  • Microsoft Certified: Azure Administrator Associate (AZ‑104) or Azure DevOps Engineer Expert (AZ‑400).
  • Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD).
  • Splunk Certified Power User or Splunk Enterprise Certified Architect.
  • Experience with service mesh technologies (Istio, Linkerd) on AKS.
  • Familiarity with FinOps practices and Azure cost‑management tooling.
  • Experience in a regulated environment (SOC 2, ISO 27001, PCI‑DSS, or HIPAA).
Our Process

What you can expect from here:

  • If we feel your experience aligns with our search, we’ll reach out for a discovery call. We’ll talk about our company culture, Unlimited Systems Core Values, this position, and we’ll learn more about your background.
  • From there, we’ll ask you to view some of your previous work as well as request that you complete an assessment. This lets us know whether our expectations for the role actually align with your unique strengths.
  • If we’re still feeling good about each other, we’ll pass you on to our Hiring Manager and then to a panel interview so you can meet more of the faces of Unlimited Systems before we make an offer decision.
EEO Statement

Unlimited Systems is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic under applicable law.

Disclaimer

Applicants may be subject to a background check. Employees in this position must be able to satisfactorily perform the essential functions of the position. If requested, Unlimited Systems will make every effort to provide reasonable accommodation to enable employees with disabilities to perform the position’s essential job duties. As markets change and the Organization grows, job descriptions may change over time as requirements and employee skill levels evolve. With this understanding, Unlimited Systems retains the right to change or assign other duties to this Senior Infrastructure Administrator position.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Support Analyst (Hybrid)
Technical Support Analyst (Hybrid)

Unlimited Systems • Cincinnati (OH)

Hybrid
USD 75,000 - 95,000
DevOps Engineer (Hybrid)
DevOps Engineer (Hybrid)

Unlimited Systems • Cincinnati (OH)

Hybrid
USD 110,000 - 160,000
DevOps Engineer
DevOps Engineer

Unlimited Systems • Cincinnati (OH)

Hybrid
USD 90,000 - 125,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Chandler (AZ)

On-site
USD 140,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

Hybrid
USD 100,000 - 135,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Plano (TX)

On-site
USD 152,000 - 192,000
Discretionary incentive eligible
Benefits eligible
Human Resources Business Partner (Hybrid)
Human Resources Business Partner (Hybrid)

Unlimited Systems • Cincinnati (OH)

Hybrid
USD 110,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mikealbert • Cincinnati (OH)

Hybrid
USD 100,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Chandler (MN)

On-site
USD 100,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Axle • Frederick (MD)

On-site
USD 140,000 - 155,000
Paid Time Off
401K match
Educational Benefits
+5