Sr. Software Engineer, Site Reliability

JobCubby

Indianapolis (IN)

Hybrid

USD 115,000 - 150,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
Vision and dental insurance
401k matching
Fully remote in US

Job summary

Bloomerang is seeking an experienced Site Reliability Engineer to evolve our SRE practices, focusing on reliability, observability, and automation across services, databases, and infrastructure. You will own complex production issues, drive proactive reliability, and collaborate with Software Engineering, Product, Support, and DevOps in a fully remote role within the United States.

This is a hands-on, cross-functional role with on-call rotation and a strong emphasis on improving customer

Qualifications

  • Hands-on SRE experience applying software engineering practices to production reliability and helping establish or mature SRE practices.
  • Strong knowledge of SLIs, SLOs, error budgets, observability, automation, and toil reduction.
  • Experience building monitoring, dashboards, alerts, and telemetry using tools such as Honeycomb, New Relic, Grafana, CloudWatch, Kibana, or similar.
  • Experience with production incident management, root cause analysis, blameless post-incident reviews, and corrective-action follow-through.
  • Strong programming and scripting skills to navigate and troubleshoot application code and build automation and operational tooling.
  • Strong SQL and relational database skills for production troubleshooting; PostgreSQL experience preferred.
  • Experience troubleshooting cloud-hosted applications using source code, logs, APIs, telemetry, event streams, and databases.
  • Comfort navigating application stacks across technologies such as PHP, .NET, and Node.js.

Responsibilities

  • Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
  • Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions to recurring problems and defects.
  • Bring proven SRE practices to the team and foster proactive reliability, continuous improvement, automation, and shared ownership.
  • Lead incident response from triage and mitigation through recovery, root cause analysis, and blameless post-incident reviews, turning lessons learned into reliability improvements.
  • Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.
  • Define and mature SLIs and SLOs that measure system reliability and customer experience.
  • Develop synthetic monitoring for critical customer journeys to detect failures before they impact customers.
  • Identify sources of recurring operational toil and drive automation, tooling, process improvements, or permanent fixes that reduce manual effort.
  • Use AI-assisted tools and source code repositories to accelerate triage, troubleshooting, code analysis, automation, and technical investigation.
  • Participate in a rotating on-call schedule, primarily during business hours, with limited after-hours and weekend support.

Skills

SRE experience
Observability
Automation
Troubleshooting
Programming & scripting

Tools

Honeycomb
New Relic
Grafana
CloudWatch
Kibana

Job description

At Bloomerang, we believe change happens on purpose. We champion the power and potential of nonprofits, igniting next-level impact with the team and technology built for purpose. Our powerful giving platform and stellar support enable tens of thousands of nonprofits to raise more, recruit more, and retain more, fueling maximum impact and raising the bar on what’s possible for the nonprofit sector. That's why, even as the nonprofit sector sees declines in giving, Bloomerang customers raise more year over year.

We're also in the business of creating thriving employees. Join a mission-driven culture built on our core values of Simplify, Care and Act. We know our people are the key to our success, and we're proud to be home to some of the most innovative and skilled individuals in the workforce today. Come feel invigorated and unstoppable with us!

The Role

We are evolving our Tier 3 Support Engineering team into a Site Reliability Engineering (SRE) organization focused on improving product reliability, observability, and operational efficiency. We're looking for an experienced SRE who brings strong software engineering fundamentals and is excited to help shape this transformation and mature our SRE practices.

This is a highly collaborative, hands-on role working across application code, telemetry, databases, APIs, and infrastructure to diagnose complex production issues and improve reliability. Production support and product defects remain part of today's work as you help reduce reactive effort through observability, SLOs, automation, permanent fixes, and proactive reliability engineering.

What Success Looks Like

Success isn't measured solely by issues resolved, but by issues that no longer require manual intervention. You'll help detect problems earlier, reduce recurring issues and toil, strengthen incident response, and create more capacity for proactive reliability engineering.

What You Will Do
  • Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
  • Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions to recurring problems and defects.
  • Bring proven SRE practices to the team and foster proactive reliability, continuous improvement, automation, and shared ownership.
  • Lead incident response from triage and mitigation through recovery, root cause analysis, and blameless post-incident reviews, turning lessons learned into reliability improvements.
  • Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.
  • Define and mature SLIs and SLOs that measure system reliability and customer experience.
  • Develop synthetic monitoring for critical customer journeys to detect failures before they impact customers.
  • Identify sources of recurring operational toil and drive automation, tooling, process improvements, or permanent fixes that reduce manual effort.
  • Use AI-assisted tools and source code repositories to accelerate triage, troubleshooting, code analysis, automation, and technical investigation.
  • Participate in a rotating on-call schedule, primarily during business hours, with limited after-hours and weekend support.
What You Need to Succeed
SRE Experience & Transformation
  • Hands-on Site Reliability Engineering experience applying software engineering practices to production reliability and helping establish or mature SRE practices.
  • Strong knowledge of SLIs, SLOs, error budgets, observability, automation, and toil reduction.
Observability & Incident Management
  • Experience building monitoring, dashboards, alerts, and telemetry using tools such as Honeycomb, New Relic, Grafana, CloudWatch, Kibana, or similar.
  • Experience with production incident management, root cause analysis, blameless post-incident reviews, and corrective-action follow-through.
Technical Depth
  • Strong programming and scripting skills to navigate and troubleshoot application code and build automation and operational tooling.
  • Strong SQL and relational database skills for production troubleshooting and safe data correction; PostgreSQL experience preferred.
  • Strong code literacy and debugging skills, including navigating unfamiliar codebases, understanding application flow, reviewing code and change history, and identifying potential reliability issues.
  • Experience troubleshooting cloud-hosted applications using source code, logs, APIs, telemetry, event streams, and databases. Comfort navigating application stacks across technologies such as PHP, .NET, and Node.js; deep expertise in each is not required.
AI, Ownership & Collaboration
  • Demonstrated experience using and embracing AI-assisted tools in day-to-day engineering workflows, including troubleshooting, code analysis, scripting, and automation.
  • A collaborative self-starter who tackles difficult problems, adapts to changing priorities, and challenges the status quo.
  • Strong communication and collaboration skills across Software Engineering, Product, Support, DevOps, and other technical teams.
Benefits
Health + Wellness

You’ll have access to generous health, vision, and dental insurance options as well as HealthiestYou, a healthcare service that offers convenient, confidential access to quality doctors 24/7, anytime, anywhere.

Time Off

You’ll get a competitive PTO package that includes 20 PTO days, 3 flex days, 4 optional volunteer days, 12 paid holidays, as well as paid parental leave. More is more!

401k

You’ll receive a 401k match to help invest in your future.

Equipment

Everything you need to be successful, shipped right to your door. You got this. We got you.

Compensation

The salary range for this position is $114,800 - $150,000. You may also be eligible for a discretionary bonus. Actual compensation within the range will be dependent on your skills, experience, qualifications, and location, as well as applicable employment laws

Location

This is a permanent, full-time, fully remote position (within the U.S. and select Canadian Provinces only). Employees living in Indianapolis, IN are welcome to work from our company headquarters. We do not offer Visa sponsorship or relocation assistance at this time.

Accommodations

Applicants who require accommodations may contact careers@bloomerang.com to request an accommodation in completing an application.

Bloomerang is an Equal Opportunity Employer. Individuals seeking employment at Bloomerang are considered without regard to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, or sexual orientation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Software Engineer (Tech Lead) - Typescript/React
Sr. Software Engineer (Tech Lead) - Typescript/React

Bloomerang • United States

Remote
USD 115,000 - 191,000
Health, vision, and dental insurance
PTO + holidays
401k match
+1
Sr. Software Engineer (Tech Lead) - Typescript/React
Sr. Software Engineer (Tech Lead) - Typescript/React

Far Coder • Northern (KY)

Hybrid
USD 115,000 - 191,000
Health insurance
401k matching
Equipment provided
+2
Sr. Account Executive
Sr. Account Executive

Bloomerang • United States

On-site
USD 135,000
Generous health, vision, and dental insurance
Competitive PTO package
401k match
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Senior Site Reliability Engineer
Senior Site Reliability Engineer

National Black MBA Association • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Annual discretionary plan
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Madrona Venture Labs • United States

Remote
USD 120,000 - 150,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Jobot • Erie

Remote
USD 165,000 - 190,000
Flexible paid time off
Affordable health, dental, and vision insurance
Monthly fitness reimbursement
+4
Forward Deployed Engineer - Strategic Search Accounts
Forward Deployed Engineer - Strategic Search Accounts

Bloomreach • United States

On-site
USD 135,000 - 175,000
Health care including medical, dental,
Vision insurance
401k with employer contribution
+2
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Technical Consultant
Senior Technical Consultant

BloomReach Inc. • Northern (KY)

Hybrid
USD 100,000 - 130,000
Health insurance
401k with employer contribution