Journey-Centric Lead Site Reliability Engineer (SRE)

Solugenix

Phoenix (AZ)

Hybrid

USD 96,000 - 103,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Solugenix in Phoenix, AZ, is seeking a Journey-Centric Lead Site Reliability Engineer to drive end-to-end reliability across critical customer journeys. You will lead the design and implementation of modern SRE practices, unified observability, and autonomous operations for highly resilient digital services.

The ideal candidate brings 5+ years in SRE/Platform Engineering for BFS clients, strong cloud and networking skills, hands-on experience with Dynatrace/New Relic/Splunk, IaC tooling, and AI

Qualifications

  • 5+ years in SRE/Platform/DevOps with BFS customers.
  • 3+ years leading enterprise-scale reliability initiatives.
  • Experience with production support and critical SLOs.
  • Strong knowledge of SRE principles, distributed systems, and cloud-native platforms.
  • Hands-on with monitoring/observability tools: Dynatrace, New Relic, Splunk, OpenTelemetry.
  • Proficient in IaC and automation tools (Terraform, CloudFormation, Ansible).

Responsibilities

  • Define and implement enterprise-scale SRE best practices across critical journeys.
  • Establish reliability frameworks, standards, and governance models.
  • Lead incident management, postmortems, and reliability reviews.
  • Design unified observability across metrics, logs, traces, and user telemetry.
  • Build dashboards and enable end-to-end monitoring for complex microservices.
  • Develop self-healing and autonomous operations with AI Ops approaches.
  • Automation of deployments, monitoring, remediation, and runbooks.

Skills

SRE principles
Distributed systems
Cloud-native platforms
Dynatrace
New Relic
Splunk Observability
Elastic Stack
OpenTelemetry
Terraform
CloudFormation
Ansible
GitHub Actions
Jenkins
Python
Go
Bash/Shell
AI Ops
AWS AgentCore
Networking knowledge
TCP/IP
DNS
HTTP/HTTPS
API Gateways
Load Balancers
CDN
Service Mesh
VPC & cloud networking

Tools

Dynatrace
New Relic
Splunk Observability
Elastic Stack
OpenTelemetry
Terraform
CloudFormation
Ansible
GitHub Actions
Jenkins
Python
Go

Job description

Journey-Centric Lead Site Reliability Engineer (SRE)

Phoenix, AZ (Hybrid)

JPC - 20640

We are seeking a highly experienced Journey-Centric Lead Site Reliability Engineer (SRE) with Banking and Financial Services (BFS) domain experience to drive end-to-end reliability, observability, automation, and operational excellence across critical customer and business journeys. This role will lead to the design and implementation of modern SRE practices, unified observability platforms, self-healing capabilities, AI-driven operations, and workflow automation to ensure highly resilient, scalable, and intelligent digital services.

The ideal candidate combines deep SRE expertise with strong platform engineering, cloud operations, networking, observability, and AI Ops experience, with a particular focus on AWS AgentCore-powered operational intelligence and autonomous operations.

Qualifications:

  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, or DevOps with Banking and Financial Services customers.
  • 3+ years leading enterprise-scale reliability transformation initiatives.
  • Experience managing mission-critical digital platforms and customer journeys.
  • Experience with production support and handling critical SLOs.
  • Strong knowledge of:
  • SRE principles and practices
  • Reliability engineering frameworks
  • Distributed systems
  • Cloud-native platforms
  • Hands-on experience with one or more:
  • Dynatrace
  • New Relic
  • Splunk Observability
  • Elastic Stack
  • OpenTelemetry
  • Automation & Engineering experience with:
  • Terraform
  • CloudFormation
  • Ansible
  • GitHub Actions
  • Jenkins
  • Python
  • Go
  • Bash/Shell scripting
  • AI Ops & GenAI - Experience implementing AI Ops solutions.
  • Knowledge of autonomous operations and intelligent remediation systems.
  • Hands-on experience with AWS AgentCore and agent-based operational platforms.
  • Familiarity with AI-powered observability and incident management tools.
  • Networking Knowledge - Strong understanding of Network Layer:
  • TCP/IP
  • DNS
  • HTTP/HTTPS
  • API Gateways
  • Load Balancers
  • CDN
  • Service Mesh
  • Network Performance Engineering

Preferred Qualifications (Desired)

  • Experience in customer journey monitoring and digital experience.

Responsibilities:

SRE & Reliability Engineering

  • Define and implement enterprise-scale SRE best practices across critical applications and digital journeys.
  • Establish reliability frameworks, operational standards, and governance models.
  • Drive proactive reliability engineering initiatives to improve system availability, resilience, and performance.
  • Lead incident management, postmortem analysis, root cause investigations, and reliability reviews.

Unified Observability & Monitoring

  • Design and implement a unified observability strategy encompassing metrics, logs, traces, events, and user experience telemetry.
  • Build comprehensive observability dashboards for business and technology stakeholders.
  • Implement distributed tracing and end-to-end monitoring across complex microservices ecosystems.
  • Define observability standards and instrumentation frameworks across engineering teams.
  • Enable unified logs, metrics, and trace correlation capabilities for rapid issue detection and troubleshooting.
  • Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through observability-driven insights.
  • Establish service dependency mapping and journey-centric operational visibility.

Self-Healing & Autonomous Operations

  • Design and implement self-healing capabilities using event-driven automation and AI-assisted remediation.
  • Develop automated recovery processes for common failure scenarios.
  • Create autonomous operational workflows that minimize manual intervention.
  • Integrate predictive alerting and automated response mechanisms.

Automation & Workflow Engineering

  • Build scalable operational automation frameworks.
  • Develop infrastructure, application, observability, and operational workflows using Infrastructure as Code (IaC), Monitoring as Code (MaC), and Observability as Code (OaC).
  • Automate deployments, monitoring, remediation, and operational runbooks.
  • Reduce operational toil through intelligent engineering solutions.

SLO, SLA & Error Budget Management

  • Define and govern measurable Service Level Objectives (SLOs), Service Level Agreements (SLAs), and Error Budgets.
  • Partner with engineering and business teams to align reliability targets with customer expectations.
  • Establish service maturity metrics and reliability scorecards.
  • Drive data-driven operational decision-making through reliability KPIs.

Network & Platform Reliability

Apply deep understanding of:

  • Understanding of network layer to troubleshoot critical bandwidth/latency issues
  • No need to pass all these:
  • TCP/IP
  • DNS
  • CDN
  • API Gateway architectures
  • Service Mesh technologies
  • VPC and cloud networking
  • Troubleshoot complex network performance and availability issues.
  • Ensure end-to-end reliability across cloud and hybrid environments.

AI Ops & AWS AgentCore

  • Implement and operationalize AI Ops platforms and autonomous operations capabilities.
  • Leverage AWS AgentCore to build intelligent operational agents for functions like:
  • Root cause analysis
  • Predictive remediation
  • Capacity forecasting
  • Automated operational workflows
  • Drive adoption of GenAI-powered operational intelligence across the enterprise.
  • Integrate AI-assisted observability, automation, and service management solutions.

Pay Range for CA, CO, IL, NJ, NY, WA, and DC: $70/hour to $75/hour. Starting rate of pay offered may vary depending on factors including but not limited to, position offered, location, education, training and/or experience.

Solugenix will consider qualified applicants with a criminal history pursuant to the California Fair Chance Act and Ordinance. Applicants do not need to disclose their criminal history or participate in a background check until a conditional job offer is made to you. After making a conditional offer and running a background check, if we are concerned about conviction that is directly related to the job, applicants will be given the chance to explain the circumstances surrounding the conviction, provide mitigating evidence, or challenge the accuracy of the background report.

About Solugenix

Solugenix is a leader in IT services, delivering cutting-edge technology solutions, exceptional talent, and managed services to global enterprises. With extensive expertise in highly regulated and complex industries, we are a trusted partner for integrating advanced technologies with streamlined processes. Our solutions drive growth, foster innovation, and ensure compliance providing clients with reliability and a strong competitive edge.

Recognized as a 2024 Top Workplace, Solugenix is proud of its inclusive culture and unwavering commitment to excellence. Our recent expansion, with new offices in the Dominican Republic, Jakarta, and the Philippines, underscores our growing global presence and ability to offer world-class technology solutions. Partnering with Solugenix means more than just business it means having a dedicated ally focused on your success in today's fast-evolving digital world.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. IT Auditor
Sr. IT Auditor

Solugenix • El Monte (CA)

Remote
USD 70,000 - 110,000
Cloud Architect/Principal Engineer
Cloud Architect/Principal Engineer

Solugenix • Irvine (CA)

On-site
USD 131,000 - 158,000
Category Manager (Global Real Estate & Construction Procurement)
Category Manager (Global Real Estate & Construction Procurement)

Solugenix • New York (NY)

On-site
USD 117,000 - 167,000
Category Manager (Global Real Estate & Construction Procurement)
Category Manager (Global Real Estate & Construction Procurement)

Solugenix • Charlotte (NC)

On-site
USD 124,000 - 172,000
CI/CD Engineer
CI/CD Engineer

Solugenix • Irvine (CA)

On-site
Journey-Driven SRE Lead — AI-Powered Reliability
Journey-Driven SRE Lead — AI-Powered Reliability

Solugenix • Phoenix (AZ)

Hybrid
USD 96,000 - 103,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Director, ECMS
Director, ECMS

Solugenix • New York (NY)

Hybrid
USD 205,000 - 260,000
Inclusive culture
Recognition as a top workplace
Professional development opportunities
Genesys Cloud Engineer
Genesys Cloud Engineer

Solugenix • San Antonio (TX)

On-site
USD 229,233,000 - 243,560,000
Lead SRE/DevOps Engineer
Lead SRE/DevOps Engineer

Synechron • Dallas (TX)

Hybrid
USD 125,000 - 135,000
Medical insurance
401(k)
Paid maternity leave
+3