Site Reliability Engineering Lead

LexisNexis Risk Solutions Inc. Company

Alpharetta (GA)

On-site

USD 118,000 - 264,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

LexisNexis Risk Solutions Inc. is seeking a Site Reliability Engineering Lead to provide strategic direction for SRE initiatives across insurance tech platforms. You will lead a team of SREs to ensure reliability, scalability, and security for mission-critical applications, partnering with engineering, security, and product teams.

The role emphasizes cloud modernization, automation, and building resilient systems, with strong leadership and collaboration across multiple portfolios.

Qualifications

  • 8+ years in Cloud Engineering, DevOps, or SRE roles in large-scale environments.

Responsibilities

  • Lead and mentor a team of Site Reliability Engineers.
  • Define and drive SRE strategy, standards, and operational frameworks.
  • Improve application reliability, scalability, security, and performance across product portfolios.
  • Establish and maintain SLOs/SLIs and error budgets.
  • Lead incident management, root cause analysis, and post-incident reviews.
  • Drive cloud modernization and migrations to Azure, AWS, and containerized environments.
  • Champion automation and IaC using Terraform, Ansible, and related tools.
  • Develop observability strategies with metrics, logs, traces, and dashboards.

Skills

SRE
Cloud engineering
DevOps
Azure
AWS
Kubernetes
Terraform
Ansible
Observability
Python/Go

Education

Bachelor's degree in CS/Engineering

Tools

Grafana
Prometheus
OpenTelemetry
Splunk
Dynatrace
Datadog
Terraform
Ansible

Job description

About the Business

LexisNexis Risk Solutions is the essential partner in the assessment of risk. Within our Insurance vertical, we provide customers with solutions and decision tools that combine public and industry specific content with advanced technology and analytics to assist them in evaluating and predicting risk and enhancing operational efficiency. Our insurance risk solutions help drive better data-driven decisions across the insurance policy lifecycle, all while reducing risk. You can learn more about LexisNexis Risk at https://risk.lexisnexis.com/insurance

About our Team

TheICS (Insurance Core Services) team is responsible for establishing and driving reliability, observability, automation, and operational excellence standards across Insurance technology platforms. The team partners closely with application, infrastructure, database, and cloud engineering teams to improve platform availability, scalability, performance, and resilience. The SRE Lead will play a key role in shaping reliability strategy, mentoring engineers, driving cross-functional initiatives, and partnering with business and technology stakeholders to improve service reliability and operational maturity across the organization.

The ICU leads strategic initiatives including SLO/SLI implementation, observability platform adoption, cloud modernization, operational readiness reviews, performance engineering, and reliability automation. The team also develops reusable engineering frameworks, standards, and best practices that enable product teams to build and operate highly reliable cloud-native services at scale.

About the Role

As a Site Reliability Engineering Lead, you will provide technical leadership and strategic direction for Site Reliability Engineering initiatives across multiple product portfolios. You will lead a team of SREs responsible for ensuring the reliability, scalability, security, performance, and operational excellence of mission-critical applications and platforms. The SRE Lead will partner closely with engineering, architecture, security, operations, and business stakeholders to drive cloud modernization, operational maturity, observability excellence, automation, and continuous improvement. This role combines hands‑on technical expertise with people leadership, mentoring, strategic planning, and cross‑functional collaboration.

Responsibilities

Lead and mentor a team of Site Reliability Engineers, fostering a culture of ownership, operational excellence, collaboration, and continuous learning. Define and drive SRE strategy, standards, best practices, and operational frameworks across engineering organizations. Partner with product and platform teams to improve application reliability, scalability, security, performance, and resilience. Establish and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets. Lead major incident management, root cause analysis, problem management, and post‑incident review processes. Drive cloud modernization initiatives and support application migrations to Azure, AWS, and containerized environments. Champion automation and Infrastructure as Code (IaC) practices using tools such as Terraform, GitHub, GitLab, Jenkins, and Ansible. Develop and implement observability strategies utilizing metrics, logs, traces, alerting, and dashboards. Collaborate with security and compliance teams to ensure platform adherence to enterprise security and regulatory requirements. Lead architecture reviews and provide guidance on cloud‑native and highly resilient application designs. Drive capacity planning, performance optimization, cost management, and operational efficiency initiatives. Establish engineering guardrails, governance controls, and deployment standards for production environments. Support organizational transformation toward DevOps and SRE practices. Manage operational risk and ensure business continuity and disaster recovery preparedness. Collaborate with stakeholders to prioritize reliability improvements and platform investments. Build and maintain strong relationships with product owners, engineering leaders, vendors, and business partners.

Leadership Responsibilities

Lead, coach, mentor, and develop a high‑performing team of Site Reliability Engineers. Conduct resource planning and support hiring, onboarding, and career development activities. Establish team objectives aligned with business and technology strategies. Promote accountability, innovation, and operational excellence within the team. Act as a trusted advisor and subject matter expert for reliability engineering across the organization. Drive cross‑team collaboration and alignment on strategic initiatives.

Essential Skills and Attributes

Strong leadership experience managing technical engineering teams. Deep expertise in Site Reliability Engineering, DevOps, Cloud Engineering, or Platform Engineering disciplines. Extensive experience with Azure and/or AWS cloud platforms. Strong understanding of Kubernetes, AKS, EKS, containerization, Docker, and cloud‑native architectures. Expertise with Infrastructure as Code tools such as Terraform and Ansible. Strong background in observability platforms such as Grafana, Prometheus, OpenTelemetry, Splunk, Dynatrace, Datadog, or similar technologies. Experience managing large‑scale production environments with stringent availability requirements. Strong understanding of security, compliance, networking, and cloud governance principles. Experience designing highly available, fault‑tolerant, and resilient systems. Strong proficiency in at least one scripting or programming language such as Python, Go, PowerShell, Bash, or C#. Experience with CI/CD pipelines and software delivery automation. Exceptional troubleshooting and problem‑solving capabilities. Excellent communication and stakeholder management skills. Strong documentation and presentation skills. Ability to influence technical direction across multiple engineering organizations.

Desired Skills

Experience building and managing enterprise‑scale observability platforms. Knowledge of FinOps, cloud cost optimization, and operational efficiency practices. Experience with secret management platforms such as HashiCorp Vault, Akeyless or cloud‑native alternatives. Familiarity with Chaos Engineering and resilience testing. Experience supporting regulated environments and compliance frameworks. Experience leading cloud transformation and modernization programs. Agile and Lean delivery experience. Experience supporting global, distributed engineering teams.

Qualifications

8+ years of experience in Cloud Engineering, DevOps, Platform Engineering, Infrastructure Engineering, or Site Reliability Engineering. 2+ years of leadership or people management experience leading engineering teams. Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent practical experience. Azure, AWS, Kubernetes, Terraform, or related certifications preferred. Proven track record leading reliability and operational excellence initiatives in large‑scale enterprise environments.

Salary & Benefits

U.S. National Base Pay Range: $118,300 - $219,800. Geographic differentials may apply in some locations to better reflect local market rates. Base Pay Range for CO is $118,300 - $219,800. Base Pay Range for IL is $124,200 - $230,800. Base Pay Range for Chicago, IL is $130,200 - $241,800. Base Pay Range for MD is $124,200 - $230,800. Base Pay Range for NY is $130,200 - $241,800. Base Pay Range for New York City is $142,000 - $263,800. Base Pay Range for Rochester, NY is $118,300 - $219,800. Base Pay Range for OH is $112,400 - $208,800. Base Pay Range for NJ is $149,765- $239,235. This job is eligible for an annual incentive bonus. Application deadline is 12/24/2026. We are delighted to offer country specific benefits. We are committed to providing a fair and accessible hiring process. If you have a disability or other need that requires accommodation or adjustment, please let us know by completing our Applicant Request Support Form or please contact 1-855-833-5120. We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law. USA Job Seekers: EEO Know Your Rights.

RELX is a global provider of information‑based analytics and decision tools for professional and business customers, enabling them to make better decisions, get better results and be more productive. Our purpose is to benefit society by developing products that help researchers advance scientific knowledge; doctors and nurses improve the lives of patients; lawyers promote the rule of law and achieve justice and fair results for their clients; businesses and governments prevent fraud; consumers access financial services and get fair prices on insurance; and customers learn about markets and complete transactions. Our purpose guides our actions beyond the products that we develop. It defines us as a company. Every day across RELX our employees are inspired to undertake initiatives that make unique contributions to society and the communities in which we operate.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineering Lead
Site Reliability Engineering Lead

LexisNexis Risk Solutions • Alpharetta (GA)

On-site
USD 118,000 - 220,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

LexisNexis Risk Solutions • Town of Florida (NY)

On-site
USD 118,000 - 264,000
Medical Inpatient and Outpatient
Life Assurance Policies
Flexible Benefits Plan
+3
DBA Lead
DBA Lead

LexisNexis Risk Solutions Inc. Company • Alpharetta (GA)

On-site
USD 115,000 - 192,000
Annual incentive bonus
Site Reliability Engineering Lead
Site Reliability Engineering Lead

LexisNexis Special Services Inc. • Alpharetta (GA), Northern (KY)

Hybrid
USD 118,000 - 264,000
Annual incentive bonus
Director Product Mgmt
Director Product Mgmt

LexisNexis Risk Solutions • Boca Raton (FL), Northern (KY)

On-site
USD 136,000 - 253,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

Talent Octopusventures • United States

Hybrid
USD 118,000 - 220,000
Medical Insurance
Life Assurance Policies
Flexible Benefits Plan
+3
Product Mgr II
Product Mgr II

LexisNexis Risk Solutions • Northern (KY)

Hybrid
USD 95,000 - 159,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

LexisNexis Risk Solutions • Northern (KY)

On-site
USD 118,000 - 220,000
Medical insurance
Life insurance
Flexible benefits plan
+3
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

LexisNexis Risk Solutions • New York (NY)

On-site
USD 105,000 - 175,000
Competitive compensation
Flexible work location
Ownership of production systems
+1
Senior Software Engineer I
Senior Software Engineer I

RELX INC • Columbia (SC)

On-site
USD 87,000 - 144,000
Annual incentive bonus
Benefits package