Senior Site Reliability Engineer II

LexisNexis Risk Solutions FL Inc. Company

San Jose (CA)

On-site

USD 105,000 - 175,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health benefits
401(k) with match
Employee Wellness programs
Disability/Life Insurance

Job summary

LexisNexis Risk Solutions is hiring a Senior Site Reliability Engineer to improve reliability, availability, and performance of production systems. You will lead incident reviews, drive remediation, and collaborate across cloud, on‑prem, and security teams to implement robust operational practices.

Join a team focused on cutting-edge infrastructure, Kubernetes, automation, and monitoring to ensure high-quality service delivery and resilience.

Qualifications

  • 5+ years in Site Reliability Engineering, Systems or DevOps.
  • Bachelor’s degree in Engineering, Computer Science, IT, or equivalent experience.
  • Experience supporting highly available production systems and leading incident reviews.

Responsibilities

  • Lead incident response, postmortems, root-cause analysis and remediation planning.
  • Monitor assigned environments; diagnose issues and coordinate recovery.
  • Design and maintain automation, runbooks, and workflows for provisioning and deployments.
  • Collaborate with development, security, and operations teams to resolve incidents.

Skills

SRE
Kubernetes
Automation
Monitoring
IaC
Scripting

Education

Bachelor’s degree in Engineering/CS/IT

Tools

Kubernetes

Job description

About the Role

The SRE role is responsible for improving the reliability, availability, performance, and operational quality of production systems. This role provides technical input into project plans, schedules, methodologies, and operational strategies across multiple system environments.

Job Functions Lead and participate in incident response, postmortems, root-cause analysis, and gap assessments. Identify reliability, availability, performance, security, and operational risks across production environments. Develop, prioritize, and track corrective and preventive actions through completion. Follow up with engineering, development, security, support, and business stakeholders to ensure timely resolution of incidents and identified gaps.

Respond to system-management alerts and operational exceptions within assigned enterprise systems and product offerings. Provide technical input into project plans, schedules, implementation methodologies, and operational readiness activities. Support the triage, planning, execution, documentation, and closure of changes, service requests, and operational tasks.

Lead or contribute to Operations Team projects involving cloud, on-premises infrastructure, security, Kubernetes, automation, monitoring, and system modernization. Improve production quality and availability by creating new operational capabilities and remediating weaknesses in existing systems and processes.

Qualifications

5+ years of experience in Site Reliability Engineering, Systems Engineering, DevOps, Infrastructure Engineering, or a related field.

Bachelor’s degree in Engineering, Computer Science, Information Technology, or equivalent professional experience.

Demonstrated experience supporting highly available production systems. Experience leading incident reviews, postmortems, root-cause analysis, and remediation planning. Experience working across infrastructure, application, security, and operations teams.

Strong problem-solving, analytical, organizational, and communication skills. Ability to manage multiple priorities and drive work to completion in a fast-paced operational environment.

Technical Skills

Strong experience in Site Reliability Engineering (SRE), production operations, and IT service management processes including incident, problem, change, and service request management.

Hands-on expertise with cloud and on-premises infrastructure, Kubernetes, containerized workloads, virtualization, and distributed systems.

Advanced knowledge of Linux/UNIX and Windows environments, storage and file systems, including installation, configuration, troubleshooting, lifecycle management, backup, disaster recovery, and business continuity.

Experience with monitoring, alerting, logging, observability, and performance analysis, including the ability to analyze system diagnostics, logs, traces, resource utilization, and operational metrics.

Strong automation and infrastructure engineering skills, including Infrastructure as Code (IaC), configuration management, scripting (Python, Shell, PowerShell), system provisioning, deployments, remediation, and security risk mitigation.

Accountabilities

Monitor assigned environments, respond to alerts and incidents, diagnose system and performance issues, and coordinate escalation and recovery.

Track remediation activities and stakeholder commitments through completion to improve production quality, reliability, and availability.

Design and maintain automation, scripts, integrations, runbooks, and workflows for provisioning, health checks, deployments, remediation, and routine operations.

Install, configure, troubleshoot, and support hardware, software, storage, network, cloud, Kubernetes, and other infrastructure services.

Establish logging, monitoring, alerting, metrics, and tracing standards; improve alert quality by reducing noise and ensuring alerts are actionable.

Build dashboards and visualizations that communicate system health, availability, performance, capacity, service-level objectives, and incident trends.

Develop and maintain recovery procedures and participate in disaster-recovery, resilience, and business-continuity exercises.

Partner with development, operations, security, support teams, vendors, and stakeholders to coordinate work, resolve issues, and meet delivery commitments.

Lead or contribute to Operations Team projects from planning and implementation through documentation, transition to support, and closure.

Plan, risk-assess, obtain approval for, implement, document, and close changes, service requests, and operational tasks.

Review and improve technical procedures, scripts, automation, and operational documentation while providing guidance to less-experienced team members.

Working for you

We know that your wellbeing and happiness are key to a long and successful career. These are some of the benefits we are delighted to offer:

Health Benefits: Comprehensive, multi-carrier program for medical, dental and vision benefits

Retirement Benefits: 401(k) with match and an Employee Share Purchase Plan

Wellbeing: Wellness platform with incentives, Headspace app subscription, Employee Assistance and Time-off Programs

Short-and-Long Term Disability, Life and Accidental Death Insurance, Critical Illness, and Hospital Indemnity Family Benefits, including bonding and family care leaves, adoption and surrogacy benefits

Health Savings, Health Care, Dependent Care and Commuter Spending Accounts In addition to annual Paid Time Off, we offer up to two days of paid leave each to participate in Employee Resource Groups and to volunteer with your charity of choice.

U.S. National Base Pay Range: $104,900 - $174,700. Geographic differentials may apply in some locations to better reflect local market rates. This job is eligible for an annual incentive bonus.

We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits.

Click here to access benefits specific to your location.

We are committed to providing a fair and accessible hiring process. If you have a disability or other need that requires accommodation or adjustment, please let us know by completing our Applicant Request Support Form or please contact 1-855-833-5120.

Criminals may pose as recruiters asking for money or personal information. We never request money or banking details from job applicants. Learn more about spotting and avoiding scams here. Please read our Candidate Privacy Policy.

We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law. USA Job Seekers: EEO Know Your Rights.

At LexisNexis® Risk Solutions, our businesses span multiple industries providing customers with innovative technologies, information-based analytics, decisioning tools and data management services that provide market-specific solutions. Approximately 11,100 employees in offices throughout the world support our brands by serving customers in more than 190 countries and territories. LexisNexis® Risk Solutions is part of RELX, a global provider of information and analytics for professional and business customers across industries. For more information, please visit www.risk.lexisnexis.com and www.relx.com.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

RELX INC • Sacramento (CA)

Remote
USD 72,000 - 119,000
Health Benefits
401(k) with match and ESPP
Wellbeing program and Headspace
+3
Software Engineering Lead
Software Engineering Lead

LexisNexis Risk Solutions • Arizona

Hybrid
USD 115,000 - 192,000
Senior Software Engineer I
Senior Software Engineer I

LexisNexis Risk Solutions • Northern (KY)

On-site
USD 87,000 - 144,000
Annual incentive bonus
Country-specific benefits
Systems Engineer III
Systems Engineer III

LexisNexis Risk Solutions • Alpharetta (GA)

On-site
USD 72,000 - 119,000
Software Engineer III
Software Engineer III

LexisNexis Risk Solutions Inc. Company • Alpharetta (GA)

On-site
USD 79,000 - 131,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

LexisNexis Risk Solutions • Alpharetta (GA)

On-site
USD 118,000 - 220,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

LexisNexis Risk Solutions Inc. Company • Alpharetta (GA)

On-site
USD 118,000 - 264,000
Senior Software Engineer I
Senior Software Engineer I

RELX INC • Alexandria (VA)

On-site
USD 87,000 - 144,000
Software Engineer 1
Software Engineer 1

Talent Octopusventures • Alpharetta (GA)

On-site
USD 59,000 - 99,000
Annual incentive bonus
Country specific benefits
Senior Software Engineer I
Senior Software Engineer I

RELX INC • Charleston (WV)

On-site
USD 87,000 - 144,000