Senior Engineer, Reliability

lplfinancial

Austin (TX)

On-site

USD 130,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

LPL Financial seeks a Senior Engineer, Observability and Platform Stability to ensure reliability and performance of observability platforms and related apps. You will lead proactive monitoring, incident response, platform maintenance, automation, and continuous improvement in collaboration with SRE, Platform Engineering, DevOps, and Operations teams.

The ideal candidate has strong cloud experience, excels in production support, and delivers actionable insights through observability capabilities

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.
  • 6+ years in SRE, platform engineering, production support, or enterprise operations.
  • 5+ years supporting enterprise platforms (AWS, Dynatrace, ELK, ServiceNow, SolarWinds).
  • Experience troubleshooting complex production incidents in large environments.
  • Experience with formal change and release management processes.

Responsibilities

  • Partner with owners and engineering teams for operational and release support.
  • Maintain production stability during platform upgrades and initiatives.
  • Triage incidents, root cause analysis, and corrective action planning.
  • Collaborate to implement long-term solutions to reduce incidents.
  • Develop and maintain runbooks, documentation, and release procedures.

Skills

Observability
SRE
Platform stability
Release management
CI/CD
Cloud (AWS)

Education

Bachelor's degree in Computer Science/IT/Engineering

Tools

AWS
Dynatrace
ELK
ServiceNow
SolarWinds

Job description

Where Ambition Meets Innovation

Build a career that matches all your initiative with an impressive dose of innovation. From cutting-edge resources and a collaborative environment to the freedom to make an impact and more, you'll find the ingredients you need at LPL Financial to shape your success while helping clients pursue their financial goals.

Job Overview

The Senior Engineer, Observability and Platform Stability is responsible for ensuring the reliability, availability, and performance of enterprise observability platforms and supporting applications. This role drives operational excellence through proactive monitoring, incident response, platform maintenance, automation, and continuous improvement initiatives. The position partners closely with Site Reliability Engineering (SRE), Platform Engineering, DevOps, Operations, and Incident Management teams to enhance platform stability, mature CI/CD practices, and support strategic technology initiatives. The ideal candidate brings strong cloud technology experience and a proven ability to support production environments while delivering actionable insights through observability capabilities.

Responsibilities
Delivery Support
  • Partner with Product Owners and engineering teams to provide operational and release support for technology initiatives.
  • Ensure platform changes meet operational readiness requirements, including rollback procedures, runbook documentation, integration standards, and support handoffs.
  • Maintain production stability throughout platform upgrades, enhancements, and enterprise initiatives.
  • Support platform ownership transitions and operational readiness activities across global delivery teams.
Production Support & Incident Response
  • Troubleshoot application and platform issues to restore services and minimize business impact.
  • Serve as an escalation point for complex production incidents and operational challenges.
  • Participate in incident triage, root cause analysis, corrective action planning, and resolution activities.
  • Collaborate with engineering teams to implement long-term solutions that reduce recurring incidents.
  • Support mission-critical environments through timely incident response and service restoration.
Platform Stability & Proactive Operations
  • Conduct health checks, configuration reviews, and performance assessments to identify operational risks.
  • Support reliability initiatives focused on improving recovery times, reducing incident recurrence, and minimizing change-related defects.
  • Validate vendor releases, hotfixes, and configuration changes prior to production deployment.
  • Partner with observability and analytics teams to enhance monitoring, alerting, and issue detection capabilities.
Release Execution & CI/CD Maturation
  • Execute platform changes through established SDLC, change management, and release management processes.
  • Collaborate with Platform Engineering, DevOps, Quality Engineering, and Scrum teams to improve release and deployment practices.
  • Support the adoption of source control, environment separation, release automation, and CI/CD capabilities.
  • Ensure solutions are testable, deployable, and operationally supported before and after production implementation.
Documentation & Operational Excellence
  • Maintain runbooks, support documentation, configuration records, incident playbooks, and release procedures.
  • Document incident findings, lessons learned, and process improvement opportunities.
  • Contribute to the development of standardized, repeatable, and scalable operational practices.
What are we looking for?

We seek professionals who pursue greatness , act with integrity , are driven to help our clients succeed , win together , and create and share joy . The ideal candidate brings strong technical expertise in observability platforms, site reliability engineering, production support, release management, and operational excellence while demonstrating a commitment to platform reliability, cross-functional collaboration, and continuous improvement.

Requirements
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
  • 6+ years of experience in SRE, platform engineering, production support, enterprise application operations, or related technology environments.
  • 5+ years of experience supporting enterprise platforms, including AWS, Dynatrace, ELK, ServiceNow, and SolarWinds.
  • Experience troubleshooting complex production incidents within enterprise-scale technology environments.
  • Experience executing technology changes through formal change management and release management processes.
Preferences
  • Experience within financial services or another regulated industry.
  • Experience implementing or supporting CI/CD pipelines, release automation, and deployment processes across development, testing, and production environments.
  • Experience with observability platforms, monitoring tools, performance dashboards, or application monitoring solutions.
  • Experience collaborating
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability & Platform Reliability Engineer
Senior Observability & Platform Reliability Engineer

lplfinancial • Austin (TX)

On-site
USD 130,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Senior Engineer – Reliability
Senior Engineer – Reliability

Jobtailor • South Carolina

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior Platform Engineer - Observability
Senior Platform Engineer - Observability

Capital Group • Charlotte (NC), Northern (KY)

Hybrid
USD 150,000 - 210,000
Senior Platform Engineer - Observability
Senior Platform Engineer - Observability

Capital-Group-1 • Charlotte (NC)

On-site
USD 131,000 - 219,000
Senior Platform Engineer - Observability
Senior Platform Engineer - Observability

504 CGCG-US CG Companies Global-US • Town of Charlotte (NY)

On-site
USD 137,000 - 219,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
Senior Engineer, Reliability
Senior Engineer, Reliability

LPL Financial LLC • Austin (TX)

Hybrid
USD 102,000 - 169,000
401K matching
health benefits
employee stock options
+3