Senior Observability & Platform Reliability Engineer

LPL Financial LLC

Austin (TX)

Hybrid

USD 102,000 - 169,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

401K matching
health benefits
employee stock options
paid time off
volunteer time off
more

Job summary

LPL Financial LLC in Austin, TX seeks a Senior Engineer, Observability and Platform Stability to ensure reliability, availability, and performance of enterprise observability platforms and supporting apps.

You will drive operational excellence through proactive monitoring, incident response, platform maintenance, automation, and CI/CD maturation, collaborating with SRE, Platform Engineering, DevOps, and Operations teams.

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
  • 6+ years of experience in SRE, platform engineering, production support, enterprise application operations, or related technology environments.
  • 5+ years of experience supporting enterprise platforms, including AWS, Dynatrace, ELK, ServiceNow, and SolarWinds.
  • Experience troubleshooting complex production incidents within enterprise‑scale technology environments.
  • Experience executing technology changes through formal change management and release management processes.

Responsibilities

  • Delivery Support Partner with Product Owners and engineering teams to provide operational and release support for technology initiatives.
  • Ensure platform changes meet operational readiness requirements, including rollback procedures, runbook documentation, integration standards, and support handoffs.
  • Maintain production stability throughout platform upgrades, enhancements, and enterprise initiatives.
  • Support platform ownership transitions and operational readiness activities across global delivery teams.
  • Production Support & Incident Response Troubleshoot application and platform issues to restore services and minimize business impact.
  • Serve as an escalation point for complex production incidents and operational challenges.
  • Participate in incident triage, root cause analysis, corrective action planning, and resolution activities.
  • Collaborate with engineering teams to implement long-term solutions that reduce recurring incidents.
  • Support mission‑critical environments through timely incident response and service restoration.
  • Platform Stability & Proactive Operations Conduct health checks, configuration reviews, and performance assessments to identify operational risks.
  • Support reliability initiatives focused on improving recovery times, reducing incident recurrence, and minimizing change‑related defects.
  • Validate vendor releases, hotfixes, and configuration changes prior to production deployment.
  • Partner with observability and analytics teams to enhance monitoring, alerting, and issue detection capabilities.
  • Release Execution & CI/CD Maturation Execute platform changes through established SDLC, change management, and release management processes.
  • Collaborate with Platform Engineering, DevOps, Quality Engineering, and Scrum teams to improve release and deployment practices.
  • Support the adoption of source control, environment separation, release automation, and CI/CD capabilities.
  • Ensure solutions are testable, deployable, and operationally supported before and after production implementation.
  • Documentation & Operational Excellence Maintain runbooks, support documentation, configuration records, incident playbooks, and release procedures.
  • Document incident findings, lessons learned, and process improvement opportunities.
  • Contribute to the development of standardized, repeatable, and scalable operational practices.

Skills

Site Reliability Engineering
Platform Engineering
Production Support
Observability
Incident Response
CI/CD
Change Management
AWS
Dynatrace
ELK
ServiceNow
SolarWinds

Education

Bachelor's degree in CS/IT/Engineering

Tools

AWS
Dynatrace
ELK
ServiceNow
SolarWinds

Job description

LPL Financial LLC in Austin, TX seeks a Senior Engineer, Observability and Platform Stability to ensure reliability, availability, and performance of enterprise observability platforms and supporting apps.

You will drive operational excellence through proactive monitoring, incident response, platform maintenance, automation, and CI/CD maturation, collaborating with SRE, Platform Engineering, DevOps, and Operations teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability Automation Architect
Senior Observability Automation Architect

Q2 • Austin (TX)

Hybrid
USD 180,000 - 240,000
Hybrid Work Opportunities
Flexible Time Off
Career Development & Mentoring
Senior Site Reliability Engineer - Observability
Senior Site Reliability Engineer - Observability

LSEG (London Stock Exchange Group) • Allen (TX)

On-site
USD 140,000 - 190,000
Senior SRE Tech Lead: Reliability & Observability
Senior SRE Tech Lead: Reliability & Observability

LSEG • Allen (TX)

On-site
USD 122,000 - 176,000
Senior SRE: Reliability, Observability & Cloud Platform
Senior SRE: Reliability, Observability & Cloud Platform

LSEG • Raleigh (NC)

On-site
USD 140,000 - 190,000
Healthcare
Retirement planning
Volunteer days
+1
Platform Stability Engineer — Production Support & Release
Platform Stability Engineer — Production Support & Release

LPL Financial LLC • Town of Charlotte (NY), Fort Mill (SC)

On-site
USD 106,000 - 177,000
401(k) matching
Health benefits
Employee stock options
+2
Senior Platform Reliability & Observability Lead
Senior Platform Reliability & Observability Lead

Truist • Raleigh (NC)

On-site
USD 140,000 - 180,000
Medical Insurance
401k Plan
Paid Time Off
Senior Platform Reliability Engineer - Kubernetes & Observability
Senior Platform Reliability Engineer - Kubernetes & Observability

PLP Group • New York (NY)

On-site
USD 75,000 - 130,000
Technical Lead SRE — Reliability & Observability
Technical Lead SRE — Reliability & Observability

LSEG • Raleigh (NC)

On-site
USD 140,000 - 190,000
Senior SRE: Reliability & Observability Leader
Senior SRE: Reliability & Observability Leader

LSEG • Allen (TX)

On-site
USD 150,000 - 190,000
Senior Site Reliability Engineer — Platform & Observability
Senior Site Reliability Engineer — Platform & Observability

Jobtailor • North Carolina

On-site
USD 180,000 - 240,000