- Participate in design reviews, sprint zero, and delivery planning to define and validate reliability requirements
- Collaborate with Major Release Management to ensure releases meet SRE standards and support-readiness requirements
- Define and improve monitoring, observability, dashboards, telemetry coverage, and alert strategy
- Assist in major incident response and root cause analysis
- Drive automation, intelligent tooling, and AI-assisted remediation
- Serve as the operational readiness authority before production releases
- Lead capacity, performance, workload trend, and resiliency analysis
- Establish and track reliability metrics including availability, incident volume, MTTx, alert quality, automation coverage, and change failure rate
- Participate in application reliability governance and service reviews
- Prepare executive reporting on reliability posture, release readiness, observability maturity, incident trends, risks, and improvement outcomes
- Promote SRE practices through mentoring, standards adoption, best-practice sharing, and approved AI tools
Requirements
- Minimum of 10+ years of related technical experience across application support engineering, software engineering, site reliability engineering, production support, or application operations
- Bachelor’s degree preferred or equivalent practical experience
- Experience supporting business-critical applications in production environments
- SRE, observability, automation, or ITIL certifications are a plus
- Proven experience in one or more in-scope roles including Application Support Engineer, SDET, Software Engineer, or SRE
- Strong understanding of monitoring and observability platforms, including dashboard design, alert tuning, telemetry coverage, log analysis, metrics, traces, and event correlation
- Programming or scripting proficiency in one or more languages such as Python, Java, Go, PowerShell, or similar
- Familiarity with distributed applications, middleware, messaging, batch processing, real-time processing, and production application behavior in high-availability environments
- Experience in financial services, capital markets, regulated environments, or other high-availability operational settings
- Demonstrated participation in disaster recovery, performance testing, resiliency testing, release readiness, incident response, and root cause analysis
- Knowledge of AI concepts, data platforms, anomaly detection, incident correlation, and intelligent automation use cases
- Strong collaboration skills across application support, application development, release management, risk, security, business, and vendor stakeholders
- Ability to translate production support insights into actionable engineering improvements
Core Competencies
Demonstrates expertise in Site Reliability Engineering (SRE) practices, including monitoring, observability, and automation, while effectively collaborating with cross-functional teams to enhance production readiness and incident response. Proven ability to analyze performance metrics and drive improvements in high-availability environments.
Highest-signal resume keywords
- Site Reliability Engineering (SRE)
- Monitoring And Observability
- Automation And Intelligent Tooling
- Programming Proficiency In Python, Java, Go, Or PowerShell
- Experience In Financial Services Or Regulated Environments
ATS Optimization Keywords
Hard Skills
- Monitoring And Observability Platforms
- Dashboard Design
- Alert Tuning
- Telemetry Coverage
- Log Analysis
- Metrics And Traces
- Event Correlation
- Disaster Recovery
- Performance Testing
- Resiliency Testing
Soft Skills
- Strong Collaboration Skills
- Mentoring And Best-Practice Sharing
Certifications & Qualifications
- SRE Certification
- Observability Certification
- Automation Certification
- ITIL Certification
Industry Keywords
- Application Support Engineering
- Software Engineering
- Production Support
- Application Operations
- High-Availability Environments
- Capital Markets
- Regulated Environments
Tools & Technologies
- AI Concepts
- Data Platforms
- Anomaly Detection
- Incident Correlation
- Intelligent Automation