- Partner with Product Owners and engineering teams to provide operational and release support for technology initiatives
- Ensure platform changes meet operational readiness requirements, including rollback procedures, runbook documentation, integration standards, and support handoffs
- Maintain production stability throughout platform upgrades, enhancements, and enterprise initiatives
- Support platform ownership transitions and operational readiness activities across global delivery teams
- Troubleshoot application and platform issues to restore services and minimize business impact
- Serve as an escalation point for complex production incidents and operational challenges
- Participate in incident triage, root cause analysis, corrective action planning, and resolution activities
- Collaborate with engineering teams to implement long-term solutions that reduce recurring incidents
- Conduct health checks, configuration reviews, and performance assessments to identify operational risks
- Validate vendor releases, hotfixes, and configuration changes prior to production deployment
- Enhance monitoring, alerting, and issue detection capabilities with observability and analytics teams
- Execute platform changes through SDLC, change management, and release management processes
- Improve release and deployment practices with Platform Engineering, DevOps, Quality Engineering, and Scrum teams
- Support source control, environment separation, release automation, and CI/CD adoption
- Maintain runbooks, support documentation, configuration records, incident playbooks, and release procedures
- Document incident findings, lessons learned, and process improvement opportunities
- Contribute to standardized, repeatable, and scalable operational practices
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field
- 6+ years of experience in SRE, platform engineering, production support, enterprise application operations, or related technology environments
- 5+ years of experience supporting enterprise platforms, including AWS, Dynatrace, ELK, ServiceNow, and SolarWinds
- Experience troubleshooting complex production incidents within enterprise-scale technology environments
- Experience executing technology changes through formal change management and release management processes
- Experience within financial services or another regulated industry
- Experience implementing or supporting CI/CD pipelines, release automation, and deployment processes across development, testing, and production environments
- Experience with observability platforms, monitoring tools, performance dashboards, or application monitoring solutions
- Experience collaborating with offshore, nearshore, or global delivery teams
Core Competencies
Demonstrates expertise in platform engineering and production support, with a strong focus on operational readiness, incident management, and CI/CD practices. Proficient in collaborating with cross-functional teams to enhance system stability and implement effective monitoring and release strategies.
Highest-signal resume keywords
- Platform Engineering
- Incident Management
- CI/CD Pipeline Implementation
- AWS Experience
- Change Management
ATS Optimization Keywords
Hard Skills
Soft Skills
- Collaboration
- Problem-Solving
- Communication
Industry Keywords
- Financial Services
- Regulated Industry
Tools & Technologies
- AWS
- Dynatrace
- ELK
- ServiceNow
- SolarWinds
- Monitoring Tools
- Observability Platforms