- Drive the reliability, availability, and scalability of mission-critical applications supporting Voice Self-Service and Automation capabilities.
- Lead and support a blended team of employees and consultants operating across onshore and offshore locations in a 24x7 environment.
- Implement proactive monitoring, alerting, and observability solutions to identify and resolve issues before they impact customers or business operations.
- Conduct root cause analysis of application and system failures, implementing sustainable corrective actions to prevent recurrence.
- Develop automation solutions that streamline deployments, configuration management, operational processes, and application health checks.
- Partner with engineering teams to design, deploy, and support resilient solutions that maximize system stability and performance.
- Advance self-healing capabilities that reduce manual intervention and improve overall service availability.
- Analyze availability, reliability, and performance metrics to identify improvement opportunities and drive operational excellence.
- Participate in major incident management, service restoration activities, and post-incident reviews impacting the STP organization.
- Champion best practices in cloud operations, automation, observability, and platform resiliency.
Requirements
- 6+ years of experience supporting enterprise applications in development, site reliability engineering, operations, or related technology roles
- 3+ years of hands-on Python development experience
- Experience managing highly available, scalable, and resilient applications in production environments
- Experience defining, measuring, and improving key performance indicators (KPIs), service level objectives (SLOs), and operational metrics
- Hands-on experience with AWS services, including IAM, EC2, S3, CloudWatch, Lambda, Step Functions, SQS, SNS, Glue, Athena, and related technologies
- Experience implementing monitoring, alerting, and operational support solutions
- Strong problem-solving, analytical, and troubleshooting skills
- Proven ability to collaborate effectively across engineering, operations, and business teams
- Experience using enterprise monitoring and logging tools, including Dynatrace and Splunk
- Strong verbal and written communication skills.
- Experience with containerization and orchestration technologies such as Docker and Kubernetes (preferred)
- Experience implementing Infrastructure as Code (IaC), preferably using Terraform (preferred)
- Experience with Databricks and Snowflake platforms (preferred)
- Experience with CI/CD tools such as GitHub Actions, Jenkins, or similar automation platforms (preferred)
- AWS certification (preferred)
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field (preferred)
Core Competencies
Demonstrates expertise in driving the reliability, availability, and scalability of mission-critical applications, with a strong focus on automation, monitoring, and operational excellence. Proficient in collaborating across teams to implement resilient solutions and improve service availability through data-driven analysis.
Highest-signal resume keywords
- Site Reliability Engineering
- Python Development
- AWS Services
- Monitoring and Alerting Solutions
- Infrastructure as Code (IaC)
ATS Optimization Keywords
Hard Skills
- Site Reliability Engineering
- Python Development
- Monitoring and Alerting Solutions
- Infrastructure as Code (IaC)
- Containerization
- Orchestration Technologies
- Key Performance Indicators (KPIs)
- Service Level Objectives (SLOs)
- Automation Solutions
- Operational Metrics
Soft Skills
- Problem-Solving
- Analytical Skills
- Collaboration
- Communication Skills
- Troubleshooting Skills
Certifications & Qualifications
Industry Keywords
- Cloud Operations
- Automation
- Observability
- Platform Resiliency
- Mission-Critical Applications
Tools & Technologies
- AWS
- Dynatrace
- Splunk
- Docker
- Kubernetes
- Terraform
- GitHub Actions
- Jenkins
- Databricks
- Snowflake