Incident & Problem Management: Own the RCA process for production incidents — diagnose, resolve, and put preventive measures in place so issues don't recur
Production Monitoring & Support: Continuously monitor service health, detect anomalies early, and act before they become incidents
Deployment Execution: Design, implement, and maintain CI/CD pipelines using GitHub Actions and related tooling to automate build, test, security scanning, and deployment processes.
Environment Oversight: Keep Pre-Production and Production environments stable and aligned — not building them from scratch, but ensuring they behave as expected day to day
Runbook & Knowledge Management: Document operational procedures, known issues, and resolution steps to build a reliable knowledge base for the team
Cross-team Collaboration: Work shoulder-to-shoulder with development and platform teams to triage issues, clarify operational requirements, and close the feedback loop between prod and dev
Operational improvement:
- Identify recurring pain points and propose automation or tooling to reduce toil
- Improve observability coverage — dashboards, alerts, log queries — to catch issues faster
- Contribute to service continuity initiatives and disaster recovery drills
Must-have knowledge and experience
- 5+ years in IT operations, application support (2nd/3rd line), or a similar production-facing role
- Proven track record of owning incidents end-to-end — from alert to RCA to prevention
- 2+ years working within an ITIL framework (incident, problem, change management)
- Experience working in Agile delivery environments alongside development teams
- Excellent English communication skills — able to explain technical issues clearly to both engineers and non-technical stakeholders, C1
- Must-Have Technical Skills:
- Production operations & troubleshooting:
- Proficiency with Jenkins to build and maintain pipelines to execute and troubleshoot deployments
- Proficiency with CI/CD pipeline – introducing improvements and keeping the pipeline automated
- Excellent troubleshooting and problem-solving skills with the ability to independently investigate complex production issues.
- Proficiency with log analysis and alerting tools: Splunk, Sysdig
- Fluency in observability tooling: Prometheus, Grafana — reading dashboards, tuning alerts
- Comfortable operating services running on Kubernetes (checking pod health, reading logs, triggering restarts — not cluster administration)
- Excellent skills with Ansible for applying configuration changes in controlled operational scenarios
- Strong knowledge of Docker and Docker Compose
- Basic scripting skills (Bash, Python) for automation of repetitive operational tasks and reconciliation of data
Nice-to-have knowledge and experience
- Nice-to-have knowledge and experience:
- IBM Datastage operational experience
- Awareness or willing to learn Pega, Airflow
Application & data layer:
- Relational databases (Oracle, DB2) — querying, interpreting execution plans, identifying data-related incidents
- Working knowledge of ETL application behavior, Rest API communication
- Experience supporting distributed systems and platform services such as Kafka message flow
- Java/ development background for understanding the solution and integrations