System Reliability Engineer (Application Support + Automation)
Who are we Fulcrum Digital is an agile and next-generation digital accelerating companyproviding digital transformation and technology services right from ideation toimplementation. These services have applicability across a variety ofindustries, including banking & financial services, insurance, retail,higher education, food, healthcare, and manufacturing.
The Role
- Plan, manage,and oversee all aspects of a Production Environment
- Definestrategies for Application Performance Monitoring, Optimization in Prodenvironment
- Respond toIncidents and improvise platform based on feedback and measure thereduction of incidents over time.
- Supportdeployment of code into multiple lower environments. Supportingcurrent processes with an emphasis on automating everything as soon aspossible.
- Design, develop and standardize Monitoring andAlerting mechanism for the supported applications.
- Take aholistic approach to problem solving, by connecting the dots during aproduction event through the various technology stack that makes up theplatform, to optimize meantime to recover.
- Engage in andimprove the whole lifecycle of services—from inception and design, throughdeployment, operation and refinement.
- Analyse ITSMactivities of the platform and provide feedback loop to development teamson operational gaps or resiliency concerns.
- Supportservices before they go live through activities such as system designconsulting, capacity planning and launch reviews.
- Support theapplication CI/CD pipeline for promoting software into higher environmentsthrough validation and operational gating, and lead in DevOps automationand best practices.
- Maintainservices once they are live by measuring and monitoring availability,latency, and overall system health.
- Scale systemssustainably through mechanisms like automation and evolving systems bypushing for changes that improve reliability and velocity.
- Work with aglobal team spread across tech hubs in multiple geographies and timezones.
- Ability toshare knowledge and explain processes and procedures to others.
- Able toperform on-call duties on a rotational basis.
- Occasionaloff hours work required.
Requirements
- ShellScripting
- ITIL / ITSM
- SQL - good to have
- ApplicationTroubleshooting
- AnyMonitoring tool (Preferred Splunk/Dynatrace)
- Jenkins -CI/CD
- Ansible
- GroovyScripting/Yaml
- Git basic/bitbucket
Good To Have
- Payments Flows, Switching, Settlements, Authorisation flows.