Job Overview
To apply software engineering techniques, automation, and best practices in incident response, to ensure the reliability, availability, and scalability of the systems, platforms, and technology through them.
Accountabilities
- Availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.
- Resolution, analysis and response to system outages and disruptions, and implementation of measures to prevent similar incidents from recurring.
- Development of tools and scripts to automate operational processes, reducing manual workload, increasing efficiency, and improving system resilience.
- Monitoring and optimization of system performance and resource usage, identification and resolution of bottlenecks, and implementation of best practices for performance tuning.
- Collaboration with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle, and work closely with other teams to ensure smooth and efficient operations.
- Staying informed of industry technology trends and innovations, and actively contributing to the organization's technology communities to foster a culture of technical excellence and growth.
Responsibilities
- Plan resources, budgets, and policies; manage and maintain policies/processes; deliver continuous improvements and escalated breaches of policies/procedures.
- If managing a team, define jobs and responsibilities, plan for the department's future needs and operations, counsel employees on performance and pay decisions, and lead specialists to influence the operations of a department in alignment with strategic and tactical priorities.
- As an individual contributor, act as a subject‑matter expert within own discipline, guide technical direction, lead collaborative, multi‑year assignments, train and coach less experienced specialists, and provide information affecting long‑term profits, organisational risks, and strategic decisions.
- Advise key stakeholders, including functional leadership teams and senior management, on functional and cross‑functional areas of impact and alignment.
- Manage and mitigate risks through assessment, in support of the control and governance agenda, and demonstrate leadership and accountability for managing risk and strengthening controls.
- Create solutions based on sophisticated analytical thought, comparing and selecting complex alternatives, with in‑depth analysis and interpretative thinking to define problems and develop innovative solutions.
- Seek out, build, and maintain trusting relationships and partnerships with internal and external stakeholders to accomplish key business objectives, using influencing and negotiating skills to achieve outcomes.
Qualifications
- Experience in designing, implementing, deploying, and running highly available, fault‑tolerant, auto‑scaling and auto‑healing systems.
- Strong expertise in AWS (essential), with Azure and GCP (Google Cloud Platform) as a plus, including Kubernetes (ECS essential, Fargate and GCE a plus) and server‑less architectures.
- Strong experience in running disaster recovery, zero downtime solutions, and in designing and implementing continuous delivery across large‑scale, distributed, cloud‑based micro‑service and API service solutions with 99.9%+ uptime.
- Hands‑on experience coding in Python, Bash and JSON/Yaml (Configuration as Code).
- The ability to drive reliability best practices across engineering teams, embed SRE principles into the DevSecOps lifecycle, and partner with engineering, security and product teams to balance reliability and feature velocity.
- Experience in hands‑on configuration, deployment and operation of ForgeRock COTS‑based IAM solutions (PingGateway, PingAM, PingIDM, PingDS) with embedded security gates, HTTP header signing, access token and data at rest encryption, PKI‑based self‑sovereign identity, or open source.
Location
This role will be based out of our Pune office.