Overview
Join us as a Site Reliability Engineer (SRE) - Identity Access Management. You will be bringing to life a new digital platform capability, transforming and modernising our digital estate to build a market-leading digital offering with customer experience at its heart. You will be partnering with business aligned engineering and product teams to ensure a collaborative team culture is at the heart of what we do.
Responsibilities
- Availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.
- Resolution, analysis and response to system outages and disruptions, and implement measures to prevent similar incidents from recurring.
- Development of tools and scripts to automate operational processes, reducing manual workload, increasing efficiency, and improving system resilience.
- Monitoring and optimisation of system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
- Collaboration with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle, and work closely with other teams to ensure smooth and efficient operations.
- Stay informed of industry technology trends and innovations, and actively contribute to the organization's technology communities to foster a culture of technical excellence and growth.
Qualifications
- Experience in designing, implementing, deploying, and running highly available, fault-tolerant, auto-scaling and auto-healing systems.
- Strong expertise in AWS (essential), Azure and GCP (Google Cloud Platform) is a plus, including Kubernetes (ECS is essential, Fargate and GCE is a plus) and server-less architectures.
- Strong experience in running disaster recovery, zero downtime solutions and in designing and implementing continuous delivery across large-scale, distributed, cloud-based micro service and API service solutions with 99.9%+ uptime.
- Hands-on experience coding in Python, Bash and JSON/Yaml (Configuration as Code).
- The ability to drive reliability best practices across engineering teams, embed SRE principles into the DevSecOps lifecycle and partner with engineering, security and product teams, to balance reliability and feature velocity.
Desirable Skills
- Experience in hands-on configuration, deployment and operation of ForgeRock COTS based IAM (Identity Access management) solutions (PingGateway, PingAM, PingIDM, PingDS) with embedded security gates, HTTP header signing, access token and data at rest encryption, PKI based self-sovereign identity, or open source.
This role will be based out of our Pune office.
Purpose of the role
To apply software engineering techniques, automation, and best practices in incident response, to ensure the reliability, availability, and scalability of the systems, platforms, and technology through them.
Accountabilities
- Availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.
- Resolution, analysis and response to system outages and disruptions, and implement measures to prevent similar incidents from recurring.
- Development of tools and scripts to automate operational processes, reducing manual workload, increasing efficiency, and improving system resilience.
- Monitoring and optimisation of system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
- Collaboration with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle, and work closely with other teams to ensure smooth and efficient operations.
- Stay informed of industry technology trends and innovations, and actively contribute to the organization's technology communities to foster a culture of technical excellence and growth.
Leadership/Executive Expectations
- To contribute or set strategy, drive requirements and make recommendations for change. Plan resources, budgets, and policies; manage and maintain policies/processes; deliver continuous improvements and escalate breaches of policies/procedures.
- If managing a team, define jobs and responsibilities, plan for the department’s future needs and operations, counsel employees on performance and contribute to pay decisions. They may also lead specialists to influence operations, while balancing budgets and schedules.
- Leaders demonstrate a set of leadership behaviours to create an environment for colleagues to thrive and deliver to a consistently excellent standard.
- For individual contributors, act as a subject matter expert within your discipline and guide technical direction on multi-year assignments.
- Advise key stakeholders, including functional leadership teams and senior management on functional and cross-functional areas of impact and alignment.
- Manage and mitigate risks through assessment, in support of the control and governance agenda.
- Demonstrate leadership and accountability for managing risk and strengthening controls in relation to the work your team does.
- Demonstrate understanding of the organisation's functions to contribute to achieving the goals of the business.
- Collaborate with other areas of work to keep up to speed with business activity and strategies.
- Create solutions based on sophisticated analytical thought and problem solving to define problems and develop innovative solutions.
- Adopt and include outcomes of extensive research in problem solving processes.
- Seek out, build and maintain trusting relationships with internal and external stakeholders to accomplish key business objectives.
All colleagues will be expected to demonstrate the Barclays Values of Respect, Integrity, Service, Excellence and Stewardship – our moral compass, helping us do what we believe is right. They will also be expected to demonstrate the Barclays Mindset – to Empower, Challenge and Drive.