Service Delivery Lead

Espire Infolabs (Singapore) Pte Ltd

Região Norte

Híbrido

BRL 150 000 - 260 000

Tempo integral

14 dias+
Gerador de candidaturas

Recebe uma resposta deste empregador — um currículo e uma carta de apresentação adaptados exatamente ao que estão a contratar.

Ultrapassa os filtros ATS

Resumo da oferta

Espire Infolabs (Singapore) Pte Ltd is seeking an experienced SRE & Service Delivery Lead to oversee reliability, availability, and performance of critical systems, and to guide the production support team.

The role focuses on incident management, capacity planning, disaster recovery, and continuous improvement through automation and collaboration with development, operations, and QA teams.

Qualificações

  • Minimum 5 years as SRE or Service Delivery Lead.
  • Experience with cloud environments (AWS, Azure, Oracle Cloud).
  • Experience with Observability tools and practices.
  • Proven incident management and SRE improvements for multi-channel apps.
  • Experience in production support and disaster recovery planning.

Responsabilidades

  • Lead and mentor a team of SREs and production support engineers.
  • Oversee incident response during service interruptions and degradations.

Conhecimentos

SRE leadership
Incident management
Cloud platforms
Observability
Cross-functional collaboration

Descrição da oferta de emprego

The SRE & Service Delivery Lead responsible for overseeing the reliability, availability, and performance of our systems and applications, as well as leading the production support team to ensure the smooth operation of our production environment

  • System Reliability and Stability: The primary purpose is to guarantee that the systems and applications operated by the organization are highly reliable and stable. Enforcing service level objectives (SLOs) and service level agreements (SLAs), monitoring system health and performance
  • Incident Management and Resolution: Work with Incident Management team by leading Engineering support team for incident response efforts during service interruptions or performance degradations. Include provide timely update to Engineering management
  • Team Leadership and Development: This role involves leading and managing a team of SREs and production support engineers. The purpose here is to mentor, coach, and develop team members to foster a culture of continuous learning and improvement
  • Collaboration and Engagement: Collaboration with stakeholders is essential for prioritizing and addressing production issues effectively. The SRE and Production Support Lead should engage stakeholders in incident response efforts, problem-solving activities, and decision-making processes to ensure alignment and buy-in
  • Support Optimization: Streamlining processes, leveraging automation, fostering collaboration, and continuously improving operational efficiency
Nature of Work
  • Leading and managing a team of SREs and production support engineers is a key aspect of the role. The primary focus is to ensure the reliability and stability of the organization's systems and applications
  • The SRE and Production Support Lead is responsible for leading incident response efforts during service interruptions or performance degradations. This includes coordinating the response of the support team, diagnosing the root cause of
  • incidents, and implementing solutions to restore service as quickly as possible
  • Performance and availability management including system health checks, performance monitoring and disaster recovery planning
  • Identify and implementing Service improvements including production of improvement plans and applying software upgrades
  • Collaborate with cross-functional teams including development, operations, and quality assurance to prioritize and address production issues
  • Develop and maintain documentation, runbooks, and standard operating procedures (SOPs) to facilitate knowledge sharing and ensure consistency in support processes
  • During service interruptions or performance degradations, provide regular updates to stakeholders regarding the incident status, progress in troubleshooting, and estimated time to resolution. Problem Solving
  • Responsible for guiding the team on diagnosing and resolving complex technical issues that impact the reliability and availability of systems and application
  • Responsible to lead the team to identify the root cause of the problem & implement the fastest resolution
  • Responsible to provide timely update to stakeholders
  • Responsible to ensure documentation of problem-solving process, including the steps taken, findings, and resolutions. Share this information with relevant stakeholders, support teams, and knowledge repositories to facilitate learning and prevent similar issues in the future. Change
  • Proactive improvements to operational processes within SRE area
  • Reflect on lessons learned, successes, and areas post incident for enhancement. Refine incident response procedures, update documentation, and strengthen the team's problem-solving capabilities
Resource Complexity
  • Ensure compliance with all applicable laws and regulations relating to the above functional activities
  • Dynamic environments characterized by frequent changes, updates, and deployments introduce complexity due to the need to maintain stability and reliability amid continuous change
  • Operate in hybrid environments, leveraging a mix of on-premises infrastructure, AWS, Oracle Cloud Infrastructure, Azure, and third-party SaaS solutions
  • Managing complex systems due to interdependencies between components, where changes or failures in one area can have cascading effects on other parts of the system
  • Systems composed of diverse technologies, platforms, and services introduce complexity due to the need to understand and integrate disparate components
Experience
  • Minimum 5 years of working experience as SRE / Service Delivery Lead
  • Working experience with any programming language such as Java, Cobol, Front End technology
  • Minimum 5 years of operating on cloud environments
  • Working experience with Observability tools and practices
  • Solid experience in Incident management and SRE Improvement for applications with various target audience (back office & channels)
  • Good track record in Support Optimization
  • Preferable candidate with FI background, especially Insurance
Capability
  • Strong strategic thinking in respond to incidents and support optimization
  • Strong interpersonal and facilitation skills along with effective communication (both written and verbal) skills
  • Strong leadership and mentorship skills
  • Proactive to propose innovative solutions or alternative approaches to difficult issues
  • Ability to pick up new technology / domain knowledge quickly
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

SRE Consultant – Site Reliability Engineering
SRE Consultant – Site Reliability Engineering

Jobtailor • São Paulo

Presencial
BRL 180 000 - 320 000
SRE Pleno
SRE Pleno

Jobgether • Brasil

Presencial
BRL 120 000 - 210 000
Observability program
Multi-region exposure
OpenTelemetry stack
+2
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Riograndina

Híbrido
BRL 385 000 - 551 000
Professional growth
Competitive compensation
Flextime
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Recife

Híbrido
BRL 298 000 - 399 000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)

iTRTech Group • São Paulo

Híbrido
BRL 180 000 - 260 000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

N-iX • Brasil

Híbrido
BRL 180 000 - 360 000
Flexible work format
Competitive salary
Career growth
+3
Site Reliability Engineer ID55632
Site Reliability Engineer ID55632

AgileEngine • São Paulo

Híbrido
Professional growth opportunities
Competitive USD-based compensation
Exciting projects with top companies
+1
Site Reliability Engineer - SAP Cloud Ops
Site Reliability Engineer - SAP Cloud Ops

SAP • São Leopoldo

Presencial
BRL 180 000 - 260 000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Netskope • Brasil

Presencial
BRL 180 000 - 300 000
CX & Support Manager
CX & Support Manager

Jobtailor • São Paulo

Presencial
BRL 180 000 - 260 000