Specialist II – Site Reliability Engineering, Command Center

Jobtailor

São Paulo

Presencial

BRL 180 000 - 240 000

Tempo integral

14 dias+

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

Jobtailor in São Paulo, Brazil, seeks an experienced Site Reliability Engineer to lead incident response and reliability initiatives. You will collaborate across development, architecture, platform, and operations teams to enhance availability and drive continuous improvements.

You will own SLIs/SLOs, promote automation with Python/Shell, and apply AI/AIOps to prevent incidents and reduce toil, while guiding war rooms and technical meetings in English.

Qualificações

  • Experience with SRE practices across software lifecycle.
  • Proven ability to lead incident response and war rooms.
  • Strong knowledge of infrastructure and AWS Cloud.
  • Fluent English for leading technical meetings.
  • Knowledge of observability, monitoring and incident management.
  • Automation experience with Python or Shell Script.
  • Understanding of distributed architecture and microservices.
  • Experience with observability and monitoring tools.
  • Familiarity with container platforms and orchestration.
  • Knowledge of AI applied to operations and automation.

Responsabilidades

  • Lead troubleshooting across the software development lifecycle with cross-team collaboration.
  • Manage critical incident response to reduce MTTR and recurrence.
  • Serve as reliability expert for a Business Unit and its services.
  • Map risks and opportunities through SRE principles.
  • Lead initiatives to increase availability, stability, and resilience.
  • Define and evolve SLIs, SLOs, and error budgets.
  • Lead root cause analyses and remediation plans.
  • Promote automation to reduce toil and improve efficiency.
  • Explore AI- and AIOps-based incident prevention and resolution.
  • Develop proactive operational perspective to anticipate risks.

Conhecimentos

SRE practices
Lifecycle knowledge
AWS Cloud
English fluency
Observability
Automation (Python/Shell)
Distributed architecture
Observability tools
Containers & orchestration
AI for IT ops

Ferramentas

Observability tools
Container platforms

Descrição da oferta de emprego

Key Responsibilities
  • Lead troubleshooting across the software development lifecycle, collaborating with development, architecture, platform, and operations teams to drive continuous improvements.
  • Manage critical incident response, contributing to reduced impact, lower MTTR and decreased recurrence.
  • Act as the reliability expert for a specific Business Unit, developing deep knowledge of its journeys and critical services.
  • Map risks, weaknesses and opportunities for improvement through the lens of Site Reliability Engineering (SRE) principles.
  • Identify and lead initiatives to increase availability, stability, performance, and resilience of systems.
  • Define, track, and evolve reliability metrics such as SLIs, SLOs, and Error Budgets.
  • Lead root cause analyses and structured remediation plans for recurring issues.
  • Promote adoption of automation to reduce manual operational work (toil) and improve operational efficiency.
  • Explore and implement AI- and AIOps-based solutions for incident prevention, detection, and resolution.
  • Develop a proactive operational perspective, anticipating risks before they impact customers and the business.
Requirements
  • Experience with Site Reliability Engineering (SRE) practices and experience operating in Command Centers, NOCs, or Operations Centers.
  • Proven, in-depth experience across stages of the software development lifecycle.
  • Strong knowledge of infrastructure and AWS Cloud.
  • Fluent English for leading war rooms and technical meetings.
  • Knowledge of observability, monitoring, and incident management.
  • Experience with automation using languages such as Python, Shell Script, or similar.
  • Knowledge of distributed architecture, microservices, and critical systems.
  • Experience with observability and monitoring tools.
  • Familiarity with container platforms and orchestration.
  • Knowledge of Artificial Intelligence applied to operations, observability, and automation.
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Mid-level SRE
Mid-level SRE

Jobtailor • São Paulo

Presencial
BRL 180 000 - 240 000
Senior DevOps Analyst – SRE, Kubernetes, Cloud
Senior DevOps Analyst – SRE, Kubernetes, Cloud

Jobtailor • São Paulo

Presencial
BRL 90 000 - 130 000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • São Paulo

Presencial
BRL 180 000 - 260 000
Senior SRE
Senior SRE

Jobtailor • São Paulo

Presencial
BRL 180 000 - 260 000
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Riograndina

Híbrido
BRL 385 000 - 551 000
Professional growth
Competitive compensation
Flextime
+1
Senior Reliability and Platform Engineer
Senior Reliability and Platform Engineer

Jobtailor • São Paulo

Presencial
BRL 180 000 - 300 000
Senior SRE, Specialist
Senior SRE, Specialist

Jobtailor • São Paulo

Presencial
BRL 180 000 - 320 000
Site Reliability Engineer ID55632
Site Reliability Engineer ID55632

AgileEngine • São Paulo

Híbrido
Professional growth opportunities
Competitive USD-based compensation
Exciting projects with top companies
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Recife

Híbrido
BRL 298 000 - 399 000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems • Brasil

Presencial
BRL 250 000 - 450 000