Platform Reliability & Operations Lead, Sumaré

3M

Sumaré

Presencial

BRL 180 000 - 240 000

Tempo integral

14 dias+

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

3M is seeking a Platform Reliability & Operations Lead in Sumaré to ensure the reliability, stability, and operational excellence of our enterprise data and AI platform, including Databricks and AWS. You will own runtime health, lead incident response, and drive continuous improvement across engineering and governance teams.

You will combine platform engineering with operations, translating platform capabilities into reliable production systems and maturing observability, runbooks, and

Qualificações

  • Bachelor’s degree or higher in Computer Science, Engineering, or related technical field.
  • Experience in IT operations, SRE, or platform operations with high availability and reliability requirements.
  • Strong knowledge of incident management and root-cause analysis.

Responsabilidades

  • Ensure the stability, availability, and operational health of the platform, including Databricks, workflows, integrated systems, and AWS infrastructure.
  • Define and monitor health, performance, and reliability metrics such as platform availability, job success rates, workflow completion, and incident resolution times.
  • Lead the management of critical incidents by coordinating technical teams, running war rooms, and driving root cause analysis with a focus on continuous improvement.
  • Implement and enhance SRE practices, observability, monitoring, alerting, and automation to improve platform reliability.
  • Design and operate the platform support model, including intake, triage, escalation, SLAs, and operational KPIs.
  • Monitor and support production pipelines and workflows, acting to prevent and resolve failures, delays, and service disruptions.
  • Ensure adherence to operational and governance standards in production, including access controls, naming conventions, and approved platform policies.
  • Partner with Platform Engineering, Data Engineering, and Governance teams to ensure operational readiness, effective support, and sustainable platform evolution.
  • Maintain up-to-date operational documentation, including runbooks, incident playbooks, and troubleshooting guides.

Conhecimentos

Infrastructure operations
Incident management
Cloud operations
Cross-functional coordination
English proficiency

Formação académica

Bachelor’s degree in Computer Science or Engineering

Ferramentas

Databricks
AWS
Datadog
Grafana
CloudWatch
Dynatrace
Temporal/Workflow orchestrators

Descrição da oferta de emprego

Job Description

3M has a long-standing reputation as a company committed to innovation. We provide the freedom to explore and encourage curiosity and creativity. We gain new insight from diverse thinking, and take risks on new ideas. Here, you can apply your talent in bold ways that matter.

Platform Reliability & Operations Lead, Sumaré
Collaborate with Innovative 3Mers Around the World

Choosing where to start and grow your career has a major impact on your professional and personal life, so it’s equally important you know that the company that you choose to work at, and its leaders, will support and guide you. With a diversity of people, global locations, technologies and products, 3M is a place where you can collaborate with 93,000 other curious, creative 3Mers.

The Impact You’ll Make in this Role

3M is seeking a Platform Ops & Reliability Lead to join the Corporate Research Digital Platforms (CRDP) team to ensure the reliability, stability, and operational excellence of our enterprise data and AI platform, including Databricks and supporting cloud infrastructure (AWS). In this role, you will own the runtime health of the platform, serving as the primary leader for incident response, operational processes, and system reliability. You will work across Databricks, Temporal workflows, AWS infrastructure, and metadata systems to ensure that platform services operate predictably, scale effectively, and meet reliability expectations.

You will operate at the intersection of engineering and operations, partnering closely with Platform Engineers, Data Engineers, and Governance teams to translate platform capabilities into reliable production systems.

This role is critical in enabling scalable adoption of the platform by ensuring systems are stable, supportable, and operationally mature.

In this role, you will have the opportunity to
  • Ensure the stability, availability, and operational health of the platform, including Databricks, workflows, integrated systems, and AWS infrastructure.
  • Define and monitor health, performance, and reliability metrics such as platform availability, job success rates, workflow completion, and incident resolution times.
  • Lead the management of critical incidents by coordinating technical teams, running war rooms, and driving root cause analysis with a focus on continuous improvement.
  • Implement and enhance SRE practices, observability, monitoring, alerting, and automation to improve platform reliability.
  • Design and operate the platform support model, including intake, triage, escalation, SLAs, and operational KPIs.
  • Monitor and support production pipelines and workflows, acting to prevent and resolve failures, delays, and service disruptions.
  • Ensure adherence to operational and governance standards in production, including access controls, naming conventions, and approved platform policies.
  • Partner with Platform Engineering, Data Engineering, and Governance teams to ensure operational readiness, effective support, and sustainable platform evolution.
  • Maintain up-to-date operational documentation, including runbooks, incident playbooks, and troubleshooting guides.
Your Skills And Expertise
  • Strong experience in infrastructure operations, cloud operations, SRE, or platform operations roles.
  • Experience supporting Databricks or modern data platforms (Lakehouse architectures).
  • Proven experience managing production systems with high availability, reliability, and operational rigor.
  • Experience leading incident management, including major incident response and root cause analysis.
  • Hands‑on experience with monitoring and observability tools (e.g., Datadog, Grafana, CloudWatch, Dynatrace).
  • Experience supporting cloud‑based platforms (AWS preferred), including troubleshooting infrastructure and services.
  • Strong understanding of IT operations, support models (L1/L2/L3), and service management practices (ITIL).
  • Experience working in cross‑functional environments, coordinating between engineering, platform, and support teams.
  • Bachelor’s degree or higher in Computer Science, Engineering, or related technical field.
  • Proficiency in English.
Additional qualifications that could help you succeed even further in this role
  • Familiarity with workflow orchestration platforms such as Temporal, Airflow, or Step Functions.
  • Experience implementing SRE practices, including SLIs, SLOs, and reliability metrics.
  • Experience with cloud cost management and FinOps practices.
  • Familiarity with IAM, access control models, and security best practices in cloud environments.
  • Basic scripting or automation experience (Python, Bash, or similar).
  • Strong communication and coordination skills, with the ability to lead cross‑team operational efforts.
Work location

This role follows an on‑site working model, requiring the employee to work at least four days a week at 3M in Sumaré/SP.

Supporting Your Well‑being

3M offers many programs to help you live your best life – both physically and financially. To ensure competitive pay and benefits, 3M regularly benchmarks with other companies that are comparable in size and scope.

A 3M é um empregador que oferece oportunidades iguais à todos. A 3M não discriminará nenhum candidato baseado em sua raça, cor, idade, religião, gênero, orientação sexual, identidade ou expressão de gênero, nacionalidade ou deficiência.

Safety is a core value at 3M. All employees are expected to contribute to a strong Environmental Health and Safety (EHS) culture by following safety policies, identifying hazards, and engaging in continuous improvement.

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Backend Software Engineering Lead, Sumaré
Backend Software Engineering Lead, Sumaré

3M • Sumaré

Presencial
BRL 180 000 - 300 000
Sr Data Engineering Specialist, Sumaré
Sr Data Engineering Specialist, Sumaré

3M • São Paulo

Presencial
BRL 180 000 - 320 000
Backend Software Engineering Lead, Sumaré
Backend Software Engineering Lead, Sumaré

3M • São Paulo

Presencial
BRL 150 000 - 230 000
Sr Cloud Engineering Specialist
Sr Cloud Engineering Specialist

3m • Sumaré

Presencial
Benefícios competitivos
Oportunidades de desenvolvimento profissional
Analytics & Insights Business Partner Specialist, Sumaré
Analytics & Insights Business Partner Specialist, Sumaré

3M • Sumaré

Presencial
BRL 180 000 - 280 000
Full-Stack Developer - Sumare, Brasil.
Full-Stack Developer - Sumare, Brasil.

3M • Sumaré

Presencial
BRL 170 000 - 230 000
Sr Data Engineering Specialist, Sumaré
Sr Data Engineering Specialist, Sumaré

3M • Sumaré

Presencial
BRL 180 000 - 270 000
Engenheiro de Projetos e confiabilidade, Manaus
Engenheiro de Projetos e confiabilidade, Manaus

3m • Manaus

Presencial
Programas de benefícios
Marketing Performance Analytics Sr Specialist / Sumaré SP
Marketing Performance Analytics Sr Specialist / Sumaré SP

3M • Sumaré

Presencial
BRL 180 000 - 300 000
Analytics Delivery Sr. Specialist
Analytics Delivery Sr. Specialist

3M • Sumaré

Presencial
BRL 180 000 - 240 000