DevOps Coordinator – SRE

Jobtailor

São Paulo

Presencial

BRL 180 000 - 320 000

Tempo integral

14 dias+

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

Jobtailor in São Paulo seeks an experienced DevOps/SRE leader to define and run SLOs, manage incidents, and scale production environments across GKE, AWS, and Azure.

You will build pipelines with ArgoCD and Terraform, implement OpenTelemetry, and mentor engineers while ensuring security and reliability in high-availability contexts.

Qualificações

  • 6+ years in DevOps/SRE/Infrastructure with at least 2 years in leadership roles.
  • Hands-on with GCP/GKE, AWS, Azure, and multi-cloud environments.
  • Strong Kubernetes knowledge: RBAC, Network Policies, Admission Controllers.
  • CI/CD, GitOps and IaC: ArgoCD, Cloud Build, Terraform (modules, state).
  • Observability and SRE: Datadog, Grafana; SLI/SLO/SLA concepts.
  • DevSecOps: SAST tools (SonarQube), image scanning, incident management.
  • Scripting: Python or Shell for automation.
  • OpenTelemetry instrumentation; API gateways (Apigee/Kong).
  • English: Intermediate to advanced reading/writing for vendor interactions.

Responsabilidades

  • Define and monitor SLOs with product squads using error budgets to guide priorities.
  • Lead war rooms for P1/P2 incidents with blameless post-mortems.
  • Maintain scalable, secure production GKE clusters across AWS and Azure.
  • Design and evolve pipelines with Cloud Build + ArgoCD and quality gates.
  • Structure Terraform modules for multi-project GCP environments and remote state.
  • Ensure production readiness with logging, traces (OpenTelemetry) and alerts (Datadog/Grafana).
  • Mentor team members, conduct code reviews, and organize a sustainable on-call rota.

Conhecimentos

Python Scripting
Shell Scripting
Terraform Modules
Helm Charts Creation
OpenTelemetry Instrumentation
API Gateway Management
SAST Tools (SonarQube)
Cloud Build
ArgoCD
Multicloud Services (AWS, Azure)
Mentorship
Team Collaboration
Crisis Management
English Proficiency
Kubernetes
GKE
Cloud Platforms (GCP)

Formação académica

Bachelor’s degree in Computer Science, Software Engineering, Information Systems or related fields

Ferramentas

Kubernetes
GKE
Terraform
Helm
OpenTelemetry
Datadog
Grafana
ArgoCD
Cloud Build
APIs Gateways (Apigee/Kong)
AWS
Azure
GCP

Descrição da oferta de emprego

  • Define and monitor SLOs with product squads, actively using error budgets to guide prioritization decisions and control release velocity.
  • Lead war rooms for critical incidents (P1/P2) end-to-end (triage, diagnosis, resolution) and conduct blameless post‑mortems (5 Whys).
  • Operate and ensure the scalability, health, and security of production GKE (Google Kubernetes Engine) clusters, while maintaining visibility over workloads in AWS and Azure.
  • Design and evolve Cloud Build + ArgoCD pipelines with mandatory quality gates (SonarQube, image scanning, smoke tests) and define rollout strategies (canary, blue/green).
  • Structure and maintain Terraform modules for multi‑project GCP environments, managing remote state, drift detection, and policy‑as‑code.
  • Ensure production readiness with structured logs, traces via OpenTelemetry, alerts in Datadog/Grafana, and integrate security tools (SAST, Trivy, Secret Manager) without introducing friction into the delivery flow.
  • Provide structured mentorship to the team (internal staff and consultants), use code reviews as a teaching tool, and organize a sustainable on‑call rota.
Requirements
  • GCP & GKE: Hands‑on production cluster operation (troubleshooting, HPA, PDB, networking, node pools, Workload Identity).
  • Advanced Kubernetes: Deep understanding of workload lifecycle, RBAC, Network Policies, and Admission Controllers.
  • CI/CD, GitOps and IaC tools: ArgoCD, Cloud Build (or GitHub Actions/GitLab CI), and Terraform (modules with semantic versioning and state management).
  • Observability and SRE: Datadog and Grafana (dashboards, monitors, SLO tracking) and conceptual and practical mastery of SRE concepts (SLI/SLO/SLA).
  • DevSecOps & Incident Management: Experience with SAST (SonarQube), image scanning and crisis/incident handling.
  • Scripting: Python and/or Shell scripting for automation.
  • Multicloud knowledge of equivalent services in AWS (EKS, S3, IAM) and Azure (AKS, Azure AD).
  • Creation and maintenance of Helm charts.
  • Instrumenting services with OpenTelemetry.
  • Experience with API Gateways (Apigee X/Edge and/or Kong).
  • Active certifications: CKA, Google Cloud Professional Cloud Architect, AWS Solutions Architect or AZ‑104.
  • Education: Bachelor's degree in Computer Science, Software Engineering, Information Systems or related fields (postgraduate degree is a plus).
  • Experience: Minimum of 6 years in DevOps/SRE/Infrastructure, with at least 2 years in coordination or technical leadership roles.
  • Industry experience: Prior work in high‑availability environments, dealing with financial systems or high‑volume e‑commerce (critical transactions).
  • Language: Intermediate to advanced English for reading technical documentation and interacting with vendors.
ATS Optimization Keywords

Below are skills and terms extracted directly from this job posting to improve Applicant Tracking System (ATS) visibility. This unique feature helps candidates tailor their applications more effectively — a feature exclusive to JobTailor job listings.

Hard Skills
  • Python Scripting
  • Shell Scripting
  • Terraform Modules
  • Helm Charts Creation
  • OpenTelemetry Instrumentation
  • API Gateway Management
  • SAST Tools (SonarQube)
  • Cloud Build
  • ArgoCD
  • Multicloud Services (AWS, Azure)
Soft Skills
  • Mentorship
  • Team Collaboration
  • Crisis Management
Certifications & Qualifications
  • CKA
  • Google Cloud Professional Cloud Architect
  • AWS Solutions Architect
  • AZ‑104
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

DevOps Engineer I
DevOps Engineer I

Jobtailor • Blumenau

Presencial
BRL 90 000 - 150 000
Senior Infrastructure and Cloud Analyst
Senior Infrastructure and Cloud Analyst

Jobtailor • São Paulo

Presencial
BRL 180 000 - 240 000
Mid-level SRE
Mid-level SRE

Jobtailor • São Paulo

Presencial
BRL 180 000 - 240 000
Platform Engineering Specialist, Cloud
Platform Engineering Specialist, Cloud

Jobtailor • São Paulo

Presencial
BRL 180 000 - 300 000
Senior DevOps Analyst – SRE, Kubernetes, Cloud
Senior DevOps Analyst – SRE, Kubernetes, Cloud

Jobtailor • São Paulo

Presencial
BRL 90 000 - 130 000
Senior SRE, GCP
Senior SRE, GCP

Jobtailor • São Paulo

Presencial
BRL 180 000 - 260 000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • São Paulo

Presencial
BRL 180 000 - 260 000
Senior Reliability and Platform Engineer
Senior Reliability and Platform Engineer

Jobtailor • São Paulo

Presencial
BRL 180 000 - 300 000
Senior Software Engineer – C++, Python
Senior Software Engineer – C++, Python

Jobtailor • São Paulo

Presencial
BRL 150 000 - 210 000
Middleware – Cloud Operations
Middleware – Cloud Operations

Jobtailor • São Paulo

Presencial
BRL 120 000 - 180 000