Site Reliability Engineer

AgileEngine

Brasil

Híbrido

BRL 250 000 - 450 000

Tempo integral

Há 10 dias
Gerador de candidaturas

Uma candidatura completa num minuto — currículo e carta de apresentação personalizados, prontos a enviar.

Ultrapassa os filtros ATS

Vantagens oferecidas por esta oferta de emprego

Professional growth
Competitive compensation
A selection of exciting projects
Flextime

Resumo da oferta

AgileEngine is seeking a DevOps/SRE with focus on operational resilience across Azure, AWS, and GCP. You will lead major incidents, own remediation follow-through, and develop playbooks to guide response.

Expertise in Kubernetes, Terraform, CI/CD orchestration, and Python/Go is required, with 5+ years of experience and knowledge of PCI-DSS/SOC2. You will drive security baselines using IaC, manage post-incident reviews, and communicate effectively with technical teams and executives.

Qualificações

  • 5+ years of relevant experience in DevOps or SRE.
  • Strong expertise in multi-cloud security and zero-trust design.
  • Hands-on experience with Kubernetes, Terraform, CI/CD orchestration, and scripting (Python/Go).
  • Senior-level incident-command experience in 24x7 environments.
  • Proven ability to own remediation follow-through and run post-incident reviews.

Responsabilidades

  • Scale and maintain operational stability across Azure, AWS, and GCP.
  • Engineer IaC to enforce secure baselines and guardrails.
  • Design and optimize enterprise CI/CD pipelines for secure deployments.
  • Respond to monitoring alerts using CSPM tools (e.g., Wiz).
  • Act as Incident Commander for major incidents and coordinate cross-functional teams.
  • Own post-incident remediation tracking and systemic fixes.
  • Draft clear incident notifications for technical and executive audiences.
  • Develop playbooks and runbooks for incident management.

Conhecimentos

Kubernetes
Terraform
CI/CD pipelines
Python/Go
Incident management
Multi-cloud
Zero trust
Wiz CSPM
PCI-DSS/SOC2
Communication

Ferramentas

Wiz CSPM
CI/CD tooling
Cloud security tooling

Descrição da oferta de emprego

We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. This role blends platform engineering with incident command, using Terraform, CI/CD pipelines, and CSPM tools like Wiz. You will lead major-incident calls, own remediation follow-through, and build the playbooks that guide response.

What you will do
  • Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).
  • Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.
  • Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.
  • Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.
  • Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure.
  • Own the post-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.
  • Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.
  • Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.
Must haves
  • 5+ years of experience.
  • In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles.
  • Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting.
  • Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment.
  • Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed.
  • Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.
  • Direct experience authoring divisional/group incident-management playbooks and escalation procedures.
  • Fully autonomous.
  • Drives the architecture of complex automated runbooks and mentors Middle-level SREs.
  • Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz.
  • Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2).
Nice to haves
  • PagerDuty — hands-on experience with on-call scheduling, alert routing, and incident orchestration.
  • ServiceNow — familiarity with incident, problem, and change management workflows and reporting.
Perks and Benefits
  • Professional growth

Accelerate your professional journey with mentorship, TechTalks, and personalized growth roadmaps

  • Competitive compensation

We match your ever-growing skills, talent, and contributions with competitive USD-based compensation and budgets for education, fitness, and team activities

  • A selection of exciting projects

Join projects with modern solutions development and top-tier clients that include Fortune 500 enterprises and leading product brands

  • Flextime

Tailor your schedule for an optimal work-life balance, by having the options of working from home and going to the office – whatever makes you the happiest and most productive.

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Senior DevSecOps Engineer
Senior DevSecOps Engineer

AgileEngine • Brasil

Híbrido
BRL 456 000 - 609 000
Professional growth
Competitive compensation
Flextime
+1
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine • Campinas

Híbrido
BRL 150 000 - 200 000
Professional growth
Competitive compensation
Exciting projects
+1
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine • São Bernardo do Campo

Híbrido
BRL 354 000 - 507 000
Professional growth
Competitive compensation
Exciting projects
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Recife

Híbrido
BRL 298 000 - 399 000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Riograndina

Híbrido
BRL 385 000 - 551 000
Professional growth
Competitive compensation
Flextime
+1
Site Reliability Engineer ID55632
Site Reliability Engineer ID55632

AgileEngine • São Paulo

Híbrido
Professional growth opportunities
Competitive USD-based compensation
Exciting projects with top companies
+1
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • São Bernardo do Campo

Híbrido
BRL 250 000 - 360 000
Professional growth
Competitive compensation
Exciting projects
+1
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)

iTRTech Group • São Paulo

Híbrido
BRL 180 000 - 260 000
Site Reliability Engineer ID62591
Site Reliability Engineer ID62591

AgileEngine • Salvador

Híbrido
BRL 120 000 - 180 000
Mentorship
Education budgets
Flexible schedule
+1
Senior/Lead DevSecOps Engineer (Azure, Terraform, Sentinel)
Senior/Lead DevSecOps Engineer (Azure, Terraform, Sentinel)

N-iX • Brasil

Híbrido
BRL 362 000 - 467 000
Flexible working format
Competitive salary and compensation package
Personalized career growth
+2