SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)

iTRTech Group

São Paulo

Híbrido

BRL 180 000 - 260 000

Tempo integral

14 dias+
Gerador de candidaturas

Destaca-te para esta função — gera um currículo e uma carta de apresentação personalizados em cerca de um minuto.

Ultrapassa os filtros ATS

Resumo da oferta

iTRTech Group in Brazil seeks a Senior Site Reliability Engineer to improve reliability, scalability, and operability of enterprise cloud environments. You will automate operations, enhance observability, and ensure high availability across mission-critical systems.

The role requires advanced English, 6+ years of experience, and strong SRE/DevOps skills with AWS/Azure/GCP, Terraform, Docker, Kubernetes, and CI/CD expertise.

Qualificações

  • Advanced/Fluent English.
  • Proven experience as an SRE.
  • Strong background in Site Reliability Engineering, DevOps and cloud infrastructure.
  • Hands-on with Terraform, Docker, Kubernetes, CI/CD pipelines, Python or Bash.

Responsabilidades

  • Automate operational processes, deployments, and infrastructure provisioning.
  • Develop internal automation tools using Python and Bash.
  • Design, implement, and optimize CI/CD pipelines.
  • Provision and manage cloud infrastructure using IaC.
  • Monitor applications and infrastructure; define SLIs/SLOs/SLAs.
  • Respond to incidents and conduct RCA for continuous improvement.
  • Collaborate with software engineering and platform teams.

Conhecimentos

Cloud infrastructure
DevOps
Observability
Incident management
Automation
SRE practices

Formação académica

Computer Science
Engineering
Related field

Ferramentas

Terraform
Docker
Kubernetes
CI/CD pipelines
Python
Bash
Linux/Unix

Descrição da oferta de emprego

SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE BRAZIL)

Brazilian company hires for hybrid or remote position

Location: Brazil (any location)

Only candidates already based in Brazil will be considered

Work Model: Hybrid for candidates living in state capitals and Remote for candidates living in countryside/cities outside the state capitals

Language Requirements: Advanced/Fluent English – Mandatory

Seniority: Senior (6+ years)

Compensation: Please inform your salary expectations when applying.

Build Highly Reliable Cloud Platforms at Enterprise Scale

We are looking for an experienced Site Reliability Engineer (SRE) to improve the reliability, scalability, performance, and operational excellence of enterprise cloud environments.

You will work closely with development, platform, and infrastructure teams to automate operations, improve observability, optimize deployments, and ensure high availability across mission-critical systems.

If you enjoy solving complex infrastructure challenges through engineering, automation, and modern cloud technologies, this opportunity is for you.

The Professional We Are Looking For

We are seeking a highly skilled Site Reliability Engineer with extensive experience in cloud infrastructure, DevOps practices, automation, and distributed systems.

The ideal candidate combines strong infrastructure knowledge with software engineering principles, helping organizations improve operational efficiency through automation, Infrastructure as Code (IaC), observability, and continuous improvement.

You should be comfortable working in high-availability environments, responding to critical incidents, and continuously enhancing platform resilience.

Key Responsibilities
Automation & Operational Excellence
  • Automate operational processes, deployments, and infrastructure provisioning.
  • Develop internal automation tools using Python and Bash.
  • Design, implement, and optimize CI/CD pipelines.
  • Reduce repetitive operational activities through automation.
Cloud Infrastructure & Infrastructure as Code
  • Provision and manage cloud infrastructure.
  • Implement Infrastructure as Code using Terraform (preferred) and other IaC tools.
  • Ensure scalability, resilience, and high availability.
  • Design auto-scaling and self-healing solutions.
Reliability & Performance
  • Monitor applications and infrastructure using metrics, logs, and distributed tracing.
  • Define and monitor SLIs, SLOs, and SLAs.
  • Perform capacity planning.
  • Troubleshoot performance bottlenecks and system issues.
Incident Management
  • Respond to critical production incidents.
  • Conduct Root Cause Analysis (RCA).
  • Drive continuous improvement initiatives.
Observability
  • Implement monitoring, alerting, and observability platforms.
  • Build operational dashboards.
  • Improve end-to-end system visibility.
Platform Engineering
  • Collaborate closely with software engineering teams.
  • Support DevOps and SRE best practices.
  • Improve deployment processes, system architecture, and application resilience.
Business Continuity
  • Support Disaster Recovery strategies.
  • Execute load and stress testing.
  • Implement Chaos Engineering practices.
  • Improve operational resilience.
Governance
  • Define operational standards.
  • Document infrastructure and architecture.
  • Promote security best practices and DevSecOps principles.
Mandatory Requirements (Eliminatory)

All requirements below are mandatory and eliminatory. Candidates who cannot clearly demonstrate these qualifications in their CV are unlikely to proceed in the recruitment process.

Education
  • Computer Science
  • Engineering
  • or a related field
Mandatory Experience
  • Advanced/Fluent English.
  • Proven experience as a Site Reliability Engineer (SRE).
  • Strong background in:
  • Site Reliability Engineering
  • DevOps
  • Cloud Infrastructure
  • Cloud Engineering
  • Hands-on experience with at least one major cloud platform:
  • AWS
  • Microsoft Azure
  • Google Cloud Platform (GCP)
  • Practical experience with:
  • Terraform (preferred)
  • Docker
  • Kubernetes
  • CI/CD Pipelines
  • Python or Bash
  • Linux/Unix
  • Strong experience with monitoring and observability platforms.
  • Incident Management and Production Support.
  • Strong troubleshooting and problem-solving skills.
  • Experience building highly available and scalable systems.
Nice-to-Have Skills
  • Large-scale enterprise environments.
  • Consulting experience.
  • Chaos Engineering.
  • Highly distributed architectures.
  • Advanced observability platforms.
  • Apache Kafka.
  • Financial Services or mission-critical environments.
  • DevSecOps.
  • Platform Engineering.
What You'll Find in This Opportunity
  • Enterprise-scale cloud infrastructure projects.
  • Modern Site Reliability Engineering culture.
  • Cloud-native and Platform Engineering initiatives.
  • Strong DevOps and Infrastructure as Code practices.
  • Collaboration with international engineering teams.
  • High-impact role supporting mission-critical platforms.
  • Modern engineering environment focused on automation and operational excellence.
Before Applying, Ask Yourself These 5 Questions
  • Have I worked as a Site Reliability Engineer (SRE) or in an equivalent role focused on cloud reliability, automation, and operational excellence?
  • Does my CV clearly demonstrate hands-on experience with AWS, Azure, or GCP, as well as Terraform, Docker, Kubernetes, Linux, Python/Bash, and CI/CD pipelines?
  • Have I implemented observability solutions, managed production incidents, conducted Root Cause Analysis (RCA), and worked with SLIs, SLOs, and SLAs?
  • Do I have experience designing highly available, scalable, and resilient cloud environments using Infrastructure as Code and DevOps best practices?
  • Am I fluent in English and comfortable collaborating with international engineering teams in mission-critical production environments?
Important

This position is intended for a Senior Site Reliability Engineer (SRE) with extensive experience in cloud infrastructure, automation, observability, DevOps, and platform reliability.

Candidates whose experience is primarily focused on traditional infrastructure administration, system support, or operations without demonstrated expertise in Infrastructure as Code, Kubernetes, cloud platforms, automation, CI/CD, and Site Reliability Engineering practices are unlikely to meet the expectations for this role.

Keywords That Should Appear in Your CV

Site Reliability Engineer, SRE, DevOps Engineer, Platform Engineer, Cloud Engineer, Infrastructure Engineer, Cloud Infrastructure, AWS, Amazon Web Services, Microsoft Azure, Google Cloud Platform, GCP, Terraform, Infrastructure as Code, IaC, Docker, Kubernetes, Linux, Unix, Python, Bash, Shell Scripting, CI/CD, Jenkins, GitLab CI, GitHub Actions, Monitoring, Observability, Prometheus, Grafana, ELK Stack, OpenTelemetry, Distributed Tracing, Logging, Metrics, SLI, SLO, SLA, Incident Management, Root Cause Analysis, RCA, Capacity Planning, Auto Scaling, Self-Healing, Disaster Recovery, Chaos Engineering, DevSecOps, Platform Engineering, Apache Kafka, High Availability, Scalability, Enterprise Infrastructure, Financial Services

#EY BR

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Riograndina

Híbrido
BRL 385 000 - 551 000
Professional growth
Competitive compensation
Flextime
+1
Senior Site Reliability / Gitops Engineer
Senior Site Reliability / Gitops Engineer

Jobgether • Brasil

Presencial
BRL 240 000 - 420 000
Annual learning budget
In-person team sprints twice a year
Performance-based bonus or commission
+2
SRE Pleno
SRE Pleno

Jobgether • Brasil

Presencial
BRL 120 000 - 210 000
Observability program
Multi-region exposure
OpenTelemetry stack
+2
Site Reliability Engineer ID55632
Site Reliability Engineer ID55632

AgileEngine • São Paulo

Híbrido
Professional growth opportunities
Competitive USD-based compensation
Exciting projects with top companies
+1
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • São Bernardo do Campo

Híbrido
BRL 250 000 - 360 000
Professional growth
Competitive compensation
Exciting projects
+1
Site Reliability Engineer - Remote Work | REF#279922
Site Reliability Engineer - Remote Work | REF#279922

BairesDev • Belo Horizonte

Teletrabalho
BRL 120 000 - 150 000
Excellent compensation in USD or local currency
Hardware and software setup for home office
Flexible working hours
+2
Site Reliability Engineer - SAP Cloud Ops
Site Reliability Engineer - SAP Cloud Ops

SAP SE • São Leopoldo

Presencial
BRL 201 000 - 312 000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

N-iX • Brasil

Híbrido
BRL 180 000 - 360 000
Flexible work format
Competitive salary
Career growth
+3
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Recife

Híbrido
BRL 298 000 - 399 000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
DevOps/SRE Engineer - São Paulo, State of São Paulo
DevOps/SRE Engineer - São Paulo, State of São Paulo

MissionHires • Brasil

Presencial
BRL 120 000 - 160 000