Senior Site Reliability Engineer

Jobgether

Brasil

Presencial

BRL 180 000 - 320 000

Tempo integral

Há 8 dias

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Vantagens oferecidas por esta oferta de emprego

Remote-work equipment
Home office budget
Training & development allowance
Wellness budget
Paid vacation
Volunteer day
Global collaboration
Career growth

Resumo da oferta

Jobgether is seeking a Senior Site Reliability Engineer based in Brazil to design, operate, and continuously improve large-scale distributed infrastructure across cloud, Kubernetes, networking, storage, observability, and Linux systems.

You will automate, monitor, and optimize platforms, collaborating with internal teams, clients, and AI/ML specialists to ensure reliability and performance for demanding workloads.

Qualificações

  • Extensive experience in SRE/DevOps for cloud-based platforms.
  • Hands-on with GCP and Terraform for IaC and deployments.
  • Strong knowledge of Kubernetes, Docker, Istio, and modern networking.
  • Linux administration and troubleshooting capabilities.
  • Automation-focused with emphasis on observability and reliability.
  • Willingness to participate in on-call rotations.

Responsabilidades

  • Operate, scale, and maintain large-scale Kubernetes clusters and Istio service mesh.
  • Design automation workflows to reduce manual tasks and improve efficiency.
  • Build and maintain monitoring with Prometheus, Grafana, and Loki.
  • Troubleshoot complex infrastructure, networking, storage, and performance issues.
  • Collaborate with AI/ML teams to support training workloads and data pipelines.
  • Lead postmortems and drive improvements to prevent recurrence.
  • Engage with clients and internal teams to deliver reliable infrastructure solutions.

Conhecimentos

SRE/DevOps
Cloud infrastructure
GCP
Networking
Automation
On-call

Ferramentas

Terraform
Kubernetes
Docker
Istio
Prometheus
Grafana

Descrição da oferta de emprego

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Brazil.

This role offers the opportunity to design, operate, and continuously improve large-scale, distributed infrastructure in a highly technical environment.

You will work across cloud, Kubernetes, networking, storage, observability, and Linux systems to deliver resilient and high-performing platforms.

The position combines hands‑on engineering with automation, intelligent monitoring, and architectural problem-solving.

You will collaborate closely with internal engineering teams, clients, and AI/ML specialists to ensure infrastructure is reliable and ready for demanding workloads.

The role is ideal for an experienced SRE who enjoys solving complex operational challenges and improving systems through automation.

You will contribute to a culture focused on scalability, reliability, continuous learning, and operational excellence.

This is a strong opportunity to make a direct impact on modern cloud and AI/ML infrastructure while expanding your technical expertise.

Accountabilities
  • Operate, maintain, and optimize Kubernetes clusters, Istio service mesh, and Linux-based systems to ensure reliability, scalability, and performance.
  • Design and implement automation workflows using Go, Python, and Shell scripting to reduce manual operations and improve efficiency.
  • Build and maintain monitoring and observability solutions using Prometheus, Grafana, and Loki.
  • Diagnose and resolve complex issues involving networking, storage, infrastructure, and system performance.
  • Collaborate with AI/ML teams to ensure infrastructure is prepared to support model training, data pipelines, and other demanding workloads.
  • Contribute to infrastructure architecture, deployment, automation, and continuous improvement initiatives.
  • Participate in on‑call rotations and incident response activities to maintain system availability and reliability.
  • Lead and contribute to postmortem reviews, identifying root causes and implementing improvements to prevent recurring incidents.
  • Work collaboratively with clients and multidisciplinary engineering teams to deliver resilient, high-performing infrastructure solutions.
Requirements
  • Strong professional experience in Site Reliability Engineering, DevOps, cloud infrastructure, or a closely related field.
  • Hands‑on experience with Google Cloud Platform (GCP) and Infrastructure as Code tools such as Terraform.
  • Strong knowledge of microservices, containers, Kubernetes, Docker, and modern networking concepts.
  • Practical experience with Linux systems administration and troubleshooting.
  • Experience with PKI and service mesh technologies, preferably including Istio.
  • Strong understanding of SRE principles, with a focus on automation, scalability, availability, observability, and reliability.
  • Experience troubleshooting complex infrastructure, networking, storage, and performance problems.
  • Strong scripting and automation capabilities using tools such as Python and Shell.
  • Experience with Golang is considered an asset.
  • Ability to work effectively in fast‑paced, problem‑solving environments and collaborate with both technical teams and clients.
  • Strong ownership, analytical thinking, communication, and continuous‑improvement mindset.
  • Willingness and ability to participate in an on‑call rotation.
Benefits
  • Competitive total rewards package.
  • Remote‑work equipment, including a laptop with your choice of operating system.
  • Annual budget to personalize and improve your home‑work environment.
  • Substantial training and professional development allowance.
  • Opportunities to attend training, pursue certifications, and participate in professional development days.
  • Annual wellness budget that can be used for activities such as gym memberships, fitness, massages, and other wellness initiatives.
  • Generous paid vacation and sick leave.
  • Paid day off to volunteer for a charity of your choice.
  • Opportunity to collaborate with highly experienced professionals in a global technical environment.
  • Support for continuous learning, career development, and technical growth.
  • Background‑check requirements apply to the successful candidate.
  • Reasonable accommodations are available upon request during the selection process.
How Jobgether Works

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top‑fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?
Data Privacy Notice

By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre‑contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Triwill Group • Brasil

Presencial
BRL 250 000 - 380 000
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)

iTRTech Group • São Paulo

Híbrido
BRL 180 000 - 260 000
Site Reliability Engineer - Remote Work | REF#279922
Site Reliability Engineer - Remote Work | REF#279922

BairesDev • Belo Horizonte

Teletrabalho
BRL 120 000 - 150 000
Excellent compensation in USD or local currency
Hardware and software setup for home office
Flexible working hours
+2
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Riograndina

Híbrido
BRL 385 000 - 551 000
Professional growth
Competitive compensation
Flextime
+1
Senior Back-end Engineer (C#/.NET) - Software
Senior Back-end Engineer (C#/.NET) - Software

Jobgether • Brasil

Presencial
BRL 622 000 - 933 000
100% remote work
USD-based compensation
Paid time off
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Zoomcar • Brasil

Presencial
BRL 422 000 - 485 000
Stock options
Health benefits
Unlimited PTO
+2
Site Reliability Engineer ID55632
Site Reliability Engineer ID55632

AgileEngine • São Paulo

Híbrido
Professional growth opportunities
Competitive USD-based compensation
Exciting projects with top companies
+1
Site Reliability Engineer - SAP Cloud Ops
Site Reliability Engineer - SAP Cloud Ops

SAP SE • São Leopoldo

Presencial
BRL 201 000 - 312 000
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • São Bernardo do Campo

Híbrido
BRL 250 000 - 360 000
Professional growth
Competitive compensation
Exciting projects
+1
Site Reliability Engineer - SAP Cloud Ops
Site Reliability Engineer - SAP Cloud Ops

SAP • São Leopoldo

Presencial
BRL 180 000 - 260 000