Cloud Reliability Engineer

Infios US, Inc.

Brasil

Presencial

BRL 120 000 - 150 000

Tempo integral

14 dias+

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

A leading technology provider in Brazil is seeking a Cloud Infrastructure Engineer. The ideal candidate will have strong expertise in cloud platforms and Kubernetes management. Responsibilities include improving cloud infrastructure, automating processes, and ensuring system reliability. This position requires a minimum of 5 years of relevant experience. Candidates with a Bachelor's degree in computer science or equivalent are preferred. Join us in enhancing supply chains through innovative technology solutions.

Qualificações

  • 5+ years of experience in Cloud Engineering, DevOps, or Site Reliability roles.
  • Strong knowledge of Kubernetes deployment and troubleshooting.
  • Hands-on experience with cloud platforms (AWS, Azure, GCP).
  • Proficiency in scripting and automation tools.
  • Experience with incident response and postmortem processes.

Responsabilidades

  • Operate and maintain cloud infrastructure in AWS, Azure, or GCP.
  • Manage Kubernetes clusters, ensuring system availability and performance.
  • Build automated pipelines for deployment and configuration.
  • Implement SRE principles and monitor critical services.
  • Collaborate with teams to maintain reliable systems.

Conhecimentos

Cloud Engineering
DevOps
Site Reliability Engineering
Kubernetes management
Scripting (Python, Bash)
Automation
Incident response
Monitoring (Dynatrace, DataDog)

Formação académica

Bachelor’s degree in Computer Science or related field

Ferramentas

AWS
Azure
GCP
Terraform
Ansible

Descrição da oferta de emprego

If you are looking for a meaningful career where people work and act with passion, rethink the existing and always strive to find the best solution - you have come to the right place. We develop future technologies to relentlessly make supply chains better.We are a leader in supply chain software solutions, helping organizations streamline operations, reduce costs, and improve efficiency.**Key Responsibilities** Cloud Infrastructure Operationso Operate, maintain, and improve cloud infrastructure in AWS, Azure, or GCP environments.o Manage and optimize Kubernetes clusters — deployment, scaling, patching, and upgrades.o Ensure system availability, scalability, and performance through proactive monitoring and optimization.o Maintain infrastructure-as-code (IaC) for consistent and repeatable deployments. Automation & Continuous Improvemento Identify opportunities for operational automation to eliminate manual processes (“reduce toil”).o Build and maintain automated pipelines for deployments, configuration, and remediation.o Develop self-healing mechanisms to automatically detect and resolve common service issues.o Participate in continuous improvement initiatives around reliability, performance, and efficiency. Reliability Engineeringo Implement SRE principles: define and track SLIs, SLOs, and error budgets.o Perform incident analysis and postmortems to identify root causes and prevent recurrence.o Design proactive monitoring, alerting, and observability dashboards (Dynatrace, DataDog).o Collaborate with DevOps and development teams to build reliable, observable, and resilient systems. CI/CD and Release Operationso Manage and optimize CI/CD pipelines to ensure reliable and consistent delivery.o Support deployment strategies (blue/green, canary, rolling) to reduce downtime risk.o Collaborate with Product and DevOps teams on release readiness and rollback automation. Incident Response & Troubleshootingo Monitor, troubleshoot, and resolve infrastructure and application issueso Respond to production incidents and ensure rapid mitigation and resolution.o Troubleshoot complex cloud, container, and networking issues across distributed systems.o Drive a culture of proactive monitoring, data-driven analysis, and preventive action.**Required Qualifications** Bachelor’s degree in computer science, Engineering, or related field (or equivalent experience). 5+ years of experience in experience in Cloud Engineering, DevOps, or Site Reliability roles. Hands-on experience with cloud platforms (OCI, AWS, Azure, or GCP). Strong knowledge of Kubernetes deployment, management, and troubleshooting Solid understanding of observability and monitoring (e.g., Dynatrace, DataDog) and incident management platforms. Proficiency in scripting and automation (e.g., Python, Bash, Terraform, Ansible). Strong troubleshooting and analytical skills across infrastructure and applications. Experience with incident response, RCA, and postmortem processes. A mindset of continuous improvement, reliability, and self-healing automation. Understanding of SRE principles, SLAs/SLOs/SLIs, and chaos engineering practices.**Preferred Skills** Experience in conducting resilience assessments and recovery drills. Familiarity with ServiceNow and Dynatrace or other observability and ITSM tools. Experience with chaos engineering or resiliency testing frameworks Background in networking, load balancing, and performance tuning Strong communication and stakeholder management skills.**Soft Skills & Mindset** Strong collaboration skills — comfortable working with developers, ops, and management. Clear communicator; able to translate technical issues into business impact. Self-starter with a problem-solving and automation-first mentality. Resilient under pressure — thrives in a dynamic, fast-paced environment. Passionate about operational excellence and continuous learning.**Key Success Metrics** SLA/SLO compliance for critical services Reduction in MTTR (Mean Time to Recover) Increase in automated incident resolution rates Reduction in customer-impacting incidents Frequency and outcomes of resilience testing exercises Service uptime / availability**Why join us**At Infios, we're not just looking for employees; we're looking for partners in innovation, growth, and purpose. Meeting you where you are to create the future you need is at the core of who we are and what we do. Whether you're at the beginning of your career or a seasoned expert, we meet you on your journey, equipping you with the tools and opportunities to build the future you envision. Together, we will relentlessly work toward one common goal - making supply chains better.**We believe the future is better when supply chains work better.**We are an equal-opportunity employer and committed to inclusion in the workplace.At Infios, we believe that inclusion is a fundamental cornerstone of our success. We are committed to creating a safe and welcoming environment where every individual’s unique experiences and perspectives are valued—whether they look, think, move, believe, or love differently.All qualified applicants will receive consideration for employment without regard to race, color, ethnicity, national origin, sex, sexual orientation, gender identity, marital status, pregnancy, religion, age, disability, veteran status, genetic information, or any other characteristic protected by law. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions of this role. If you require assistance or accommodation due to a disability during the recruiting process, please let us know at jobs@infios.com Disclaimer: This job advertisement is not designed to cover a comprehensive listing of all duties or responsibilities that are required for this job. Please note that any salary information is a general guideline only. Individual compensation will be determined by various factors such as the scope and responsibilities of the position, experience, education, skills, location, and market and business considerations. Applications must be submitted via our career site.
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Senior Software Engineer
Senior Software Engineer

Infios US, Inc. • Brasil

Teletrabalho
BRL 120 000 - 180 000
Equipped with tools for professional growth
Commitment to inclusion and diversity
Software Engineer
Software Engineer

Infios • Brasil

Presencial
BRL 90 000 - 120 000
Senior System Developer
Senior System Developer

Infios US, Inc. • Brasil

Híbrido
BRL 120 000 - 160 000
Technical Support L1
Technical Support L1

Infios • Brasil

Presencial
Inclusive work environment
Opportunities for professional growth
Senior Systems Developer
Senior Systems Developer

Infios • Blumenau

Híbrido
BRL 120 000 - 150 000
Inclusive workplace
Career growth opportunities
Remote work flexibility
Service Desk Analyst
Service Desk Analyst

Infios • Brasil

Presencial
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Recife

Híbrido
BRL 298 000 - 399 000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Support Operations Analyst
Support Operations Analyst

Infios • Brasil

Presencial
Technical Support L1
Technical Support L1

Infios US, Inc. • Brasil

Teletrabalho
BRL 208 000 - 364 000
Site Reliability Engineer ID55632
Site Reliability Engineer ID55632

AgileEngine • São Paulo

Híbrido
Professional growth opportunities
Competitive USD-based compensation
Exciting projects with top companies
+1