Cloud Site Reliability Engineer Specialist

ScotiaTech

Bogotá

Presencial

COP 90.000.000 - 150.000.000

Jornada completa

Hace 7 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

ScotiaTech is seeking a Cloud Site Reliability Engineer Specialist to ensure availability, performance, and reliability of critical corporate apps hosted on Google Cloud Platform. This role blends software and systems engineering to run large-scale, fault-tolerant systems with proactive monitoring and automation.

Responsibilities include designing scalable GCP infrastructure with Terraform, deploying containerized apps with Kubernetes, implementing GitOps with Argo CD/Flux, and leading incident

Formación

  • Post-secondary degree in CS, engineering or related field.
  • Experience in financial services or regulated industries.
  • Expert knowledge of Google Cloud Platform services, cloud networking and security.
  • Proficiency with Terraform and IaC, automation via Python.
  • Hands-on with Kubernetes and GitOps (Argo CD/Flux).
  • CI/CD design and implementation for automated deployments.
  • Strong understanding of SRE concepts: SLOs, error budgets, blameless post-mortems.

Responsabilidades

  • Design, build, and maintain scalable cloud infrastructure on GCP using Terraform.
  • Develop automation to eliminate toil and improve efficiency with Python.
  • Deploy, scale, and manage containerized apps with Kubernetes.
  • Implement secure cloud networking, VPCs, firewalls, and load balancing.
  • Enforce security protocols and secret management in cloud environments.
  • Lead GitOps adoption with Argo CD or Flux for declarative infra management.
  • Define and report SLOs/SLIs to meet reliability targets.
  • Build and optimize CI/CD pipelines for rapid deployments.
  • Lead incident response for GCP and perform blameless post-mortems.
  • Collaborate with development teams to embed reliability in lifecycle.

Conocimientos

Python
System reliability
Cloud engineering

Educación

Bachelor's degree in Computer Science or Engineering

Herramientas

Google Cloud Platform
Terraform
Kubernetes
Argo CD
Flux

Descripción del empleo

The System Reliability Engineering team focuses on System Reliability, Reliability Engineering, and Resilience across Global Corporate Functions, with a specific emphasis on Google Cloud Platform engineering and onboarding to Cloud infrastructure.

The Cloud Site Reliability Engineer Specialist is responsible for the availability, performance, and reliability of critical corporate function applications hosted on Google Cloud Platform. This role combines software and systems engineering to build and run large-scale, distributed, fault-tolerant systems. The primary objective is to enhance system resilience and reduce operational work through proactive monitoring, automation, and continuous improvement. The position's mandate is to drive the adoption of SRE principles and practices, collaborating with development teams to engineer scalable and reliable solutions from inception through to production.

Accountabilities
  • Design, build, and maintain scalable and resilient infrastructure on Google Cloud Platform using Terraform.
  • Develop and manage automation solutions with Python to eliminate operational toil and improve system efficiency.
  • Deploy, manage, and scale containerized applications using Kubernetes, ensuring optimal performance and availability.
  • Implement and administer secure cloud networking architectures, including virtual private clouds, firewall rules, and load balancing.
  • Establish and enforce robust security protocols and secret management practices within the cloud environment.
  • Drive the adoption of GitOps methodologies using tools like Argo CD or Flux for declarative infrastructure and application management.
  • Define, track, and report on Service Level Indicators and Service Level Objectives to maintain reliability targets.
  • Construct and optimize CI/CD pipelines to enable rapid, reliable, and automated software delivery.
  • Lead incident response efforts for Google Cloud Platform, facilitate blameless post-mortems, and implement corrective actions to prevent future occurrences.
  • Collaborate with application development teams to integrate reliability best practices into the software development lifecycle.
Education / Experience / Other Information
  • Completion of a post-secondary degree in Computer Science, Engineering, or a related technical field.
  • Experience working within the financial services or a similarly regulated industry.
  • Expert-level knowledge of Google Cloud Platform services, cloud networking, and security principles.
  • Familiarity with regulatory and compliance standards applicable to the banking sector.
  • Extensive experience with Infrastructure as Code, specifically using Terraform.
  • Advanced proficiency in scripting and automation using Python.
  • Demonstrated expertise in container orchestration with Kubernetes.
  • Strong understanding and practical application of GitOps principles and tools such as Argo CD or Flux.
  • Proven experience designing, building, and maintaining CI/CD pipelines for automated deployments.
  • In\u2011depth knowledge of SRE principles, including Service Level Objectives, error budgets, and blameless post-mortems.
Working Conditions

Work in a standard office-based environment; non-standard hours are a common occurrence

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Cloud Site Reliability Engineer Specialist
Cloud Site Reliability Engineer Specialist

Scotiabank • Bogotá

Presencial
COP 120.000.000 - 190.000.000
Senior Cloud SRE Engineer – Google Cloud & Automation
Senior Cloud SRE Engineer – Google Cloud & Automation

Scotiabank • Bogotá

Presencial
COP 120.000.000 - 190.000.000
Cloud SRE Specialist: Kubernetes, Terraform & Resilience
Cloud SRE Specialist: Kubernetes, Terraform & Resilience

ScotiaTech • Bogotá

Presencial
COP 90.000.000 - 150.000.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Publicis Sapient • Colombia

Presencial
COP 284.270.000 - 473.784.000
Cloud Site Reliability Engineer Associate
Cloud Site Reliability Engineer Associate

Scotiabank • Bogotá

Híbrido
COP 25.000.000 - 42.000.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MPS Group LLC • Bogotá

Presencial
COP 200.880.000 - 312.480.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Publicis Sapient • Bogotá

Presencial
COP 180.000.000 - 300.000.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Colombia

Presencial
COP 90.000.000 - 150.000.000
SRE Software Engineer
SRE Software Engineer

Jobtailor • Bogotá

Presencial
COP 100.000.000 - 180.000.000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Metropolitana

Híbrido
COP 142.369.000 - 213.554.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1