DevOps Engineer

WatersEdge Solutions

South Africa

Hybrid

ZAR 900,000 - 1,350,000

Full time

34 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Hybrid work
Wellness support
Learning opportunities
Employee share ownership opportunities

Job summary

WatersEdge Solutions is seeking a senior DevOps/SRE Cloud Engineer to own the production Azure-based Kubernetes platform. You will build, operate and secure resilient cloud infrastructure, improve CI/CD pipelines and establish observability to run workloads reliably at scale.

You’ll work across Azure, Kubernetes, Terraform, Docker, and GitHub Actions, owning from deployment to incident response while collaborating with engineering teams in a hybrid/remote setup.

Qualifications

  • 5+ years of DevOps, SRE or platform engineering experience.
  • 2–3+ years operating Kubernetes workloads in production.
  • 2+ years hands-on Azure experience.
  • Strong production experience with Azure Kubernetes Service (AKS).
  • Hands-on with ADLS Gen2, Azure Key Vault and Entra ID.
  • Understanding of Azure networking, VNets and firewalls.
  • Experience with Terraform and CI/CD pipelines (GitHub Actions).
  • Proficient in Bash scripting and Linux administration.
  • Observability with Prometheus and Grafana; incident response and SLOs.

Responsibilities

  • Operate and maintain production Kubernetes clusters and Azure cloud infrastructure.
  • Manage modular Terraform code across environments.
  • Own deployments, configuration, upgrades and troubleshooting.
  • Design Helm charts and environment-specific deployment configurations.
  • Maintain secure CI/CD pipelines using GitHub Actions.
  • Manage container image builds and health checks.
  • Maintain platform observability with dashboards and alerts.
  • Define SLOs and alerting strategies; participate in incident response.
  • Capacity planning based on production workloads.
  • Manage cloud secrets and network security across environments.

Skills

DevOps
SRE
Kubernetes production
Azure
Terraform
Docker
GitHub Actions
Prometheus
Grafana
Bash scripting
Linux administration
Secrets management

Education

Bachelor’s degree in Computer Science or Engineering

Tools

AKS
Terraform
Docker
GitHub Actions
Prometheus
Grafana
Azure Key Vault
Azure AD / Entra ID
Helm
Terraform State Management

Job description

Industry: Cloud Infrastructure | DevOps | Site Reliability Engineering | Data Technology

WatersEdge Solutions is partnering with a growing technology business to appoint a DevOps / SRE Cloud Engineer (Site Reliability Engineer) to take end-to-end ownership of its Azure-based Kubernetes platform.

This is a senior, hands-on engineering role focused on building and operating resilient cloud infrastructure, strengthening CI/CD, and creating the observability needed to run production workloads reliably at scale. You’ll work across Azure, Kubernetes, Terraform, Docker, GitHub Actions and SRE practices, with the autonomy to own the environment and continuously improve its reliability, security and performance.

About the Role

As DevOps / SRE Cloud Engineer, you’ll take ownership of the production Kubernetes and Azure environment, managing infrastructure as code and ensuring workloads are secure, observable and dependable.

You’ll operate Kubernetes clusters, maintain modular Terraform infrastructure, build and improve deployment pipelines, manage cloud identity and secrets, and establish effective monitoring, alerting and SLOs.

This role requires someone who is comfortable going beyond maintaining infrastructure. You’ll be expected to understand how the platform behaves under real production load, troubleshoot complex issues alongside engineering teams, respond effectively to incidents and make practical improvements that prevent problems from recurring.

Key Responsibilities
  • Operate and maintain production Kubernetes clusters and Azure cloud infrastructure.
  • Manage infrastructure through modular Terraform code across multiple environments.
  • Own Kubernetes deployments, configuration, upgrades and operational troubleshooting.
  • Design and maintain Helm charts and environment-specific deployment configurations.
  • Manage Kubernetes node pools, resource requests and limits, and capacity requirements.
  • Maintain secure and reliable CI/CD pipelines using GitHub Actions.
  • Implement linting, testing and build quality gates for production deployments.
  • Manage container image build and publishing workflows.
  • Maintain branch protection and required-status-check processes.
  • Own platform observability across dashboards, metrics, alerting and structured logging.
  • Define and maintain meaningful SLOs and alerting strategies.
  • Participate in and improve incident response processes.
  • Perform capacity planning based on measured production workloads.
  • Manage cloud secrets, identity and network security across environments.
  • Maintain secure integration with Azure Key Vault and appropriate identity mechanisms.
  • Manage Azure networking requirements, including VNets, private endpoints and firewall rules.
  • Maintain Docker images, container health checks and local Compose-based development environments.
  • Build and maintain operational tooling and scripts using Bash.
  • Partner with software and data engineering teams to troubleshoot production issues and improve platform reliability.
  • Ensure operational scripts and infrastructure code are maintained with the same discipline as application code.
What You’ll Bring
  • 5+ years of DevOps, SRE or platform engineering experience.
  • 2–3+ years of experience operating Kubernetes workloads in production.
  • 2+ years of hands-on Azure experience.
  • Strong production experience with Azure Kubernetes Service (AKS).
  • Hands-on experience with ADLS Gen2, Azure Key Vault and Entra ID.
  • Understanding of Azure workload identities and managed identities.
  • Strong Azure networking experience, including VNets, private endpoints and firewall rules.
  • Experience with Azure Service Bus or an equivalent message broker.
  • Strong production Kubernetes knowledge, including Helm chart authoring and environment overlays.
  • Experience working with Kubernetes operators.
  • Understanding of node-pool design, capacity sizing and resource requests/limits.
  • Strong Kubernetes pod troubleshooting and upgrade experience.
  • Advanced practical experience with Terraform and modular Infrastructure as Code.
  • Experience managing Terraform state and disciplined plan/apply processes across environments.
  • Strong Docker experience, including image builds, multi-architecture builds, health checks and memory-limit tuning.
  • Experience maintaining Docker Compose-based local development environments.
  • Strong CI/CD experience using GitHub Actions.
  • Experience implementing automated lint, test and build gates.
  • Hands-on experience with Prometheus and Grafana.
  • Practical understanding of SRE principles, including SLOs, alerting, incident response and capacity planning.
  • Strong Linux administration and Bash scripting skills.
  • Experience implementing secure secrets-management practices.
  • Ability to independently own a production cloud environment.
  • Strong communication skills and the ability to collaborate effectively within hybrid and remote engineering teams.
Nice to Have
  • Experience operating Apache Spark on Kubernetes.
  • Knowledge of Spark Operator, Spark Connect and executor/driver tuning.
  • Exposure to open data-platform technologies including Hive Metastore, Trino, Apache Ranger and Delta Lake.
  • Experience with S3-compatible object storage such as MinIO.
  • Understanding of governed data and query architectures.
  • OpenTelemetry metrics and tracing experience.
  • Python experience for operational tooling and infrastructure testing.
  • Experience using pytest for environment or infrastructure test suites.
  • Operational experience with SQL Server and PostgreSQL.
  • Database backup and container deployment experience.
  • Experience within multi-tenant or regulated-data environments.
  • Understanding of tenant isolation patterns.
  • Exposure to POPIA, GDPR or ISO 27001-aligned controls.
  • Experience with secret scanning, dependency auditing and container image provenance.
  • Azure and/or Kubernetes certifications such as CKA.
Qualifications
  • Bachelor’s degree in Computer Science, Engineering or a related discipline, or equivalent practical experience.
  • Proven experience operating production‑grade cloud infrastructure at scale.
  • Azure, Kubernetes or related cloud certifications are advantageous but not essential.
What’s On Offer
  • Flexible hybrid and remote working arrangements.
  • End-to-end ownership of a production Azure and Kubernetes environment.
  • Opportunity to work with modern cloud, containerisation and infrastructure technologies.
  • Exposure to complex data-platform infrastructure and distributed workloads.
  • Scope to shape platform reliability, observability, automation and security practices.
  • Wellness initiatives and home‑office support.
  • Continuous learning and professional development opportunities.
  • Supportive and inclusive team environment.
  • Regular team‑building activities and social events.
  • Performance recognition and employee share ownership opportunities.
  • A culture that values transparency, accountability and work‑life balance.
Company Culture

You’ll be joining a technically ambitious environment where engineers are trusted to take ownership and make meaningful improvements to the systems they manage.

The team values reliability, automation and strong engineering discipline, while maintaining a collaborative and supportive working style. You’ll have the autonomy to identify weaknesses, improve infrastructure and tooling, and work closely with engineering teams to create a platform that can scale reliably as the business grows.

Please Note: If you have not been contacted with 10 working days, consider your application unsuccessful
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps / SRE Cloud Engineer (Site Reliability Engineer)
DevOps / SRE Cloud Engineer (Site Reliability Engineer)

Recruit-It • South Africa

Hybrid
ZAR 900,000 - 1,500,000
Hybrid work
Home office stipend
Learning budget
+2
DevOps / SRE Cloud Engineer (Site Reliability Engineer) - Remote
DevOps / SRE Cloud Engineer (Site Reliability Engineer) - Remote

Recruit-It • South Africa

Hybrid
ZAR 900,000 - 1,500,000
Flexible hours
Hybrid/Remote options
Wellness program
+3
Senior DevOps & Site Reliability Engineer at Datonomy Solutions
Senior DevOps & Site Reliability Engineer at Datonomy Solutions

Datonomy Solutions • Emfuleni Local Municipality

On-site
ZAR 900,000 - 1,500,000
Senior DevOps & Site Reliability Engineer at Datonomy Solutions
Senior DevOps & Site Reliability Engineer at Datonomy Solutions

Datonomy Solutions • Emfuleni Local Municipality

Hybrid
ZAR 900,000 - 1,300,000
Biz Dev Ops Engineer (Contract)
Biz Dev Ops Engineer (Contract)

The Focus Group • Sandton

On-site
ZAR 1,000,000 - 1,800,000
Senior DevOps & Site Reliability Engineer
Senior DevOps & Site Reliability Engineer

Datonomy Solutions (Pty) Ltd • Sandton

On-site
ZAR 1,200,000 - 2,000,000
Senior DevOps & SRE Cloud Engineer (Azure/Kubernetes)
Senior DevOps & SRE Cloud Engineer (Azure/Kubernetes)

WatersEdge Solutions • Cape Town

On-site
ZAR 900,000 - 1,800,000
Flexible hybrid/remote work
End-to-end ownership
Professional development
+2
Senior DevOps & SRE — Azure Kubernetes Platform
Senior DevOps & SRE — Azure Kubernetes Platform

WatersEdge Solutions • South Africa

Hybrid
ZAR 900,000 - 1,350,000
Hybrid work
Wellness support
Learning opportunities
+1
Senior DevOps & Site Reliability Engineer
Senior DevOps & Site Reliability Engineer

DeARX • Sandton

On-site
ZAR 1,000,000 - 1,400,000
DevOps / SRE Cloud Engineer Site Reliability Engineer - WatersEdge Solutions
DevOps / SRE Cloud Engineer Site Reliability Engineer - WatersEdge Solutions

OpenTalent • Cape Town

On-site
ZAR 600,000 - 1,000,000