Get more replies from employers
Send a job-specific resume in minutes.
Datonomy Solutions (Pty) Ltd is seeking a senior DevOps & Site Reliability Engineer to design, build and operate highly available enterprise platforms. You will work across engineering and delivery teams to boost reliability, deployment velocity, resilience and production performance.
The role spans DevOps, SRE, Azure Cloud, Platform Engineering, Kubernetes, IaC, CI/CD, Observability and DevSecOps. You will lead incident response, monitor SLIs/SLOs/SLAs and drive automation across hybrid-cloud
This is a senior hands-on engineering role spanningDevOps, Site Reliability Engineering, Azure Cloud, Platform Engineering, Kubernetes, Infrastructure as Code, CI/CD, Observability and DevSecOps.
The successful candidate will work across engineering and delivery teams to improveplatform reliability, deployment velocity, resilience, automation, operational efficiency and production performance, while supporting mission-critical enterprise applications.
Design, build and maintaincloud-native infrastructure and platform services.
Develop and maintainInfrastructure as Code (IaC)solutions.
Automate infrastructure provisioning, configuration and operational processes.
Build reusable engineering tools, deployment templates and platform components.
Establish and standardise platform engineering practices across multiple delivery teams.
Identify opportunities to reduce manual intervention and increase engineering automation.
Design, implement and maintain enterprise-gradeCI/CD pipelinesfor application and infrastructure deployments.
Implement automated testing, security scanning, code-quality controls and release automation.
Enable automated deployments, rollback and recovery processes.
Improve deployment frequency while reducing change and deployment risk.
Continuously optimise software delivery and release-management processes.
Implement and matureSite Reliability Engineering practicesacross production environments.
Define, monitor and manageService Level Indicators (SLIs), Service Level Objectives (SLOs) and Service Level Agreements (SLAs).
Improve application and platformavailability, scalability, resilience and performance.
Lead production incident response, troubleshooting, problem management andRoot Cause Analysis (RCA).
Drive proactive reliability improvements and reduction of technical debt.
ImproveMean Time to Detect (MTTD)andMean Time to Recover (MTTR).
Design, implement and operate enterpriseMicrosoft Azureenvironments.
Work extensively with technologies such as:
Azure Kubernetes Service (AKS)
Azure App Services
Azure Networking
Azure Monitor
Azure Storage
Azure Identity Services
Design highly available and disaster-recovery-capable environments.
Optimise cloud environments forperformance, resilience, security and cost.
Support hybrid-cloud and multi-cloud environments where required.
Build, deploy and support containerised applications usingDockerandKubernetes.
Manage Kubernetes environments, particularlyAzure Kubernetes Service (AKS).
Develop and maintain deployment configurations usingHelm.
Support container-platform reliability, scalability and operational performance.
OpenShift experience would be advantageous.
Hands-on experience with technologies such as:
Terraform
Bicep
ARM Templates
Ansible
Candidates should be comfortable using Infrastructure as Code to build repeatable, scalable and governed enterprise infrastructure.
Implement comprehensivelogging, monitoring, metrics, tracing and alerting.
Build operational dashboards and platform insights.
Establish enterprise observability standards.
Implement proactive and predictive monitoring.
Use observability information to improve application and infrastructure reliability.
Relevant technologies may include:
Dynatrace
Grafana
Prometheus
Elastic Stack / ELK
Splunk
Azure Monitor
OpenTelemetry
EmbedDevSecOpspractices throughout the software-delivery lifecycle.
Integrate security scanning and controls into CI/CD pipelines.
Support vulnerability identification, remediation and risk reduction.
Ensure cloud and platform environments comply with enterprise security and regulatory requirements.
Work closely with information-security teams to continuously improve platform security.
Provide technical leadership across DevOps, Cloud, Platform and SRE teams.
Mentor and coach junior and intermediate engineers.
Contribute to architecture decisions and technology roadmaps.
Promote engineering standards and operational best practice.
Lead cross-functional initiatives aimed at improving engineering productivity and reliability.
8+ years' experienceacross software engineering, infrastructure engineering, cloud engineering, DevOps or platform engineering.
5+ years' hands-on DevOps engineering experience.
3+ years' Site Reliability Engineering or production-operations experience.
Proven experience supportingmission-critical production systems.
Experience operatinglarge-scale enterprise technology platforms.
Strong exposure to highly available and business-critical environments.
Microsoft Azure
Azure Kubernetes Service (AKS)
Azure Networking
Azure App Services
Azure Monitor
Azure Storage
Azure Identity
Azure DevOps
GitHub / Git
Jenkins
SonarQube
Artifactory and/or Nexus
Terraform
Bicep
ARM Templates
Ansible
Kubernetes
Docker
Helm
Dynatrace
Grafana
Prometheus
Elastic Stack
Splunk
Azure Monitor
OpenTelemetry
Strong scripting or programming ability using technologies such as:
Python
PowerShell
Bash
C#
Java
Go experience would be advantageous.
DevOps Engineering
Site Reliability Engineering
Azure Cloud Engineering
Platform Engineering
Infrastructure Automation
Kubernetes / Container Orchestration
CI/CD
Infrastructure Automation
DevSecOps
Cloud Architecture
Observability
Continuous Delivery
Systems Integration
Capacity Planning
Performance Optimisation
Incident & Problem Management
Root Cause Analysis
Strong technical problem-solving ability
Strategic thinking
Strong decision-making skills
Collaboration across engineering disciplines
Stakeholder management
Continuous-improvement mindset
Coaching and mentoring capability
Accountability and ownership
Customer-centric approach
A Bachelor's Degree or equivalent technical qualification in one of the following areas is preferred:
Computer Science
Information Technology
Software Engineering
Information Systems
Relevant certifications would be advantageous, including:
Microsoft Certified:Azure DevOps Engineer Expert
Microsoft Certified:Azure Solutions Architect Expert
Certified Kubernetes Administrator (CKA)
Certified Kubernetes Application Developer (CKAD)
HashiCorp Terraform Associate
AWS Certified DevOps Engineer
ITIL Foundation
SRE Foundation Certification
The ideal candidate is a senior, hands-on engineer who can bridgesoftware development, cloud infrastructure, DevOps, platform engineering and production operations.
They should have deep experience building and running highly available enterprise environments and possess a strongautomation-first and reliability-focused mindset.
This person should be equally comfortable troubleshooting a critical production issue, building Terraform infrastructure, improving a Kubernetes platform, designing a CI/CD pipeline, implementing observability, defining SLOs and mentoring other engineers.