Senior DevOps and Site Reliability Engineer (SRE)

Indsafri

South Africa

On-site

ZAR 900,000 - 1,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Indsafri is seeking a Senior DevOps and Site Reliability Engineer (SRE) to design, implement, and operate scalable, secure, and automated cloud platforms. You will drive best practices across delivery teams, ensuring reliable deployments and robust platform resilience.

The role emphasizes IaC, CI/CD, monitoring, security, and cross-functional leadership to improve availability, performance, and customer experience. Strong Kubernetes or OpenShift experience is a plus.

Qualifications

  • Bachelor's degree in a technical field or equivalent experience.
  • 8+ years in software engineering, infrastructure, cloud, DevOps, or platform engineering.
  • 5+ years hands-on DevOps engineering experience.
  • 3+ years in Site Reliability Engineering (SRE) or production operations.
  • Proven track record managing mission-critical production systems.

Responsibilities

  • Design, implement, and maintain highly available, scalable platforms.
  • Develop IaC solutions and reusable deployment templates.
  • Build and maintain CI/CD pipelines with automated testing and security checks.
  • Lead incident response, problem management, and root cause analysis.
  • Define and manage SLIs, SLOs, and SLAs.
  • Drive reliability, performance, and cost optimization initiatives.
  • Mentor DevOps, platform, cloud, and SRE engineers.

Skills

Cloud experience
CI/CD
Monitoring & observability
Security & compliance
Leadership & mentoring
Problem solving

Education

Bachelor's Degree in Computer Science
Bachelor's Degree in Information Technology
Bachelor's Degree in Software Engineering
Bachelor's Degree in Information Systems

Tools

Git
Jenkins
SonarQube
Nexus
Terraform
Ansible
Docker
OpenShift

Job description

Senior DevOps and Site Reliability Engineer (SRE)

The Senior DevOps and Site Reliability Engineer is responsible for designing, implementing, and maintaining highly available, scalable, secure, and automated technology platforms that support mission-critical business applications. The role combines software engineering, platform engineering, cloud infrastructure, automation, observability, and operational excellence to improve system reliability, deployment velocity, platform resilience, and customer experience.

The incumbent will drive DevOps and SRE best practices across delivery teams, ensuring that systems are built, deployed, monitored, and operated efficiently while maintaining stringent availability, security, and performance standards.

Key Responsibilities

Platform Engineering and Automation

  • Design, build, and maintain cloud-native infrastructure and platform services.
  • Develop Infrastructure as Code (IaC) solutions using modern automation frameworks.
  • Build reusable deployment templates, pipelines, and automation tooling.
  • Standardize platform engineering practices across teams.
  • Automate provisioning, configuration management, and operational processes.

DevOps and CI/CD

  • Design and maintain CI/CD pipelines supporting both application and infrastructure deployments.
  • Implement automated testing, security scanning, code quality controls, and release automation.
  • Drive continuous improvement of release management processes.
  • Enable fully automated deployment and rollback capabilities.
  • Improve deployment frequency while reducing deployment risk.

Site Reliability Engineering (SRE)

  • Establish reliability engineering practices and operational standards.
  • Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
  • Improve system availability, performance, resilience, and scalability.
  • Lead incident response, problem management, and root cause analysis activities.
  • Drive proactive reliability improvements and technical debt reduction.

Cloud Operations

  • Design and manage cloud infrastructure environments.
  • Optimize performance, availability, security, and cost management.
  • Implement high-availability and disaster recovery solutions.
  • Support hybrid-cloud and multi-cloud environments where applicable.

Monitoring and Observability

  • Implement enterprise monitoring and observability platforms.
  • Establish logging, metrics, tracing, and alerting standards.
  • Build operational dashboards and platform insights.
  • Reduce mean time to detect (MTTD) and mean time to recover (MTTR).
  • Drive predictive monitoring and proactive issue detection.

Security & Compliance

  • Implement DevSecOps practices throughout the software delivery lifecycle.
  • Integrate security controls into CI/CD pipelines.
  • Support vulnerability management and remediation activities.
  • Ensure compliance with organizational security and regulatory requirements.
  • Collaborate with security teams to improve platform security posture.
  • Provide technical leadership across engineering teams.
  • Mentor DevOps, platform, cloud, and reliability engineers.
  • Promote engineering excellence and operational best practices.
  • Contribute to architectural decisions and technology roadmaps.
  • Lead cross-functional initiatives to improve engineering productivity.

Minimum Qualifications

Preferred:

  • Bachelor's Degree in:
  • Computer Science
  • Information Technology
  • Software Engineering
  • Information Systems
  • 8+ years of software engineering, infrastructure, cloud, DevOps, or platform engineering experience.
  • 5+ years of hands-on DevOps engineering experience.
  • 3+ years in Site Reliability Engineering (SRE) or production operations environments.
  • Proven experience managing mission-critical production systems.

Technical Skills

Strong experience with:

  • Azure Identity Services

Exposure to AWS and Google Cloud is advantageous.

DevOps Tooling

  • Git
  • Jenkins
  • SonarQube
  • Nexus

Infrastructure as Code

  • Terraform
  • Bicep
  • ARM Templates
  • Ansible

Containerisation and Orchestration

  • Docker
  • OpenShift (advantageous)

Observability

  • Dynatrace
  • Grafana
  • Elastic Stack
  • Splunk
  • OpenTelemetry

Programming and Scripting

  • Python
  • PowerShell
  • Bash
  • C#
  • Java
  • Go (advantageous)

Core Competencies

  • Reliability Engineering
  • Infrastructure Automation
  • DevSecOps
  • Systems Integration
  • Capacity Planning
  • Performance Optimization
  • Strategic Thinking
  • Problem Solving
  • Decision Making
  • Stakeholder Management
  • Coaching and Mentoring
  • Customer Centricity
  • Accountability

Key Performance Indicators (KPIs)

The role will be measured against:

  • Platform availability targets
  • SLO/SLA compliance
  • Deployment success rate
  • Mean Time to Detect (MTTD)
  • Mean Time to Recover (MTTR)
  • Automation coverage
  • Security and compliance adherence
  • Cost optimization targets
  • Engineering productivity improvements

Preferred Certifications

  • Microsoft Certified: Azure Solutions Architect Expert
  • HashiCorp Terraform Associate
  • ITIL Foundation
  • SRE Foundation Certification

Ideal Candidate Profile

An experienced engineering professional who can bridge software development, cloud infrastructure, platform engineering, and operations. The successful candidate will possess deep technical expertise, a strong automation mindset, and a passion for building reliable, secure, scalable platforms while enabling engineering teams to deliver business value rapidly and safely.

Skills
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Biz Dev Ops Engineer (Contract)
Biz Dev Ops Engineer (Contract)

The Focus Group • Sandton

On-site
ZAR 1,000,000 - 1,800,000
Senior DevOps & Site Reliability Engineer
Senior DevOps & Site Reliability Engineer

Datonomy Solutions (Pty) Ltd • Sandton

On-site
ZAR 1,200,000 - 2,000,000
Senior DevOps & Site Reliability Engineer
Senior DevOps & Site Reliability Engineer

DeARX • Sandton

On-site
ZAR 1,000,000 - 1,400,000
Senior DevOps & Site Reliability Engineer at Datonomy Solutions
Senior DevOps & Site Reliability Engineer at Datonomy Solutions

Datonomy Solutions • Emfuleni Local Municipality

On-site
ZAR 900,000 - 1,500,000
Senior DevOps & Site Reliability Engineer
Senior DevOps & Site Reliability Engineer

Indsafri • City of Johannesburg Metropolitan Municipality

On-site
ZAR 1,200,000 - 1,800,000
Senior DevOps Engineer
Senior DevOps Engineer

ATS Client • Johannesburg

On-site
ZAR 900,000 - 1,300,000
Senior DevOps Engineer
Senior DevOps Engineer

Boardroom Appointments • Cape Town

On-site
ZAR 600,000 - 900,000
Cloud / SRE Platform Engineer
Cloud / SRE Platform Engineer

Syspro • Johannesburg

Hybrid
ZAR 600,000 - 1,000,000
25 days annual leave
30 days paid sick leave over 3-year
Hybrid working environment
Biz Dev Ops Engineer II
Biz Dev Ops Engineer II

Ovations Technologies • Johannesburg

Hybrid
ZAR 900,000 - 1,300,000
Senior DevOps Engineer
Senior DevOps Engineer

Blue Pearl HQ • Johannesburg

On-site
ZAR 600,000 - 850,000