SRE Platform Engineer (Microsoft Azure)

Syspro

Johannesburg

Hybrid

ZAR 650,000 - 700,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Hybrid work environment
Pension 10%
Medical aid
Life insurance
Income continuation protection
Bonus/commission programmes
Gym Facilities
Free Daily Lunch
Free Fruit & Snacks

Job summary

Syspro is seeking an SRE Platform Engineer to design, build, and maintain a reliable, scalable cloud platform infrastructure. You will apply SRE principles with modern AIOps tooling to automate workflows and partner with development teams to embed reliability into the software lifecycle.

The role requires 4–6 years in platform engineering/DevOps/SRE, hands-on Kubernetes and CI/CD experience, and strong scripting skills.

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, or a related field; or equivalent NQF Level 7 qualification.
  • 4–6 years in platform engineering/DevOps/SRE roles with hands-on Kubernetes, CI/CD pipelines and cloud infra management.

Responsibilities

  • Design scalable, reliable platform infra using SRE principles (SLOs, error budgets, reliability targets).
  • Automate workflows with IaC and CI/CD to reduce toil and speed deployments.
  • Monitor platform health with observability tools; drive RCA and prevent recurrence.
  • Maintain and improve Kubernetes-based environments for availability and security of workloads.
  • Apply AI-assisted monitoring/AIOps to identify risks and reduce MTTR.
  • Collaborate with development teams and Architects to embed reliability early in design.
  • Participate in on-call rotations and post-incident reviews to strengthen resilience.
  • Contribute to capacity planning, DR design, and runbooks.

Skills

Kubernetes
CI/CD pipelines
IaC
Observability
Python/Bash
Incident response
Collaboration

Education

Bachelor’s degree in Computer Science/IT or related (NQF7)

Tools

Kubernetes
Terraform
Ansible
Prometheus
Grafana
Datadog

Job description

Job Description

The SRE Platform Engineer is responsible for designing, building, and maintaining the reliability, availability, and scalability of Syspro’s cloud platform infrastructure. Applying site reliability engineering principles alongside modern AIOps tooling, the role ensures platform resilience, automates operational workflows, and partners with development teams to embed reliability into the software delivery lifecycle.

Job Requirements

Minimum Level of Education

  • Bachelor’s Degree in Computer Science, Information Technology, or a related field; or equivalent NQF Level 7 qualification.

Beneficial Qualifications

Certified Kubernetes Administrator (CKA); AWS, GCP, or Azure cloud certifications; Google SRE or DORA-related training advantageous.

Minimum Experience

  • 4–6 years in a platform engineering, DevOps, or site reliability engineering role. Demonstrated hands‑on experience with Kubernetes, CI/CD pipelines, and cloud infrastructure management. Proven ability to own and resolve production incidents independently.

Beneficial Experience

  • Experience with multitenancy cloud platform environments. Exposure to AIOps or AI assisted observability platforms. Terraform or Ansible infrastructure‑as‑code experience. Background in B2B SaaS or ERP software environments.

Special Skills and Knowledge

  • Site reliability engineering: working knowledge of SLOs, SLIs, error budgets, and reliability engineering practices.
  • Kubernetes and container orchestration: hands‑on experience managing Kubernetes clusters, workloads, and networking configurations.
  • CI/CD and IaC: proficiency with pipeline tooling (GitLab CI, GitHub Actions, Jenkins) and infrastructure‑as‑code (Terraform, Helm, Ansible).
  • Observability and monitoring: competency with tools such as Prometheus, Grafana, Datadog, or equivalent platforms.
  • AI and AIOps tooling: ability to apply AI‑assisted monitoring and anomaly detection tools to improve platform health and reduce MTTR.
  • Solid scripting languages (Python, Bash).
  • Incident response: disciplined approach to on‑call responsibilities, structured incident management, and blameless post‑mortems.
  • Collaboration and communication: ability to work effectively with development teams, architects, and operations leadership.

Self-driven learning: commitment to staying current with evolving SRE practices, cloud‑native tooling, and platform technologies.

Job Responsibilities
  • Design and implement scalable, reliable platform infrastructure using SRE principles — including service‑level objectives (SLOs), error budgets, and reliability targets — to drive consistent uptime and performance.
  • Automate operational workflows through infrastructure‑as‑code (IaC) tooling and CI/CD pipeline configuration, eliminating manual toil and accelerating deployment velocity.
  • Monitor and observe platform health across services using observability tooling (Prometheus, Grafana, or Datadog equivalents), responding to incidents and driving root‑cause analysis to prevent recurrence.
  • Maintain and improve Kubernetes‑based container orchestration environments, ensuring optimal resource allocation, availability, and security posture across all workloads.
  • Leverage AI‑assisted monitoring, AIOps platforms, and ML‑based anomaly detection tools to proactively identify platform risks, reduce mean time to resolution (MTTR), and surface predictive insights for the team.
  • Collaborate with development teams and the Senior Kubernetes Architect to embed reliability requirements into system design from the outset, participating in design reviews and pre‑production readiness assessments.
  • Participate in on‑call rotations and incident response, contributing to post‑incident reviews and implementing corrective actions to strengthen platform resilience.

Contribute to capacity planning, disaster recovery design, and business continuity efforts, maintaining up‑to‑date runbooks and operational documentation.

Job Benefits
  • 25 annual leave days
  • 30 days paid sick leave over 3-year cycle
  • Hybrid working environment (3 office days as determined by the function/ manager)
  • Pension 10%
  • Medical aid
  • Life insurance
  • Income continuation protection
  • Funeral Benefit
  • Provident Fund Admin Fee
  • Global Education Protection
  • Bonus/ commission programmes
  • Maternity: 6 months at half pay
  • Paternity: 2 weeks full pay
  • Employee Assistance Programme
  • Free Barista Coffee Daily
  • Gym Facilities
  • Free Daily Lunch
  • Free Fruit on a Tuesday & Thursday
  • Free Refreshments & Snacks Daily
Salary

650000 - 700000 ZAR (yearly)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Platform Engineer (Microsoft Azure)
SRE Platform Engineer (Microsoft Azure)

Syspro • Sandton

Hybrid
ZAR 900,000 - 1,500,000
Annual leave 25 days
Sick leave 30 days
Hybrid work model
+16
Senior Kubernetes Architect
Senior Kubernetes Architect

Syspro • Johannesburg

Hybrid
ZAR 900,000 - 1,200,000
Hybrid working environment
Pension
Medical aid
+3
Junior Kubernetes Architect (Microsoft Azure)
Junior Kubernetes Architect (Microsoft Azure)

Syspro • Johannesburg

Hybrid
ZAR 348,000 - 525,000
Hybrid work
Pension
Medical aid
+5
Junior Kubernetes Architect (Microsoft Azure)
Junior Kubernetes Architect (Microsoft Azure)

Syspro • Sandton

Hybrid
ZAR 420,000 - 640,000
25 annual leave days
30 days paid sick leave over 3-year
Hybrid working environment
+16
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Placements24 • LegKraal Gate

Hybrid
ZAR 1,469,000 - 2,285,000
Home office stipend
Health, dental, and vision insurance
Remote work flexibility
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Sana Commerce • Cape Town

Hybrid
ZAR 700,000 - 900,000
Impact from day one
Continuous learning and mentorship
Hybrid flexibility
Senior DevOps and Site Reliability Engineer (SRE)
Senior DevOps and Site Reliability Engineer (SRE)

Indsafri • South Africa

On-site
ZAR 900,000 - 1,500,000
DevOps / SRE Cloud Engineer (Site Reliability Engineer)
DevOps / SRE Cloud Engineer (Site Reliability Engineer)

Recruit-It • South Africa

On-site
ZAR 900,000 - 1,500,000
Hybrid work
Home office stipend
Learning budget
+2
DevOps / SRE Cloud Engineer (Site Reliability Engineer) - Remote
DevOps / SRE Cloud Engineer (Site Reliability Engineer) - Remote

Recruit-It • South Africa

Remote
ZAR 900,000 - 1,500,000
Flexible hours
Hybrid/Remote options
Wellness program
+3
Software Engineer - Intermediate: Product Engineering
Software Engineer - Intermediate: Product Engineering

Syspro • Sandton

Hybrid
ZAR 700,000 - 1,100,000
25 annual leave days
Sick leave 30 days
Hybrid working environment
+15