Site Reliability & Production Support Engineer

Lesaka Technologies

Gauteng

Presencial

ZAR 600.000 - 900.000

Jornada completa

Hace 9 días
Generador de candidaturas

No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

Lesaka Technologies seeks a Site Reliability & Production Support Engineer to maintain the availability, performance, and resilience of production systems in South Africa. You will own incidents, investigate root causes, and coordinate with teams and providers to resolve issues.

You will enhance monitoring, automate support tasks, and contribute to robust disaster recovery and observability practices across cloud and on-prem environments.

Formación

  • Bachelor's degree or diploma in Computer Science, Information Technology, Software Engineering, Engineering, or a related technical field.
  • 5+ years' relevant experience in SRE, production engineering, DevOps, or similar production-focused role.
  • Experience investigating and resolving production incidents.
  • Ability to troubleshoot applications, databases, OS, networks, and cloud infrastructure.
  • Experience supporting cloud-hosted and on-prem workloads.
  • SQL skills including connectivity, performance, locking, and connection pooling issues.
  • Experience with monitoring, logging, alerting, and distributed tracing tools.
  • Proficiency in at least one scripting or programming language (Python, Bash, PowerShell, C#, or Go).
  • Ability to read application code and work with developers on fixes.
  • Familiarity with CI/CD, version control, infrastructure as code, and safe deployments.
  • Clear written and verbal communication, especially during incidents.
  • Ability to take ownership and prioritise under pressure.

Responsabilidades

  • Production support and incident resolution: supporting production apps, APIs, databases, infrastructure, and integrations.
  • Own incidents from detection through investigation, restoration, and closure.
  • Prioritise incidents by severity and business impact.
  • Investigate issues using logs, metrics, traces, queries, and diagnostics.
  • Implement changes within your area of responsibility and coordinate with teams/external providers.
  • Carry out controlled rollbacks, failovers, and recovery steps.
  • Maintain incident records and communicate progress to stakeholders.
  • Participate in on-call rotation, including outside business hours.

Conocimientos

Incident investigation
Cloud & on-prem workloads
SQL skills
Monitoring & tracing tools
Scripting (Python)
CI/CD & IaC
Ownership & problem solving
Read application code
Version control

Educación

Bachelor's degree or diploma in Computer Science / IT / Software Engineering or related field

Herramientas

Docker
Kubernetes
OpenTelemetry
Prometheus
Grafana
Terraform

Descripción del empleo

Lesaka EasyPay provides accessible financial services to consumers across South Africa, including banking, lending, insurance, payments and value-added services. We use technology to make everyday financial services simpler, more accessible and more convenient for our customers.

We are looking for a Site Reliability & Production Support Engineer to help ensure the availability, performance and resilience of the systems that support our customers and our business.

The Opportunity

We are looking for a Site Reliability & Production Support Engineer to maintain and improve the availability, performance, and resilience of our production systems.

You will own production incidents from investigation through to resolution, implementing fixes yourself where possible and coordinating with other teams or external providers where needed.

You will also address root causes, prevent recurring issues, improve monitoring, and automate support and recovery tasks.

Key Responsibilities

Production support and incident resolution

  • Support production applications, APIs, databases, infrastructure, and integrations.
  • Own incidents from detection and triage through investigation, service restoration, and closure.
  • Prioritise incidents based on severity, business impact, and affected services.
  • Investigate issues using logs, metrics, traces, database queries, application code, and infrastructure diagnostics.
  • Implement configuration, script, infrastructure, and application code changes within your area of responsibility.
  • Carry out controlled rollbacks, failovers, restarts, and message reprocessing with appropriate safeguards.
  • Coordinate engineering teams and external providers where their expertise or access is needed, and drive issues through to resolution.
  • Keep incident records and communicate impact, progress, and recovery status to stakeholders.
  • Participate in an agreed on-call rotation, including support for critical incidents outside business hours.

Root cause analysis and prevention

  • Conduct root cause investigations and blameless incident reviews.
  • Identify causes and contributing factors across applications, infrastructure, processes, and dependencies.
  • Document incident timelines, recovery actions, findings, and preventative measures.
  • Assign owners to corrective actions, track completion, and verify that fixes address the underlying problem.
  • Identify recurring incidents and support requests, and implement changes to prevent them.

Reliability and observability

  • Define and monitor service level indicators and objectives with engineering and business stakeholders.
  • Build and maintain dashboards, alerts, centralised logging, and distributed tracing.
  • Monitor system health, performance, dependencies, and customer impact.
  • Reduce alert noise and improve failure detection.
  • Address reliability risks, performance bottlenecks, capacity constraints, and single points of failure.
  • Define and support resilience tests, disaster recovery exercises, and backup and restore testing.

Automation and operational improvement

  • Automate repetitive support tasks, diagnostics, health checks, and recovery procedures.
  • Maintain operational tools, scripts, runbooks, and troubleshooting guides.
  • Contribute to infrastructure as code, CI/CD pipelines, and deployment safeguards.
  • Ensure services have the monitoring, health checks, rollback plans, and documentation needed for production support.
  • Work with developers to improve timeout handling, retries, idempotency, and graceful failure.
  • Follow production access controls, change management, and audit requirements.
Required Experience and Skills
  • A Bachelor's degree or diploma in Computer Science, Information Technology, Software Engineering, Engineering, or a related technical field.
  • 5+ years' relevant experience in SRE, production engineering, DevOps, technical application support, or a similar production-focused technical role.
  • Experience investigating and resolving production incidents.
  • Ability to troubleshoot applications, databases, operating systems, networks, and cloud infrastructure.
  • Experience supporting cloud-hosted and on-premises workloads.
  • Practical SQL skills, including diagnosing connectivity, query performance, locking, and connection pooling issues.
  • Experience with monitoring, logging, alerting, and distributed tracing tools.
  • Proficiency in at least one scripting or programming language, such as Python, Bash, PowerShell, C#, or Go.
  • Ability to read application code, investigate defects, and work with developers on fixes.
  • Familiarity with CI/CD, version control, infrastructure as code, and safe production deployments.
  • Clear written and verbal communication, especially during incidents.
  • Ability to take ownership, solve problems systematically, and prioritise under pressure.

Practical production experience is important to us. While 5+ years' relevant experience is preferred, we will consider the depth and relevance of hands-on experience alongside formal qualifications and years of experience.

Advantageous Experience
  • Microsoft Azure.
  • One or more of .NET, Laravel, Go, Python, or Rust.
  • Container platforms and orchestration tools, such as Docker and Kubernetes.
  • OpenTelemetry, Prometheus, Grafana, or similar tools.
  • Terraform and Git-based delivery pipelines.
  • High-volume transactional systems, distributed architectures, or customer-facing financial services platforms.
  • Disaster recovery, high availability, and multi-region architectures.
What Success Looks Like
  • Faster incident detection and service restoration, with fewer recurring incidents.
  • Clear ownership of incidents and corrective actions through to completion.
  • Permanent fixes that are verified to address root causes.
  • Measurable reliability objectives that inform engineering priorities.
  • Useful dashboards, actionable alerts, and up-to-date runbooks.
  • Less manual support work through automation and safe, automated recovery.
  • Timely, accurate stakeholder communication during incidents.
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Site Reliability & Production Support Engineer
Site Reliability & Production Support Engineer

Lesaka Technologies • Johannesburg

Presencial
ZAR 800.000 - 1.200.000
Site Reliability Engineer: Production & Incident Response
Site Reliability Engineer: Production & Incident Response

Lesaka Technologies • Johannesburg

Presencial
ZAR 800.000 - 1.200.000
Senior DevOps and Site Reliability Engineer (SRE)
Senior DevOps and Site Reliability Engineer (SRE)

Indsafri • Sudáfrica

Presencial
ZAR 900.000 - 1.500.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Placements24 • LegKraal Gate

Híbrido
ZAR 1.469.000 - 2.285.000
Home office stipend
Health, dental, and vision insurance
Remote work flexibility
+2
SRE: Production Uptime & Resilience Specialist
SRE: Production Uptime & Resilience Specialist

Lesaka Technologies • Gauteng

Presencial
ZAR 600.000 - 900.000
Senior DevOps Engineer
Senior DevOps Engineer

Blue Pearl PTY • Johannesburg

Presencial
ZAR 700.000 - 900.000
Site Reliability Engineer
Site Reliability Engineer

Nanolabs Health Services • Sudáfrica

Presencial
ZAR 500.000 - 700.000
Technical Support Engineer
Technical Support Engineer

Network International • Sudáfrica

Presencial
ZAR 480.000 - 720.000
Biz Dev Ops Engineer (Contract)
Biz Dev Ops Engineer (Contract)

The Focus Group • Sandton

Presencial
ZAR 1.000.000 - 1.800.000
Senior Operations Engineer
Senior Operations Engineer

Sabenza IT & Recruitment • Pretoria

Presencial
ZAR 800.000 - 1.200.000