Software Engineer - Reliability/SRE

Electrum

Cape Town

On-site

ZAR 900,000 - 1,200,000

Full time

28 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible Work
Generous Leave
Cape Town Office Perks

Job summary

Electrum is seeking a Software Engineer focused on reliability to join our Cape Town team. You will build and ship software that keeps our payments platforms healthy, scalable, and secure, collaborating across engineering to design for resilience.

In this role you’ll participate in on-call rotations, help define incident response and monitoring practices, and contribute to a culture of observability, automation, and continuous improvement within a fast-growing payments company.

Qualifications

  • Bachelor's degree in Computer Science or Information Technology.
  • 3+ years building and operating production systems (SRE/DevOps/Platforms) with ownership.
  • Experience with AWS, observability tools, and CI/CD pipelines.

Responsibilities

  • Contribute to codebases and shared platforms with the same standards of review, testing, and delivery as the wider engineering team
  • Design, build, and maintain software, tooling, and automation that improve the reliability, availability, and scalability of our applications and services
  • Work closely with other engineers to understand, address, and prevent technical issues in the systems we own and depend on
  • Develop and maintain incident response processes and alerting mechanisms
  • Build and maintain the tooling that measures our SLIs and SLOs, and make the results visible to the teams that own them
  • Participate in on‑call rotations to provide 24/7 operational support as necessary

Skills

SRE/DevOps
Cloud AWS
Observability
Terraform
Kubernetes
CI/CD
Troubleshooting
Teamwork

Education

Bachelor's degree in Computer Science or IT

Tools

DataDog
ELK Stack
Grafana

Job description

Electrum Is a Next-generation Payment Software Technology Company.

Since 2012, we've delivered trusted, enterprise-grade, cloud-native software to optimise financial transaction processing. Our deep expertise has established us as a respected partner in high-volume, low-value payment schemes, enabling clients to deliver services to millions of South Africans daily.

At Electrum, we are grounded in impact - designing solutions that matter, acting with urgency, and continuously learning as we scale. We believe in creating together - working side by side with our clients and teams to build meaningful, lasting solutions. We prioritise making it safe - encouraging open communication, smart risk-taking, and trust so that creativity and alignment thrive. And we back empowered strong teams - hiring brilliant people, collaborating hard, and holding each other to high standards while leading with empathy and kindness.

The Role

As a Software Engineer with a focus on reliability, you'll write and ship software that keeps our payments systems healthy, scalable, and secure. You'll work as part of the engineering organisation: building tooling and services, improving how we deploy and operate what we ship, and solving hard problems in production with the same engineering craft you bring to code. You'll collaborate closely with other development teams to design for scale, address issues before they become incidents, and raise the bar on how we observe and run our systems. You'll take part in on‑call rotations, help manage critical incidents, and contribute to response processes so we resolve issues quickly and learn from them. You'll also help shape decisions around security, system optimisation, rollout strategies, and monitoring for health and availability.

Responsibilities
Software Engineering for Reliability
  • Contribute to codebases and shared platforms with the same standards of review, testing, and delivery as the wider engineering team
  • Design, build, and maintain software, tooling, and automation that improve the reliability, availability, and scalability of our applications and services
  • Work closely with other engineers to understand, address, and prevent technical issues in the systems we own and depend on
  • Develop and maintain incident response processes and alerting mechanisms
  • Build and maintain the tooling that measures our SLIs and SLOs, and make the results visible to the teams that own them
System Troubleshooting and Problem Resolution
  • Diagnose and resolve infrastructure and system-level issues, ensuring minimal downtime and swift problem resolution
  • Respond to and investigate incidents related to infrastructure and applications, utilising diagnostic tools to track down and remediate issues
  • Participate in on‑call rotations to provide 24/7 operational support as necessary
  • Develop and maintain incident response processes, alerting, and runbooks, and drive the preventative actions that come out of post-incidents
Observability and Automation
  • Utilise technologies to develop and maintain effective log management and monitoring solutions for internal and external customers
  • Evaluate system health, identify performance bottlenecks and proactively optimise performance and cost‑effectiveness
  • Implement automation tools and frameworks for deployment, configuration, and monitoring processes
  • Capacity management and planning for systems to ensure continued reliability
Continuous Improvement
  • Evaluate and integrate emerging technologies, cloud services and automation tools to improve operational efficiency
  • Drive cost‑optimisation initiatives by identifying opportunities for resource right‑sizing, efficiency and other cost‑saving measures
  • Design and test disaster recovery strategies, including backup and restoration, and run or facilitate DR exercises
Requirements
  • Bachelor's degree in Computer Science, Information Technology
  • 3+ years building and operating production systems - SRE, DevOps, Platforms, or a software engineering role with real operational ownership
  • Working knowledge of cloud services across compute, object storage, databases, serverless, and monitoring (AWS preferred)
  • Demonstrable experience with observability tooling and pipelines, e.g. DataDog, Elastic/ELK Stack or Grafana
  • Infrastructure-as-code and CI/CD experience (Terraform, or equivalent)
  • Experience with containers and orchestration (e.g. Kubernetes)
  • Strong troubleshooting instincts under pressure, and the judgement to know when to stop investigating and start mitigating
  • Attention to detail and ability to work effectively in a team environment
Benefits
Why Join Electrum?
  • We believe in a People First approach, ensuring a culture where you can thrive and make a real difference
Your Career & Culture
  • Career Growth: Delivering world‑class financial software is challenging, but your effort will earn you hands‑on experience with products used by millions, accelerating your career.
  • Strong Teams: We keep teams small, focused, and collaborative to maximize impact
  • Transparency: We openly discuss strategy, finances, and salaries. Mistakes are viewed as learning opportunities that we actively discuss
  • Autonomy: We trust you. You're expected to seek out the data needed for informed decisions and manage your own time—knowing when to focus and when to recharge
  • Shared Vision: You'll have the power to shape the vision of how we build the future of financial services
Practical Perks
  • Here's how we support our culture:
  • Flexible Work: Office‑first environment with flexible hours
  • Generous Leave: Starting at 20 days per year.
  • Office Perks (Cape Town): Fully‑stocked kitchen and daily catered lunch
  • Social Life: Regular team activities like hikes, getaways, and dinners
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - Reliability/SRE
Software Engineer - Reliability/SRE

Electrum Software • Cape Town

On-site
ZAR 600,000 - 900,000
Flexible Work hours
Generous leave: 20 days/year
Cape Town office perks: daily catered/
+1
Software Engineer - Reliability/SRE
Software Engineer - Reliability/SRE

Electrum Payments • Cape Town

On-site
ZAR 600,000 - 900,000
Flexible work hours
20 days leave per year
Cape Town office lunch
+1
Software Engineer - Platform Engineer
Software Engineer - Platform Engineer

Electrum • Cape Town

On-site
ZAR 900,000 - 1,300,000
Flexible Work
20 days leave
Cape Town office lunch
+1
Product Support Engineer - Junior
Product Support Engineer - Junior

Electrum • Cape Town

On-site
ZAR 480,000 - 720,000
Generous Leave: Starting at 20 days
Office Perks (Cape Town): Fully-stocke
Social Life: Regular team activities
Software Developer - Java - Senior
Software Developer - Java - Senior

ATS Client • Cape Town

On-site
ZAR 900,000 - 1,500,000
Flexible Work
Generous Leave
Office Perks (Cape Town)
+1
Application Security Engineer - Intermediate
Application Security Engineer - Intermediate

Electrum • Cape Town

On-site
ZAR 600,000 - 800,000
Flexible work hours
Generous leave starting at 20 days per year
Office perks (fully-stocked kitchen and daily catered lunch)
+1
Software Engineer - Java - Platform Track
Software Engineer - Java - Platform Track

Electrum • Cape Town

On-site
ZAR 600,000 - 900,000
Flexible Work
Generous Leave
Cape Town Office
+1
Intermediate Front-end Software Engineer
Intermediate Front-end Software Engineer

Electrum Software • Cape Town

On-site
ZAR 600,000 - 900,000
Flexible hours
Cape Town office lunch
Office kitchen & meals
Intermediate Front-end Software Engineer
Intermediate Front-end Software Engineer

Electrum • Cape Town

On-site
ZAR 650,000 - 1,100,000
Flexible hours
20 days leave
Daily catered lunch
+1
Software Developer Team Lead - Cloud
Software Developer Team Lead - Cloud

Electrum Payments • Cape Town

On-site
ZAR 900,000 - 1,500,000
Flexible Work Hours
Generous Leave
Cape Town Office Perks
+2