Site Reliability Engineer

ATS Client

Cape Town

On-site

ZAR 900,000 - 1,300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible Working
Health and Wellness
RSUs
Coursera access
Parental leave
Tech stipend

Job summary

Impact.com is seeking a Site Reliability Engineer to own and grow the reliability practice for content indexing, monitoring, and enrichment systems. You will implement SRE disciplines, define SLOs/RCAs, and partner with Cloud Ops and security teams to ensure secure, auditable, and reliable services on GCP.

You will read and write Java/Spring, tune JVM, harden infrastructure, and improve batch processes. This role emphasizes ownership and long-term growth within the Content Intelligence &

Qualifications

  • 3+ years in SRE, software engineering, or production systems roles.
  • Strong grasp of reliability, performance, and security practices.
  • Experience with Java/Spring and cloud environments (GCP preferred).
  • Hands-on with observability, metrics, tracing, and alerting.
  • Experience tuning MySQL/SQL and batch/workflow orchestration.

Responsibilities

  • Define and own SLO/SLI frameworks for services on GCP with squads.
  • Build security-by-default and auditability in production systems.
  • Drive vulnerability management and patching; surface security findings.
  • Lead blameless RCAs and durable fixes across teams.
  • Develop and mature observability, dashboards, tracing, and alerting.
  • Monitor performance and resources, guide optimization efforts.
  • Improve CI/CD, GitOps, and deployment safety for finance services.
  • Collaborate with Cloud Operations and Security to ensure compliance.
  • Analyze client workloads and ensure space for reliability improvements.
  • Establish runbook standards and reliable on-call practices.

Skills

SRE experience
Java/Spring
Cloud & Kubernetes
Observability
Database performance
Automation scripting
Security & compliance
SLO/SLI discipline
Collaboration
Problem solving
Education-focused

Education

B.S. in Computer Science or related field

Tools

Vault
GCP Secret Manager
SOPS
Jenkins
ArgoCD
Helm
Prometheus/Grafana

Job description

About impact.com

impact.com is the world's leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer journey. From affiliates and influencers to content publishers, brand ambassadors, and customer advocates, impact.com empowers brands to drive trusted, performance-based growth through authentic relationships. Its award‑winning products – Performance (affiliate), Creator (influencer), and Advocate (customer referral) – unify every type of partner into one integrated platform. As consumers increasingly rely on recommendations from people and communities they trust, impact.com helps brands show up where it matters most. Today, over 5,000 global brands, including Walmart, Uber, Shopify, Lenovo, L'Oréal, and Fanatics, rely on impact.com to power more than 225,000 partnerships that deliver measurable business results.

Your Role at impact.com:

As the Site Reliability Engineer for the Content Intelligence & Regulatory Apps Group, you will own and grow our reliability practice for the systems that index, monitor and enrich social and web content on the impact.com platform. This role focuses on building and professionalizing rather than firefighting. Since the platform is stable, your mission is to implement SRE engineering disciplines (service‑level objectives, observability, runbooks, and structured root‑cause analysis) to ensure reliability is measurable, repeatable, and owned.

You will work with Squad leads, Platform Engineering and Cloud Operations and will own the SLO and RCA practice for the group while partnering closely with other internal and external squads. Your job is to provide them with the framework, tooling, and habits needed to run reliable services, and to act as the central point of contact for reliability across those teams. This is a software‑engineering‑led SRE role where you will read and write Java, instrument Spring services, tune the JVM, help harden infrastructure, batch processes, data and orchestration flows.

Our guiding principle is to prioritize system stability and data integrity above all else. Because these are business‑critical systems, security and compliance are part of the reliability mandate, not an afterthought: you will build observability, audit trails, and operational practices that are secure and auditable by default, working alongside the central security and DevOps teams. Reliability, resilience, data integrity, and compliance take precedence over short‑term feature velocity. This is a strong opportunity for an engineer ready to step into ownership and grow the role and themselves over time, with the support of the group.

What You'll Do:
  • Own the SLO/SLI practice for CIRA: Define meaningful service‑level objectives and indicators for services on GCP with each squad. Establish error budgets and necessary baselines, applying extra rigor to flows that affect critical business processes.
  • Build in security and compliance by default. Treat security and auditability as reliability properties: ensure critical business processes and data flows have the necessary audit trails, and if applicable, forensic‑replay history needed for transactional‑correctness and compliance obligations (e.g. SOX, and PCI‑adjacent concerns). Champion secrets hygiene (HashiCorp Vault, GCP Secret Manager, SOPS), least‑privilege access to production and data, and audited break‑glass procedures. Partner with the Cloud Security and Cloud Platform teams rather than duplicating their function.
  • Manage vulnerability and patch posture for services: Track and drive remediation of vulnerabilities across the JVM, Spring/Spring Boot dependencies, and container images; help establish patching expectations and surface security‑relevant findings from quality gates (SonarQube) and secret scanning (ggshield) so they get prioritized alongside reliability work.
  • Own and run root‑cause analysis: Establish a consistent, blameless RCA practice for the group. Drive investigations toward durable fixes and preventative actions while partnering with the owning squad rather than working in isolation. Over time, improve the squads' own troubleshooting and post‑incident habits.
  • Build and mature observability for critical services: Become the group's point person for the observability stack (e.g., Grafana). Build and standardize monitoring, dashboards, tracing, and alerting using Open Telemetry that surface the health of systems, services, infrastructure and business/transactional processes – without leaking PII data into logs, traces, or dashboards – while creating reusable patterns that squads can adopt.
  • Keep performance and resource metrics within thresholds: Track response latency, JVM heap/GC behavior, thread‑pool saturation, CPU/memory consumption, cloud costs, error rates, and uptime across all services (e.g. Java, Spring Boot, NodeJs etc). Turn findings into prioritized improvements in collaboration with the squads.
  • Improve batch and orchestration reliability: Help harden batch jobs and support the migration toward durable workflow orchestration to reduce partial‑failure windows and manual idempotency.
  • Support database performance and query optimization: Monitor the groups databases (e.g. MySql, SingleStore, Elastic, GCP Spanner) for slow queries, indexing, lock contention, and connection‑pool health. Flag operational or performance risks in Liquibase migrations before they ship.
  • Understand client‑generated workloads: Learn how brand, agency, and partner workloads translate into resource consumption. Help confirm that consumption aligns with contract tiers and surface anomalous traffic patterns as both a reliability and a security signal, escalating suspected abuse to the security team.
  • Establish alerting and runbook standards: Define the group's conventions for actionable alerts and runbooks to ensure on‑call engineers across squads can respond quickly and consistently.
  • Troubleshoot across the stack: Address issues in the JVM/application, Spring container, MySQL, Pub/Sub messaging, Feign/REST and legacy RMI/HTTP‑Invoker integrations, GKE/GCE, network, and client‑generated workloads. This requires growing depth in JVM and database performance (profiling and query optimization). Remediation may include code/query optimizations, JVM/connection‑pool tuning, rate‑limit or auto‑scaling configuration, retry/backoff changes, or escalating heavy‑request patterns.
  • Make the delivery path safer: Partner with squads to improve CI/CD (Jenkins/Github Actions) and GitOps deploys (ArgoCD + Helm), including health checks, readiness/liveness probes, progressive rollouts, and rollbacks for finance services.
  • Inform capacity and cost: Help analyze platform usage to attribute costs to customers and workloads, inform capacity planning, and identify efficiency improvements. Partner with Cloud Operations (FinOps).
  • Leave things more reliable than you found them: Improve tests, observability, and operability each time a service is touched, and grow the group's reliability guardrails into a shared, durable practice.
  • General: Other duties as assigned by the Company. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.
What You Bring:
  • Experience: 3+ years in SRE, software engineering, or systems/operations roles supporting production services, with the appetite to own and grow a reliability practice.
  • Systems Design: Solid understanding of systems and application design, with the ability to reason about reliability, failure modes, and performance.
  • Java: Able to read, debug, and make changes to Java/Spring code today, with the willingness and aptitude to deepen JVM expertise (garbage collection, memory, thread pools) on the job.
  • Cloud & Kubernetes: Experience operating services on a major cloud (ideally GCP/GKE) and exposure to Kubernetes and containers.
  • Observability: Proficiency with metrics, logging, tracing/APM, and alerting.
  • Database: Experience with relational database monitoring and SQL tuning (MySQL preferred), including basic indexing and query optimization.
  • Automation: Proficiency in shell scripting and comfort automating routine operational tasks.
  • Security & Compliance Awareness: Practical understanding of operational security for production systems, including secrets management, least‑privilege access, and handling sensitive data safely in logs and telemetry.
  • SLO/SLI: Familiarity with defining or operating against SLOs/SLIs, or a strong desire and aptitude to build that practice.
  • Collaboration: A collaborative working style with the ability to influence and enable other engineers.
  • Problem Solving: Ability to prioritize, work independently, and focus on simple, efficient, and reliable solutions.
  • Education: B.S. in Computer Science or a related field, or equivalent practical experience.
Preferred Qualifications
  • Spring Ecosystem: Experience with the Spring/Spring Boot ecosystem and JVM application servers.
  • Batch/Workflow Orchestration: Experience with tools like Quartz, Temporal, or similar.
  • Event-Driven Messaging: Familiarity with Pub/Sub or Kafka and service-to-service integration (REST/Feign, gRPC).
  • Telemetry & Metrics Analysis: Experience with Prometheus/PromQL, Grafana, and log analytics.
  • AI-assisted engineering: Familiarity with using AI tools to assist in engineering workflows.
  • Tooling: Familiarity with secrets management (Vault, GCP Secret Manager, SOPS), code-quality gates (SonarQube), and CI/CD + GitOps (Jenkins, ArgoCD/Helm).
Benefits and Perks:

At impact.com, we believe that when you're happy and fulfilled, you do your best work. That's why we've built a benefits package that supports your well‑being, growth, and work‑life balance.

  • Flexible Working: Our Responsible PTO policy means you can take the time off you need to rest and recharge. We're committed to a positive work‑life balance and provide a flexible environment that allows you to be happy and fulfilled in both your career and your personal life.
  • Health and Wellness: Your well‑being is a priority. Our mental health and wellness benefit includes up to 12 fully covered therapy/coaching sessions per year, with additional dependent coverage. We also offer a monthly gym reimbursement policy to support your physical health.
  • A Stake in Our Growth: We offer Restricted Stock Units (RSUs) as part of our total compensation, giving you a stake in the company's growth with a 3‑year vesting schedule, pending Board approval.
  • Investing in Your Growth: We're committed to your continuous learning. Take advantage of our free Coursera subscription and our PXA courses.
  • Parental Support: We offer a generous parental leave policy, 26 weeks of fully paid leave for the primary caregiver and 13 weeks fully paid leave for the secondary caregiver.
  • Technology Financial Support: We provide a technology stipend to help you set up your home office and a monthly allowance to cover your internet expenses.

impact.com is proud to be an equal opportunity workplace.

All employees and applicants for employment shall be given fair treatment and equal employment opportunity regardless of their race, ethnicity or ancestry, color or caste, religion or belief, age, sex (including gender identity, gender reassignment, sexual orientation, pregnancy/maternity), national origin, weight, neurodivergence, disability, marital and civil partnership status, caregiving status, veteran status, genetic information, political affiliation, or other prohibited non‑merit factors.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Impact • Cape Town

On-site
ZAR 900,000 - 1,500,000
Flexible Working
Health and Wellness
RSU Stock
+3
Platform Engineer
Platform Engineer

impact.com • Cape Town

On-site
ZAR 600,000 - 900,000
Flexible PTO
Mental health support
RSUs with 3-year vesting
+4
Platform Engineer
Platform Engineer

Silversmith Capital Partners • Cape Town

On-site
ZAR 600,000 - 900,000
Flexible PTO
Mental health support
RSUs & vesting
+3
Associate Platform Infrastructure Engineer
Associate Platform Infrastructure Engineer

impact.com • Cape Town

On-site
ZAR 900,000 - 1,200,000
Flexible PTO
Mental health support
RSUs with 3-year vesting
+3
Senior Software Engineer
Senior Software Engineer

Silversmith Capital Partners • Cape Town

Hybrid
ZAR 600,000 - 900,000
Hybrid work environment
Medical Aid and Provident Fund
Internet Allowance
+1
Senior Software Engineer
Senior Software Engineer

impact.com • Wes-Kaap

On-site
ZAR 1,200,000 - 2,000,000
Flexible PTO
Health & Wellness
RSUs
+3
Senior Platform Engineer
Senior Platform Engineer

impact.com • Cape Town

On-site
ZAR 900,000 - 1,500,000
Flexible PTO
Mental health support
RSUs with 3-year vesting
+3
Associate Platform Infrastructure Engineer
Associate Platform Infrastructure Engineer

Silversmith Capital Partners • Cape Town

Hybrid
ZAR 600,000 - 1,000,000
Flexible Working
Health & Wellness
RSUs
+3
Software engineer
Software engineer

ATS Client • Cape Town

Hybrid
ZAR 900,000 - 1,300,000
Hybrid, Casual work environment
Unlimited PTO policy
Medical Aid and Provident Fund
+6
Senior QA Engineer
Senior QA Engineer

impact.com • Cape Town

On-site
ZAR 700,000 - 950,000
Flexible Working
Health and Wellness
RSUs
+3