DevOps Support Engineer

Connect

Johannesburg

On-site

ZAR 500,000 - 900,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Connect is seeking a DevOps Support Engineer to provide hands-on operational support for cloud-native and containerised solutions in modern cloud, Docker, and Kubernetes environments. You will troubleshoot issues, monitor platform health, support live services, and help keep production systems stable and observable.

You’ll work with engineering and platform teams to investigate incidents, respond to alerts, analyze logs and metrics, support deployments, maintain runbooks, and escalate as needed

Qualifications

  • Must have hands-on experience troubleshooting production or near-production cloud-based and containerised systems.
  • Experience supporting Docker containers, Docker images, and Kubernetes workloads, ideally on EKS.
  • Working knowledge of AWS infrastructure and operational support patterns.
  • Ability to investigate issues using logs, metrics, traces, dashboards, and alerting tools.
  • Experience supporting CI/CD pipelines using GitHub CI and/or Jenkins; Docker image build and deployment issues.

Responsibilities

  • Troubleshoot issues across cloud infrastructure, Docker containers, Kubernetes workloads, CI/CD processes, networking, and platform services.
  • Investigate production alerts, service degradation, failed deployments, and environment issues.
  • Use logs, metrics, traces, dashboards, and evidence to identify causes and support resolution.
  • Escalate complex issues with clear analysis, evidence, and impact summaries.
  • Monitor platform health using Prometheus, Grafana, VictoriaMetrics, Elasticsearch/Kibana, Sentry, and OpenTelemetry.
  • Support incident response activities for production and non-production environments.
  • Maintain runbooks and knowledge articles; contribute to post-incident reviews.

Skills

Problem solving
Incident management
Documentation
Communication
Troubleshooting

Tools

Docker
Kubernetes
AWS
GitHub CI
Jenkins
Terraform
Ansible
Packer
ArgoCD
Helm
Prometheus
Grafana
VictoriaMetrics
Elasticsearch
Kibana
Sentry
OpenTelemetry
Postgres
Kafka

Job description

Role purpose:

We’re looking for a DevOps Support Engineer to provide hands‑on operational support for cloud‑native and containerised solutions running across modern cloud, Docker, and Kubernetes environments. This role is primarily focused on troubleshooting issues, monitoring platform health, supporting live services, and helping ensure production systems remain stable, observable, and reliable.

You’re someone who enjoys solving operational problems, investigating alerts, and keeping services

stable. You are comfortable working through unclear issues methodically, including Docker image, container, deployment, and Kubernetes workload problems, using monitoring and diagnostic data to narrow down causes. You communicate clearly with engineers and stakeholders during support activities and value good runbooks, accurate documentation, and practical improvements that reduce repeated incidents.

You’ll work closely with engineering, platform, and support teams to investigate incidents, respond to alerts, analyse logs and metrics, support Docker and Kubernetes deployments, maintain runbooks, and escalates issues where deeper engineering input is required.

Main duties and key responsibilities:
Production Support & Troubleshooting
  • Troubleshoot issues across cloud infrastructure, Docker containers, Kubernetes workloads, CI/CD
    processes, networking, and platform services
  • Investigate production alerts, service degradation, failed Docker/Kubernetes deployments, and
    environment issues
  • Use logs, metrics, traces, dashboards, and system evidence to identify likely causes and support
    resolution
  • Escalate complex issues to engineering or platform teams with clear analysis, evidence, and impact
    summary
  • Track recurring issues and contribute to operational improvements, runbooks, and knowledge articles
Monitoring & Platform Health
  • Monitor platform health using Prometheus, Grafana, VictoriaMetrics, Elasticsearch / Kibana,
    Sentry, and OpenTelemetry
  • Review alerts, dashboards, logs, container events, and telemetry to detect service issues and
    infrastructure risks
  • Support alert triage, noise reduction, and monitoring coverage improvements for Docker and
    Kubernetes workloads
  • Validate that services, containers, deployments, jobs, and platform components are operating
    as expected
Incident Response & Service Recovery
  • Support incident response activities for production and non-production environments
  • Assist with initial triage, impact assessment, workaround identification, and service recovery
  • Capture incident timelines, evidence, symptoms, and actions taken during support activities
  • Contribute to post-incident reviews by identifying recurring patterns, gaps in monitoring, and
    opportunities to improve support processes
Deployment, Pipeline & Environment Support
  • Support GitHub CI and Jenkins pipeline issues, failed jobs, deployment errors, Docker image
    build failures, and release workflow problems
  • Help troubleshoot Docker, Helm, ArgoCD, Kubernetes, and environment configuration issues
  • Support infrastructure provisioning and configuration issues involving Terraform, Ansible, and
    Packer
  • Assist engineering teams with Docker deployment validation, container runtime checks,
    rollback support, and environment readiness checks
Operational Maintenance & Continuous Improvement
  • Maintain and improve runbooks, support procedures, Docker deployment troubleshooting
    guides, and operational documentation
  • Identify repeated manual support activities that could be automated or simplified
  • Support routine maintenance, health checks, capacity checks, container image checks, and
    service validation activities
  • Help improve reliability, observability, and supportability of cloud, Docker, and Kubernetes
    based solutions
Stakeholder & Engineering Support
  • Provide technical support to engineering and operations teams for platform, Docker
    deployment, container, and environment issues
  • Communicate clearly during incidents, handovers, escalations, and support updates
  • Work with engineers to validate fixes, confirm service recovery, and close support actions
  • Maintain knowledge articles and known-issue documentation to improve future support
    efficiency
Required Skills & Experience
  • Experience troubleshooting production or near-production cloud-based and containerised
    systems
  • Hands-on experience supporting Docker containers, Docker images, and Kubernetes
    workloads, ideally on EKS
  • Working knowledge of AWS infrastructure and common operational support patterns
  • Ability to investigate issues using logs, metrics, traces, dashboards, and alerting tools
  • Experience supporting CI/CD pipelines using GitHub CI and/or Jenkins, including Docker
    image build and deployment issues
  • Working knowledge of Terraform, Ansible, Packer, Docker, Helm, and ArgoCD from a support
    and troubleshooting perspective
  • Familiarity with observability tools such as Prometheus, Grafana, VictoriaMetrics,
    Elasticsearch / Kibana, Sentry, and OpenTelemetry
  • Understanding of Linux, networking, VPNs, service isolation, container networking, and Zero
    Trust concepts
  • Ability to document incidents, known issues, Docker deployment procedures, runbooks, and
    support procedures clearly
Desirable Experience
  • Experience working in an application, infrastructure, platform, or DevOps support
    environment
  • Exposure to incident management, service recovery, or operational support processes
  • Experience supporting stateful services such as Postgres and Kafka
  • Familiarity with load testing outputs, performance symptoms, and capacity-related issues
  • Experience maintaining Docker registries, Docker image repositories, artifact repositories, or
    internal developer tooling
  • Exposure to regulated environments, compliance checks, or audit-support activities
  • Ability to identify recurring support issues and suggest practical improvements
  • Example Tech Stack Exposure
Example Tech Stack Exposure
  • AWS Docker Kubernetes Terraform Ansible Packer GitHub CI Jenkins ArgoCD Helm Prometheus Grafana VictoriaMetrics Elasticsearch Kafka Postgres OpenTelemetry Zero Trust Networking
Typical technologies in this environment include:

AWS Docker Kubernetes Terraform Ansible Packer GitHub CI Jenkins ArgoCD Helm Prometheus Grafana VictoriaMetrics Elasticsearch Kafka Postgres OpenTelemetry Zero Trust Networking

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps, Platform & Support Engineer
DevOps, Platform & Support Engineer

MVIA • Wes-Kaap

On-site
ZAR 600,000 - 900,000
DevOps Lead
DevOps Lead

Boardroom Appointments • Cape Town

On-site
ZAR 500,000 - 700,000
DevOps Engineer
DevOps Engineer

Placements24 • East London

Hybrid
ZAR 600,000 - 900,000
Competitive salary and benefits
Intermediate Platform Engineer
Intermediate Platform Engineer

AiR • Cape Town

On-site
ZAR 600,000 - 900,000
Senior Platform Engineer
Senior Platform Engineer

Boardroom Appointments • Cape Town

On-site
ZAR 700,000 - 900,000
Senior DevOps Engineer | Cloud, Kubernetes & Automation
Senior DevOps Engineer | Cloud, Kubernetes & Automation

Network Finance • Randburg

Hybrid
ZAR 700,000 - 1,100,000
Cloud & Kubernetes Ops Support Engineer
Cloud & Kubernetes Ops Support Engineer

Connect • Johannesburg

On-site
ZAR 500,000 - 900,000
Biz Dev Ops Engineer (Contract)
Biz Dev Ops Engineer (Contract)

The Focus Group • Sandton

On-site
ZAR 1,000,000 - 1,800,000
DevSecOps / Cloud Engineer
DevSecOps / Cloud Engineer

Indsafri • South Africa

Hybrid
ZAR 1,200,000 - 1,800,000
Platform Engineer (AWS, GitHub Actions, Heroku CI) (JHB)
Platform Engineer (AWS, GitHub Actions, Heroku CI) (JHB)

Datafin • Johannesburg

On-site
ZAR 600,000 - 800,000