Our clients reserves the right not to make an appointment. In considering candidates for appointment into advertised posts, preference will be accorded to persons from a designated group in accordance with the approved Employment Equity Plan.

SRE Engineering (AM/SRE/13/07/2026)

Overview

Reference
AM/SRE/13/07/2026

Salary
ZAR/annum

Job Location
South Africa -Johannesburg Metro -Johannesburg

Job Type
Contract

Posted
04 September 2026

Closing date
07 Sep 2026 06:17


Site Reliability Engineer

Location

Gauteng,Johannesburg

Job Type

Contract – Full-Time hours

Job Description

We are looking for an SRE who has worked with on Prem applications and built out SRE Dashboards such as AppMon and DynaTrace or used Splunk or Open Source tools to create these Dashboards. 

Background & Objectives:

  • The initiative aims to establish and strengthen Site Reliability Engineering capability across Trade Online and related platforms to:
  • Improve service availability, resilience, and production stability.
  • Reduce incident frequency, severity, and mean time to restore service.
  • Establish measurable reliability practices through SLIs, SLOs, error budgets, and operational dashboards.
  • Enable proactive monitoring, observability, automation, and continuous improvement of production services.

Scope of Services:

  • SRE support for critical Trade Finance products, Trade Online services, and related integration points.
  • Definition and implementation of service reliability measures including SLIs, SLOs, error budgets, and production health indicators.
  • Incident response support, problem management, post-incident reviews, and corrective action tracking.
  • Reliability improvement initiatives focused on capacity, performance, availability, monitoring, and operational readiness.

Technical Implementation:

  • Implementation of observability practices covering logs, metrics, traces, alerts, and dashboards.
  • Integration with operational tooling including monitoring platforms, Jira, Jenkins, incident management processes, and release pipelines.
  • Automation of repetitive operational tasks, health checks, alert routing, and recovery procedures where feasible.
  • Establishment of production readiness checks, runbooks, escalation paths, and service ownership practices.

Experience:

  • 3-5 years of experience


Contact information

Ayesha Mohamed