Software Engineer - Reliability/SRE

Electrum Software· Engineering Platforms
Apply Now ↗
📍 Cape Town, South AfricaFull time

About this role

Electrum is a next-generation payment software technology company.

Since 2012, we've delivered trusted, enterprise-grade, cloud-native software to optimise financial transaction processing. Our deep expertise has established us as a respected partner in high-volume, low-value payment schemes, enabling clients to deliver services to millions of South Africans daily.

At Electrum, we are grounded in impact – designing solutions that matter, acting with urgency, and continuously learning as we scale. We believe in creating together – working side by side with our clients and teams to build meaningful, lasting solutions. We prioritise making it safe – encouraging open communication, smart risk-taking, and trust so that creativity and alignment thrive. And we back empowered strong teams – hiring brilliant people, collaborating hard, and holding each other to high standards while leading with empathy and kindness.

The Role

As a Software Engineer with a focus on reliability, you'll write and ship software that keeps our payments systems healthy, scalable, and secure. You'll work as part of the engineering organisation: building tooling and services, improving how we deploy and operate what we ship, and solving hard problems in production with the same engineering craft you bring to code. You'll collaborate closely with other development teams to design for scale, address issues before they become incidents, and raise the bar on how we observe and run our systems. You'll take part in on-call rotations, help manage critical incidents, and contribute to response processes so we resolve issues quickly and learn from them. You'll also help shape decisions around security, system optimisation, rollout strategies, and monitoring for health and availability.

Responsibilities

Software Engineering for Reliability

  • Design, build, and maintain software, tooling, and automation that improve the reliability, availability, and scalability of our applications and services.
  • Work closely with other engineers to understand, address, and prevent technical issues in the systems we own and depend on.
  • Contribute to codebases and shared platforms with the same standards of review, testing, and delivery as the wider engineering team.
  • Participate in on-call rotations and manage critical incidents.
  • Develop and maintain incident response processes and alerting mechanisms.
  • Develop and maintain tools to monitor application and service SLIs and SLOs.

System Troubleshooting and Problem Resolution

  • Diagnose and resolve infrastructure and system-level issues, ensuring minimal downtime and swift problem resolution.
  • Respond to and investigate incidents related to infrastructure and applications, utilising diagnostic tools to track down and remediate issues.
  • Participate in on-call rotations to provide 24/7 operational support as necessary.

Observability and Automation

  • Utilise technologies to develop and maintain effective log management and monitoring solutions for internal and external customers.
  • Evaluate system health, identify performance bottlenecks and proactively optimise performance and cost-effectiveness.
  • Implement automation tools and frameworks for deployment, configuration, and monitoring processes.
  • Capacity management and planning for systems to ensure continued reliability.

Process Improvements

  • Offer recommendations and improvements to enhance performance, security, and scalability.
  • Evaluate and integrate emerging technologies, cloud services and automation tools to improve operational efficiency.
  • Drive cost-optimization initiatives by identifying opportunities for resource right-sizing, efficiency and other cost-saving measures.

Disaster Recovery

  • Design and implement disaster recovery strategies, including backup and restoration processes, to ensure business continuity.
  • Develop and update incident management procedures, ensuring effective incident response by providing technical solutions and implementing preventative measures.
  • Regularly assess system performance, identify irregularities, troubleshoot issues, and ensure high system availability. This includes performing or facilitating Disaster Recovery tests.

  • Bachelor's degree in Computer Science, Information Technology, or related field.
  • 3+ years experience as a Java software engineer
  • Experience with observability tooling and pipelines, e.g. DataDog, Elastic/ELK Stack or Grafana.
  • Hands-on Software Engineering and scripting experience,
  • Proficient troubleshooting and problem-solving skills.
  • Excellent prioritisation and time management skills.
  • Attention to detail and ability to work effectively in a team environment.

Advantage

  • Experience in SRE, DevOps, Platforms or similar role.
  • Familiarity with Cloud services dealing with Computer, Object Storage, Databases, Serverless Computer, Monitoring & Observability.

Why Join Electrum?

  • We believe in a People First approach, ensuring a culture where you can thrive and make a real difference

Your Career & Culture

  • Career Growth: Delivering world-class financial software is challenging, but your effort will earn you hands-on experience with products used by millions, accelerating your career.
  • Strong Teams: We keep teams small, focused, and collaborative to maximize impact.
  • Transparency: We openly discuss strategy, finances, and salaries. Mistakes are viewed as learning opportunities that we actively discuss.
  • Autonomy: We trust you. You're expected to seek out the data needed for informed decisions and manage your own time—knowing when to focus and when to recharge.
  • Shared Vision: You'll have the power to shape the vision of how we build the future of financial services.

Practical Perks

  1. Here's how we support our culture:
    • Flexible Work: Office-first environment with flexible hours.
    • Generous Leave: Starting at 20 days per year.
    • Office Perks (Cape Town): Fully-stocked kitchen and daily catered lunch.
  2. Social Life: Regular team activities like hikes, getaways, and dinners

Frequently Asked Questions

Is the salary disclosed for the Software Engineer - Reliability/SRE position at Electrum Software?
The salary for this Software Engineer - Reliability/SRE role at Electrum Software is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Software Engineer - Reliability/SRE position at Electrum Software located?
This Software Engineer - Reliability/SRE role at Electrum Software is based in Cape Town, South Africa. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Software Engineer - Reliability/SRE role at Electrum Software full-time or part-time?
This is listed as a Full time position. It is posted as a Software Engineer - Reliability/SRE role in the Engineering Platforms department at Electrum Software.
Which team or department does the Software Engineer - Reliability/SRE at Electrum Software belong to?
This Software Engineer - Reliability/SRE position is part of the Engineering Platforms department at Electrum Software. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Software Engineer - Reliability/SRE position at Electrum Software?
Click the "Apply Now" button on this page. You will be redirected to Electrum Software's official application portal hosted on workable where you can submit your application directly.
When was the Software Engineer - Reliability/SRE job at Electrum Software posted?
This Software Engineer - Reliability/SRE position at Electrum Software was posted on May 22, 2026. Apply as soon as possible — early applications are often reviewed first.
Software Engineer - Reliability/SRE
Electrum Software
Apply for this role ↗

You'll be redirected to Electrum Software's official application page on workable.