Work Remotely
  • Post Date: February 17, 2026
  • Apply Before: August 31, 2026
Job Description

Job title:  Lead DevOps Engineer

Location: Remote (Johannesburg)

A vacancy is available for a Lead DevOps Engineer, skilled in GCP (Google Cloud).

Purpose of the role:

As the Lead DevOps Engineer specialising in GCP – Google Cloud, you will inspire and lead a high-performing Site Reliability Engineering (SRE) team to achieve reliability and scalability in production environments.

By implementing innovative monitoring, automation, and DevOps practices on Google Cloud Platform, you will elevate system uptime, efficiency, and performance.

Your mentorship and dedication will cultivate a culture of engineering excellence, empowering your team to thrive and make a lasting impact.

Education and Certifications required:

  • Degree or Diploma in Information Technology, Computer Science, or equivalent experience.
  • Google Cloud certifications (e.g., Professional Cloud DevOps Engineer, Professional Cloud Architect) are highly advantageous.

Experience required:

  • 5+ years of experience working on GCP infrastructure and services.
  • Experience in a management/leadership capacity within SRE/DevOps teams (3+ Years).
  • Experience with Kubernetes, Docker, and container orchestration at scale.
  • Familiarity with incident management, post-mortem processes, and production monitoring tools.
  • Hands-on experience with IaC tools such as Terraform, Ansible, or Deployment Manager.
  • Experience working with CI/CD pipelines and automation tools.
  • UNIX/Linux administration expertise.
  • Familiarity with security, compliance, and cost optimisation on GCP.

Key Responsibilities and Outputs:

  • Lead and mentor a team of SRE engineers, promoting knowledge sharing and growth.
  • Act as the technical authority on SRE practices for GCP, ensuring system reliability and uptime across environments.
  • Oversee team workload distribution and manage stakeholder expectations.
  • Champion and implement DevOps and SRE best practices with emphasis on automation and scalability.
  • Drive monitoring and observability initiatives, leveraging tools like Grafana, Prometheus, and Stackdriver.
  • Design, maintain and optimise CI/CD pipelines using GCP-native tools and industry standards.
  • Troubleshoot complex production incidents, ensuring root cause analysis and long-term fixes.
  • Collaborate with cross-functional teams to ensure consistent platform performance.
  • Apply Infrastructure as Code (IaC) principles using tools such as Terraform or Deployment Manager.
  • Stay abreast of emerging technologies to continually evolve our tooling and architecture continually.
  • Foster a proactive and blameless incident management culture.

IMPORTANT INFO:  

  • South African citizenship is essential
  • By submitting your application and personal information, you explicitly consent to Let’s Recruit processing your personal data solely for the purposes of evaluating your suitability for this position and other potential opportunities. All personal information provided will be handled in compliance with applicable South African data protection laws and will be securely retained or destroyed as required by legislation.
  • While we strive to provide responses to all applicants, if you do not hear from us within 14 days of your application, please consider your application unsuccessful.
  • Successful candidates will be notified within 14 days of application.
  • Let’s Recruit reserves the right to withdraw or modify this vacancy at any time without notice.

To apply, send your detailed CV to cv@letsrecruit.co.za