Site Reliability Engineer (SRE) (GovTech)
Avepoint · Singapore · 2026-06-24
About this role
<p>We are seeking a skilled and passionate Engineer to join our team to build and operate a Whole-of-Government (WoG) runtime platform.</p> <p>As a Site Reliability Engineer, you will be responsible for designing and operating GitLab, AWS&nbsp;and Kubernetes-based infrastructure and solutions that power our platform, to ensure the stability, scalability, and performance of our runtime platform.</p> <p><strong>Responsibilities:</strong></p> <p>As a Site Reliability Engineer, you will be responsible for:<br><span style="text-decoration: underline;">Toil Reduction &amp; Automation</span><br>• Identify repetitive tasks and develop automation via CI/CD pipelines, ensuring integration with cross-functional teams to reduce manual intervention and improve operational efficiency.<br><span style="text-decoration: underline;">Observability &amp; System Health</span><br>• Implement comprehensive observability solutions (logs, metrics, traces, alerts) around the four Golden Signals (latency, traffic, errors, saturation), and build automation for proactive system health assessments and self-remediation.<br><span style="text-decoration: underline;">Production Support &amp; Incident Management</span><br>• Participate in on-call rotations, promptly respond to incidents to minimize MTTR, and conduct thorough post-incident reviews to implement preventive measures and improve system resilience.<br><span style="text-decoration: underline;">Security &amp; Compliance</span><br>• Design and implement solutions that are secure and compliant by collaborating with dedicated security teams, conducting regular audits, and integrating advanced vulnerability scanning tools.</p> <p><span style="text-decoration: underline;">Maintenance, Optimisation &amp; Performance</span><br>• Identify and resolve performance bottlenecks and operational issues, define and track KPIs (e.g., MTTR, system uptime, cost efficiency), and drive ongoing optimisation efforts.<br><span style="text-decoration: underline;">Strategic Customer Engagement</span><br>• Act as a technical advisor for tenants, guiding them on containerization, and best practices for cloud-native deployments, and participating in strategic initiatives to enhance platform scalability and performance.<br><span style="text-decoration: underline;">Knowledge Sharing &amp; Documentation</span><br>• Develop and maintain detailed playbooks, runbooks, and documentation to facilitate team-wide knowledge sharing, streamline incident response, and ensure that critical processes are well understood across the team.<br><span style="text-decoration: underline;">Continuous Learning &amp; Innovation</span><br>• Stay current with the latest AWS, Kubernetes, and industry developments, and proactively recommend improvements and innovative solutions to maintain a competitive and reliable platform.</p> <p><strong>Requirements:</strong></p> <p>• Bachelor's degree or Diploma in Computer Science, Engineering, or a related field (or equivalent experience).<br>• Proven experience as a Site Reliability Engineer or similar role, with a strong background in containerization, orchestration, and cloud-native technologies.<br>• Proven ability to troubleshoot and resolve complex technical issues in containerized applications.<br>• Demonstrated experience with incident management, including post-incident reviews and continuous improvement.<br>• Strong documentation skills and experience in knowledge sharing across teams.<br>• Deep understanding of AWS, Kubernetes (including AWS EKS), and operational best practices, with familiarity in multi-cloud or hybrid environments.<br>• Solid grasp of networking, security, and storage in both AWS and Kubernetes contexts.<br>• Experience integrating Kubernetes with AWS cloud technologies (e.g., Secrets Manager, Load Balancers) and using infrastructure-as-code (Terraform or similar).<br>• Hands-on experience with containerization tools (Kubernetes, Kustomize, Helm) and automation scripting (Go, Python, Bash, or equivalent).<br>• Ability to write and maintain automated tests or conduct thorough manual testing for automation scripts, ensuring the reliability and effectiveness of automated solutions.<br>• Familiarity with CI/CD tools (GitLab CI/CD, ArgoCD) and version control systems (Git).<br>• Experience with observability/monitoring tools (Prometheus, Grafana, ELK Stack) and defining SLOs and Error Budgets.<br>• Certifications such as Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD) are a plus.<br>• Experience with developing Kubernetes operators using Go, service mesh technologies, and Chaos Engineering is a plus.</p> <p><strong>Soft skills:</strong></p> <p>• Proactive in identifying problems and recommending strategic solutions.<br>• Excellent problem-solving skills with a robust analytical mindset.<br>• Clear, concise, and effective communication skills; adept at collaborating across crossfunctional teams, including development, security, and customer-facing groups.<br>• Ability to remain calm and effective under pressure, especially during incident response.<br>• Adaptability to rapid change with a continuous learning mindset, sharing knowledge to foster team growth.<br>• Customer-focused with the ability to translate technical insights into understandable, actionable guidance.<br>• Leadership and mentoring capabilities, contributing to the development of a resilient and…
Skills asked for
- sre
- aws
- kubernetes
- ci/cd
- terraform
- helm
- go
- python
Similar jobs
- Site Reliability Engineer SRE - GlobalizationFuku · Singapore
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.