08 Aug
|
Talent
|
South Australia
08 Aug
Talent
South Australia
About the Role
We're looking for an experienced Site Reliability Engineer (SRE) to lead the reliability, performance, and scalability of enterprise platforms.
You'll work closely with software engineering and infrastructure teams to improve automation, cloud platforms, observability, and operational excellence while mentoring other engineers and driving continuous improvement.
Key Responsibilities
Design, build and support highly available, scalable, and resilient cloud infrastructure and platforms.
Improve system reliability through automation, Infrastructure as Code (IaC), and CI/CD best practices.
Lead major incident response, root cause analysis, and implement preventative improvements.
Develop and maintain monitoring, logging, alerting, and observability solutions.
Establish and manage SLOs, SLIs, and error budgets.
Build and maintain CI/CD pipelines and deployment automation.
Collaborate with development teams to improve application performance, reliability, and operational readiness.
Lead platform modernisation, capacity planning, and disaster recovery initiatives.
Provide technical leadership, mentor engineers, and contribute to engineering standards and best practices.
Participate in Agile ceremonies and support continuous improvement across the engineering team.
Skills & Experience Required
Essential
7+ years
of experience in Site Reliability Engineering, DevOps, Cloud Engineering, Infrastructure Engineering,
or a similar role.
Strong experience administering
Windows Server
and
Linux
environments.
Experience with cloud platforms such as
Azure
(preferred), AWS, or GCP.
Strong automation and scripting skills using
PowerShell
, plus
Python, Bash, or Go
.
Experience building and maintaining
CI/CD pipelines
using Azure DevOps, GitHub Actions, or similar.
Hands-on experience with
Infrastructure as Code
using
Terraform
, Ansible, or equivalent.
Experience supporting
Kubernetes
and containerised applications.
Solid understanding of networking fundamentals including
DNS, load balancing, firewalls, and network security
.
Experience implementing monitoring, logging, and observability platforms.
Strong troubleshooting skills with experience leading major incident investigations and root cause analysis.
Experience working within
Agile
and
DevOps
environments.
Excellent communication skills with the ability to work across technical and business stakeholders.
Desirable
Relevant tertiary qualification in IT, Computer Science, Engineering, or equivalent experience.
Microsoft, Azure, AWS, Kubernetes, Terraform, or ITIL certifications.
Experience driving platform modernisation and engineering best practices.
Previous experience mentoring engineers or leading technical initiatives.
📌 Senior Site Reliability Engineer (South Australia)
🏢 Talent
📍 South Australia