Sr Site Reliability Engineer (SRE) (Sydney)

Sr Site Reliability Engineer (SRE) (Sydney)

26 Sep
|
CareCone Group
|
Sydney

26 Sep

CareCone Group

Sydney

Role: Sr Site Reliability Engineer (SRE)

Experience: 8-10 years

Location: Sydney

Key Technologies

AWS | EKS | ECS | EC2 | Lambda | GitLab | Terraform | Ansible | Python | Bash | Linux | Kubernetes | Docker | PostgreSQL | Oracle | Vault | CloudWatch | OpenTelemetry | Grafana | Splunk | ELK | SLI | SLO | SLA | DevOps | Site Reliability Engineering

Role Summary

We are seeking an experienced AWS DevOps & Site Reliability Engineer (SRE) to build, automate, operate, and continuously improve cloud platforms and mission-critical applications. The role focuses on AWS cloud engineering, infrastructure automation, CI/CD, platform reliability, observability, incident management, resiliency, and operational excellence within a highly regulated enterprise workplace.

The successful candidate will drive automation, reliability, performance, scalability, and availability of cloud-native platforms while partnering with engineering teams to improve release quality, operational stability, and customer experience.

Key Responsibilities

Cloud & Platform Engineering

- Design, build, and maintain AWS cloud infrastructure and platform services.
- Develop and manage CI/CD pipelines using GitLab.
- Support cloud migration and platform modernization initiatives.
- Implement Infrastructure as Code (Terraform) and configuration management (Ansible).
- Automate deployments, environment provisioning, database refreshes, and operational processes.
- Manage secrets, certificates, access controls, and cloud security controls.
- Develop reusable infrastructure modules, deployment standards, and platform engineering patterns.

Site Reliability Engineering

- Define and implement reliability engineering practices, operational standards, and platform blueprints.
- Establish and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
- Improve platform reliability, scalability, resilience, and fault tolerance for mission-critical applications.




- Drive proactive reliability improvements through automation and elimination of operational toil.
- Conduct capacity planning, performance tuning, and resource optimization activities.
- Support disaster recovery planning, backup validation, resilience testing, and business continuity initiatives.
- Participate in production readiness reviews and ensure operational requirements are embedded into solution designs.

Observability & Monitoring

- Design and implement monitoring, logging, alerting, and observability frameworks.
- Build and maintain dashboards, health checks, metrics, and operational reporting.
- Enhance end-to-end system visibility using CloudWatch, Grafana, Splunk, ELK, OpenTelemetry, or similar technologies.
- Drive alert tuning and noise reduction to improve operational effectiveness.
- Establish monitoring standards across applications, infrastructure, databases, and integration components.

Incident & Operational Management

- Lead incident triage, troubleshooting, root cause analysis, and post-incident reviews.
- Support production systems and provide timely resolution of infrastructure and application issues.
- Drive reduction of Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
- Develop operational runbooks, knowledge articles, and recovery procedures.
- Collaborate with engineering and support teams to implement preventative actions and continuous improvement initiatives.
- Participate in on-call and major incident management activities where required.

DevOps & Engineering Enablement

- Promote DevOps, SRE, and Platform Engineering best practices.




- Collaborate with development teams to improve deployment reliability and automation maturity.
- Integrate security, observability, and operational controls into CI/CD pipelines.
- Support release management, change governance, and deployment strategies.
- Champion Infrastructure as Code, automation, and self-service platform capabilities.

Mandatory Skills

AWS Cloud

- AWS EC2, ECS/EKS, Lambda, VPC, IAM, S3, CloudWatch
- RDS/Aurora
- Route53
- Secrets Manager
- AWS Networking and Security Services

DevOps & Automation

- GitLab CI/CD
- Terraform
- Ansible
- Bash and Python scripting
- Linux/Unix Administration
- Infrastructure Automation

Site Reliability Engineering

- Production Support and Incident Management
- SLI / SLO / SLA implementation and management
- Reliability Engineering practices
- High Availability and Fault-Tolerant Architecture Design
- Capacity Planning and Performance Optimization
- Disaster Recovery and Business Continuity
- Root Cause Analysis and Problem Management

Observability

- Monitoring, Logging and Alerting
- CloudWatch
- Grafana
- Splunk / ELK
- OpenTelemetry
- Operational Metrics and Dashboarding

Platform & Containers

- Kubernetes (EKS preferred)
- Containerisation (Docker)
- Networking and Infrastructure Troubleshooting

Security

- Secrets Management Solutions (Vault preferred)
- IAM and Access Controls
- Cloud Security Best Practices

Preferred Skills

- AWS Certified Solutions Architect / DevOps Engineer Certification
- Certified Kubernetes Administrator (CKA)
- Experience with Aurora PostgreSQL and Oracle migration programs
- Chaos Engineering and Resilience Testing
- Experience implementing SRE operating models
- Service Mesh technologies
- Financial Services, Banking, Capital Markets, or other regulated industries
- ITIL processes including Incident, Problem, Change and Release Management
- Experience supporting large-scale cloud-native microservices platforms

📌 Sr Site Reliability Engineer (SRE) (Sydney)
🏢 CareCone Group
📍 Sydney

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sr site reliability engineer (sre) (sydney) / sydney

Subscribe to this job alert:

Get the latest job offers by email for: sr site reliability engineer (sre) (sydney) / sydney