Job Title : AWS DevOps & Site Reliability Engineer (SRE)
Location :Sydney
Mode: Permanent
We are seeking an experienced AWS DevOps & Site Reliability Engineer (SRE) to build, automate, operate, and continuously improve cloud platforms and mission-critical applications. The role focuses on AWS cloud engineering, infrastructure automation, CI/CD, platform reliability, observability, incident management, resiliency, and operational excellence within a highly regulated enterprise environment.
The successful candidate will drive automation, reliability, performance, scalability, and availability of cloud-native platforms while partnering with engineering teams to improve release quality, operational stability, and customer experience.
Key Responsibilities
Cloud & Platform Engineering
- Design, build, and maintain AWS cloud infrastructure and platform services.
- Develop and manage CI/CD pipelines using GitLab.
- Support cloud migration and platform modernization initiatives.
- Implement Infrastructure as Code (Terraform) and configuration management (Ansible).
- Automate deployments, workplace provisioning, database refreshes, and operational processes.
- Manage secrets, certificates, access controls, and cloud security controls.
- Develop reusable infrastructure modules, deployment standards, and platform engineering patterns.
Site Reliability Engineering
- Define and implement reliability engineering practices, operational standards, and platform blueprints.
- Establish and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
- Improve platform reliability, scalability, resilience, and fault tolerance for mission-critical applications.
- Drive proactive reliability improvements through automation and elimination of operational toil.
- Conduct capacity planning, performance tuning, and resource optimization activities.
- Support disaster recovery planning, backup validation,
resilience testing, and business continuity initiatives.
- Participate in production readiness reviews and ensure operational requirements are embedded into solution designs.
Observability & Monitoring
- Design and implement monitoring, logging, alerting, and observability frameworks.
- Build and maintain dashboards, health checks, metrics, and operational reporting.
- Enhance end-to-end system visibility using CloudWatch, Grafana, Splunk, ELK, OpenTelemetry, or similar technologies.
- Drive alert tuning and noise reduction to improve operational effectiveness.
- Establish monitoring standards across applications, infrastructure, databases, and integration components.
Incident & Operational Management
- Lead incident triage, troubleshooting, root cause analysis, and post-incident reviews.
- Support production systems and provide timely resolution of infrastructure and application issues.
- Drive reduction of Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
- Develop operational runbooks, knowledge articles, and recovery procedures.
- Collaborate with engineering and support teams to implement preventative actions and continuous improvement initiatives.
- Participate in on-call and major incident management activities where required.
DevOps & Engineering Enablement
- Promote DevOps, SRE, and Platform Engineering best practices.
- Collaborate with development teams to improve deployment reliability and automation maturity.
- Integrate security, observability, and operational controls into CI/CD pipelines.
- Support release management, change governance, and deployment strategies.
- Champion Infrastructure as Code, automation, and self-service platform capabilities.
Mandatory Skills
AWS Cloud
- AWS EC2, ECS/EKS, Lambda, VPC, IAM, S3, CloudWatch
- RDS/Aurora
- Route53
- Secrets Manager
- AWS Networking and Security Services
DevOps & Automation
- GitLab CI/CD
- Terraform
- Ansible
- Bash and Python scripting
- Linux/Unix Administration
- Infrastructure Automation
Site Reliability Engineering
- Production Support and Incident Management
- SLI / SLO / SLA implementation and management
- Reliability Engineering practices
- High Availability and Fault-Tolerant Architecture Design
- Capacity Planning and Performance Optimization
- Disaster Recovery and Business Continuity
- Root Cause Analysis and Problem Management
Observability
- Monitoring, Logging and Alerting
- CloudWatch
- Grafana
- Splunk / ELK
- OpenTelemetry
- Operational Metrics and Dashboarding
Platform & Containers
- Kubernetes (EKS preferred)
- Containerisation (Docker)
- Networking and Infrastructure Troubleshooting
Security
- Secrets Management Solutions (Vault preferred)
- IAM and Access Controls
- Cloud Security Best Practices
Preferred Skills
- AWS Certified Solutions Architect / DevOps Engineer Certification
- Certified Kubernetes Administrator (CKA)
- Experience with Aurora PostgreSQL and Oracle migration programs
- Chaos Engineering and Resilience Testing
- Experience implementing SRE operating models
- Service Mesh technologies
- Financial Services, Banking, Capital Markets, or other regulated industries
- ITIL processes including Incident, Problem, Change and Release Management
- Experience supporting large-scale cloud-native microservices platforms
Interested candidates can send their updated resume to
[email protected] or reach me @ M: 61283195529
📌 AWS DevOps & Site Reliability Engineer (SRE) (Sydney)
🏢 CareCone Group
📍 Sydney