06 Oct
|
CareCone Group
|
Sydney
06 Oct
CareCone Group
Sydney
Job Description
Role: Sr Site Reliability Engineer (SRE) n
Experience: 8-10 years
Location: Sydney Key Technologies n
AWS | EKS | ECS | EC2 | Lambda | GitLab | Terraform | Ansible | Python | Bash | Linux | Kubernetes | Docker | PostgreSQL | Oracle | Vault | CloudWatch | OpenTelemetry | Grafana | Splunk | ELK | SLI | SLO | SLA | DevOps | Site Reliability Engineering
Role Summary n
We are seeking an experienced AWS DevOps & Site Reliability Engineer (SRE) to build, automate, operate, and continuously improve cloud platforms and mission-critical applications. The role focuses on AWS cloud engineering, infrastructure automation, CI/CD, platform reliability, observability, incident management, resiliency, and operational excellence within a highly regulated enterprise workplace.
n
The successful candidate will drive automation, reliability, performance, scalability, and availability of cloud-native platforms while partnering with engineering teams to improve release quality, operational stability, and customer experience.
Key Responsibilities n
n
- Design, build, and maintain AWS cloud infrastructure and platform services.
n
- Develop and manage CI/CD pipelines using GitLab.
n
- Support cloud migration and platform modernization initiatives.
n
- Implement Infrastructure as Code (Terraform) and configuration management (Ansible).
n
- Automate deployments, environment provisioning, database refreshes, and operational processes.
n
- Manage secrets, certificates, access controls, and cloud security controls.
n
- Develop reusable infrastructure modules, deployment standards, and platform engineering patterns.
n
Site Reliability Engineering nn
- Define and implement reliability engineering practices, operational standards, and platform blueprints.
n
- Establish and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
n
- Improve platform reliability, scalability, resilience, and fault tolerance for mission-critical applications.
n
- Drive proactive reliability improvements through automation and elimination of operational toil.
n
- Conduct capacity planning, performance tuning, and resource optimization activities.
n
- Support disaster recovery planning, backup validation, resilience testing, and business continuity initiatives.
n
- Participate in production readiness reviews and ensure operational requirements are embedded into solution designs.
n
Observability & Monitoring nn
- Design and implement monitoring, logging, alerting, and observability frameworks.
n
- Build and maintain dashboards, health checks, metrics, and operational reporting.
n
- Enhance end-to-end system visibility using CloudWatch, Grafana, Splunk, ELK, OpenTelemetry, or similar technologies.
n
- Drive alert tuning and noise reduction to improve operational effectiveness.
n
- Establish monitoring standards across applications, infrastructure, databases, and integration components.
n
Incident & Operational Management nn
- Lead incident triage, troubleshooting, root cause analysis, and post-incident reviews.
n
- Support production systems and provide timely resolution of infrastructure and application issues.
n
- Drive reduction of Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
n
- Develop operational runbooks, knowledge articles, and recovery procedures.
n
- Collaborate with engineering and support teams to implement preventative actions and continuous improvement initiatives.
n
- Participate in on-call and major incident management activities where required.
n
- Promote DevOps, SRE,
and Platform Engineering best practices.
n
- Collaborate with development teams to improve deployment reliability and automation maturity.
n
- Integrate security, observability, and operational controls into CI/CD pipelines.
n
- Support release management, change governance, and deployment strategies.
n
- Champion Infrastructure as Code, automation, and self-service platform capabilities.
n
Mandatory Skills AWS Cloud nn
- AWS EC2, ECS/EKS, Lambda, VPC, IAM, S3, CloudWatch
n
- RDS/Aurora
n
- Route53
n
- Secrets Manager
n
- AWS Networking and Security Services
n
DevOps & Automation nn
- Terraform
n
- Ansible
n
- Bash and Python scripting
n
- Infrastructure Automation
n
Site Reliability Engineering nn
- Production Support and Incident Management
n
- SLI / SLO / SLA implementation and management
n
- Reliability Engineering practices
n
- High Availability and Fault-Tolerant Architecture Design
n
- Capacity Planning and Performance Optimization
n
- Disaster Recovery and Business Continuity
n
- Root Cause Analysis and Problem Management
n
Observability nn
- Monitoring, Logging and Alerting
n
- CloudWatch
n
- Grafana
n
- Splunk / ELK
n
- OpenTelemetry
n
- Operational Metrics and Dashboarding
n
Platform & Containers nn
- Containerisation (Docker)
n
- Networking and Infrastructure Troubleshooting
n
Security nn
- Secrets Management Solutions (Vault preferred)
n
- IAM and Access Controls
n
- Cloud Security Best Practices
n
Preferred Skills nn
- AWS Certified Solutions Architect / DevOps Engineer Certification
n
- Experience with Aurora PostgreSQL and Oracle migration programs
n
- Chaos Engineering and Resilience Testing
n
- Experience implementing SRE operating models
n
- Service Mesh technologies
n
- Financial Services, Banking, Capital Markets, or other regulated industries
n
- ITIL processes including Incident, Problem, Change and Release Management
n
#J-18808-Ljbffr
📌 Sr Site Reliability Engineer (SRE) (Sydney)
🏢 CareCone Group
📍 Sydney