31 Jul
|
Datacom
|
Melbourne
Our purposeHere at Datacom we connect people and technology in order to solve challenges, create opportunities and discover current possibilities for the communities we live in.About the RoleWe are seeking a motivated and technically capable Site Reliability Engineerto help maintain, and continuously improve the reliability, scalability, security, and performance of our enterprise platforms.You will work across cloud technologies, infrastructure automation, observability, security, and operational processes to ensure services remain resilient and available.
This role requires a strong operational mindset, a passion for automation, and the ability to troubleshoot complex technical issues across hybrid and cloud environments.To be successful in this role, you must be have full working rights (Citizen, Permanent Resident or Visa Holder), no sponsorships provided and able to get a National Police Check clearance.What you'll doReliability EngineeringImplement, and maintain highly available and resilient infrastructure.Drive continuous service improvements through performance analysis and operational metrics.Support disaster recovery planning including Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO)Observability & MonitoringDevelop and maintain monitoring, logging, and alerting solutions.Proactively identify reliability risks before they impact services.Analyse system performance and capacity trends.Create operational dashboards that provide actionable insights.Incident Response & Problem ManagementParticipate in and support on-call rotations.Lead troubleshooting and resolution of production incidents and service outages.Develop and maintain incident response runbooks and playbooks.Conduct blameless post-incident reviews and drive corrective actions.Automation & DevSecOpsImplement and maintain Infrastructure as Code (IaC) solutions.Automate operational tasks, deployments, and service recovery processes.Embed security controls throughout the technology lifecycle using DevSecOps principles.Contribute to CI/CD pipeline development and optimisation.Cloud & Platform EngineeringSupport cloud environments across Azure, AWS and Google Cloud Platform.Assist with cloud architecture, governance, landing zones, and platform standardisation.Support containerised workloads using Docker and Kubernetes.Collaborate with engineering and security teams to improve platform reliability and performance.What you'll bringTechnical Experience5+ years in Infrastructure, Cloud Engineering, Platform Engineering, Systems Engineering, or Site Reliability Engineering roles.Experience supporting enterprise cloud platforms including:Microsoft AzureAmazon Web Services (AWS)Google Cloud Platform (GCP)Experience administeringActive DirectoryMicrosoft Entra IDMicrosoft IntuneGoogle WorkspaceStrong understanding of virtualisation technologies.Experience with cloud landing zones and platform governance.Familiarity with containerisation technologies such as Docker and Kubernetes.Reliability & OperationsExperience implementing monitoring, logging and alerting solutions.Understanding of SLIs, SLOs and Error Budgets.Experience supporting high-availability environments.Knowledge of incident management and problem management practices.Understanding of business continuity, disaster recovery, RTO and RPO requirements.Security & GovernanceFamiliarity with:Infrastructure as Code (IaC)DevSecOps practicesZero Trust security principlesSecurity and governance frameworksFrameworks & StandardsWorking knowledge of:NIST Cybersecurity FrameworkITILCOBITISO *****
#J-*****-Ljbffr
📌 Site Reliability Engineer (Melbourne)
🏢 Datacom
📍 Melbourne