20 Aug
|
N2S.Global
|
New South Wales
20 Aug
N2S.Global
New South Wales
We are seeking a highly skilledSRE Observability Engineerto design, implement, and optimize observability solutions across distributed systems.
The ideal candidate will ensure system reliability, performance, and scalability through effective monitoring, logging, and tracing strategies.
Key Responsibilities
Design and implementobservability frameworksincluding monitoring, logging, and distributed tracing
Build and maintain dashboards, alerts, and SLIs/SLOs for system health and performance
Work closely withSRE, DevOps, and engineering teamsto improve system reliability and incident response
Analyze system performance, identify bottlenecks, and driveproactive improvements
Develop and maintainautomation scripts and toolingfor observability and incident management
Leadroot cause analysis (RCA)and post-incident reviews
Ensurehigh availability and resilienceof cloud-native applications
Enable end-to-end visibility across microservices and distributed systems
Required Skills & Experience
Strong experience inSite Reliability Engineering (SRE)or DevOps roles
Hands-on experience withobservability toolssuch as:
Experience withcloud platforms(AWS / Azure / GCP)
Strong understanding ofmicroservices architectureand distributed systems
Experience inmonitoring, logging, alerting, and tracing
Proficiency in scripting/programming (Python, Go, or Bash)
Knowledge ofCI/CD pipelines and containerization (Docker, Kubernetes)
Preferred Qualifications
Experience definingSLOs, SLIs, and error budgets
Exposure tochaos engineering and resilience testing
Knowledge ofInfrastructure as Code (Terraform, ARM, CloudFormation)
Strong troubleshooting and analytical skills
Soft Skills
Strong problem-solving and analytical mindset
Excellent collaboration and communication skills
Ability to work in rapid-paced, production-critical environments
#J-*****-Ljbffr
📌 Sre Observability Engineer (New South Wales)
🏢 N2S.Global
📍 New South Wales