16 Aug
|
N2S.Global
|
New South Wales
16 Aug
N2S.Global
New South Wales
We are seeking a highly skilled SRE Observability Engineer to design, implement, and optimize observability solutions across distributed systems.
The ideal candidate will ensure system reliability, performance, and scalability through effective monitoring, logging, and tracing strategies.
Key Responsibilities
Design and implement observability frameworks including monitoring, logging, and distributed tracing
Build and maintain dashboards, alerts, and SLIs/SLOs for system health and performance
Work closely with SRE, DevOps, and engineering teams to improve system reliability and incident response
Analyze system performance, identify bottlenecks, and drive proactive improvements
Develop and maintain automation scripts and tooling for observability and incident management
Lead root cause analysis (RCA) and post-incident reviews
Ensure high availability and resilience of cloud-native applications
Enable end-to-end visibility across microservices and distributed systems
Required Skills & Experience
Solid experience in Site Reliability Engineering (SRE) or DevOps roles
Hands-on experience with observability tools such as:
Experience with cloud platforms (AWS / Azure / GCP)
Strong understanding of microservices architecture and distributed systems
Experience in monitoring, logging, alerting, and tracing
Proficiency in scripting/programming (Python, Go, or Bash)
Knowledge of CI/CD pipelines and containerization (Docker, Kubernetes)
Preferred Qualifications
Experience defining SLOs, SLIs, and error budgets
Exposure to chaos engineering and resilience testing
Knowledge of Infrastructure as Code (Terraform, ARM, CloudFormation)
Strong troubleshooting and analytical skills
Soft Skills
Strong problem-solving and analytical mindset
Excellent collaboration and communication skills
Ability to work in fast-paced, production-critical environments
#J-*****-Ljbffr
📌 Sre Observability Engineer (New South Wales)
🏢 N2S.Global
📍 New South Wales