20 Aug
|
N2S.Global
|
New South Wales
20 Aug
N2S.Global
New South Wales
We are seeking a highly skilled
SRE Observability Engineer
to design, implement, and optimize observability solutions across distributed systems. The ideal candidate will ensure system reliability, performance, and scalability through effective monitoring, logging, and tracing strategies.
Key Responsibilities
Design and implement
observability frameworks
including monitoring, logging, and distributed tracing
Build and maintain dashboards, alerts, and SLIs/SLOs for system health and performance
Work closely with
SRE, DevOps, and engineering teams
to improve system reliability and incident response
Analyze system performance, identify bottlenecks, and drive
proactive improvements
Develop and maintain
automation scripts and tooling
for observability and incident management
Lead
root cause analysis (RCA)
and post-incident reviews
Ensure
high availability and resilience
of cloud-native applications
Enable end-to-end visibility across microservices and distributed systems
Required Skills & Experience
Strong experience in
Site Reliability Engineering (SRE)
or DevOps roles
Hands-on experience with
observability tools
such as:
Experience with
cloud platforms
(AWS / Azure / GCP)
Strong understanding of
microservices architecture
and distributed systems
Experience in
monitoring, logging, alerting, and tracing
Proficiency in scripting/programming (Python, Go, or Bash)
Knowledge of
CI/CD pipelines and containerization (Docker, Kubernetes)
Preferred Qualifications
Experience defining
SLOs, SLIs, and error budgets
Exposure to
chaos engineering and resilience testing
Knowledge of
Infrastructure as Code (Terraform, ARM, CloudFormation)
Robust troubleshooting and analytical skills
Soft Skills
Strong problem-solving and analytical mindset
Excellent collaboration and communication skills
Ability to work in fast-paced, production-critical environments
#J-*****-Ljbffr
📌 Sre Observability Engineer (New South Wales)
🏢 N2S.Global
📍 New South Wales