27 Aug
|
N2S.Global
|
Sydney
We are seeking a highly skilled Site Reliability Engineer (SRE) with robust expertise in Observability, Dynatrace, Application Performance Monitoring (APM), and Cloud Operations. You will play a key role in ensuring the reliability, availability, performance, and scalability of critical enterprise applications and infrastructure.
The ideal candidate will have hands-on experience with Dynatrace, monitoring and alerting frameworks, incident management, automation, and cloud-native environments.
Key Responsibilities
- Implement, configure, and support Dynatrace monitoring solutions across applications, infrastructure, and cloud platforms.
- Monitor system health, application performance, and business services through observability tools.
- Configure dashboards, alerts, synthetic monitoring, Real User Monitoring (RUM), and distributed tracing.
- Perform proactive monitoring and identify performance bottlenecks before business impact occurs.
- Lead incident triage, troubleshooting, root cause analysis (RCA), and problem management activities.
- Develop automation scripts for operational activities using Python, Shell, or PowerShell.
- Collaborate with Development, DevOps, Platform Engineering, Cloud, and Infrastructure teams.
- Support production environments and participate in major incident management processes.
- Define and maintain SLI, SLO, and SLA metrics.
- Drive observability best practices and platform reliability improvements.
- Support capacity planning and performance optimization initiatives.
Required Skills
Observability & Monitoring
- Dynatrace Administration
- Dynatrace OneAgent
- ActiveGate
- Real User Monitoring (RUM)
- Synthetic Monitoring
- Log Monitoring
- Distributed Tracing
- Application Performance Monitoring (APM)
SRE & Operations
- Site Reliability Engineering
- Incident Management
- Problem Management
- Root Cause Analysis
- Production Support
- Service Reliability
Cloud & Infrastructure
- AWS, Azure, or GCP
- Linux/Unix Administration
- Docker
- Kubernetes
- Networking Fundamentals
Automation & Scripting
- Python
- Shell Scripting
- PowerShell
Monitoring Tools (Good to Have)
- Grafana
- Prometheus
- Splunk
- ELK Stack
- Datadog
- AppDynamics
Preferred Qualifications
- Dynatrace Associate or Professional Certification
- Experience supporting enterprise-scale production environments
- Financial Services / Banking experience
- Experience with DevOps and CI/CD pipelines
- Exposure to Infrastructure as Code (Terraform, Ansible)
📌 Site Reliability Engineer (SRE) – Observability & Dynatrace (Sydney)
🏢 N2S.Global
📍 Sydney