25 Aug
|
AI Talent
|
Sydney
We are partnering with an enterprise client to strengthen platform resilience, validate fault tolerance, and ensure zero-downtime reliability for mission-critical digital systems. We are seeking a skilled Chaos Engineer (Resilience Automation Specialist) to represent our organisation and take technical ownership of fault injection experiments, automated resilience testing, and blast-radius verification.
In this role, you will bridge Site Reliability Engineering (SRE) and systems architecture within our client’s environment. You will intentionally simulate real-world failure scenarios—such as network latency, node dropouts, database failovers, and region outages—to uncover hidden failure modes and prove that distributed systems self-heal under duress.
? Visa & Sponsorship Options As the employer of record, we provide full visa and migration support for qualified engineering talent deployed to our clients:
- 482 On-Hire Sponsorship Transfers: Fully supported for qualified candidates currently in Australia on an existing 482 visa looking to transfer sponsorship to work with our clients.
- Recent 482 Visa Sponsorship: Available for qualified candidates meeting the commercial experience and technical requirements.
- Temporary & Working Visa Holders: Open to all working visa holders seeking a direct pathway to employer sponsorship.
Core Responsibilities
- Chaos Experimentation: Design, execute, and automate controlled chaos experiments across distributed microservices and containerized environments.
- Fault Injection Tooling: Deploy and manage enterprise chaos engineering platforms (e.g., Chaos Mesh, Gremlin, LitmusChaos, AWS Fault Injection Simulator).
- GameDays &
• Failure Drills:
Organize and facilitate systematic GameDay scenarios to test system recovery, automated failover mechanisms, and operational runbooks.
- Resilience in CI/CD: Embed automated resilience assertions and steady-state validation directly into continuous deployment pipelines.
- Observability &
• Blast Radius Control:
Monitor real-time telemetry (metrics, traces, logs) during experiments using tools like Prometheus, Grafana, Datadog, or OpenTelemetry, ensuring safe rollback triggers.
- Architecture Hardening: Collaborate with software engineering and cloud infrastructure teams to translate experiment findings into actionable architectural improvements (circuit breakers, retry policies, auto-healing).
Selection Criteria
- SRE &
• Resilience Mastery:
Commercial experience in Site Reliability Engineering (SRE) or platform reliability with a strong focus on distributed systems resilience.
- Chaos Engineering Tooling: Hands-on experience configuring fault-injection tools such as Chaos Mesh, Gremlin, LitmusChaos, or AWS FIS.
- Container &
• Cloud Ecosystems:
Deep understanding of Kubernetes architecture, container networking, service meshes, and cloud infrastructure (AWS/Azure).
- Scripting &
• Automation:
Strong programming or scripting skills in Python, Go, or Bash to build custom failure-injection scenarios and verification logic.
- Location Requirements: Currently residing in Australia with valid work rights or eligibility for 482 visa sponsorship/transfer.
Preferred Qualifications (Nice to Have)
- Practical experience implementing distributed tracing (OpenTelemetry, Jaeger).
- Certified Kubernetes Administrator (CKA) or relevant Cloud certifications.
- Experience in high-availability, low-latency domains (Banking, Fintech, Telecommunications, E-Commerce).
📌 Chaos Engineer (Resilience Automation Specialist) (Sydney)
🏢 AI Talent
📍 Sydney