10 Oct
|
Westbury Partners
|
Sydney
10 Oct
Westbury Partners
Sydney
Join a high-impact SRE team operating large-scale data infrastructure, distributed systems and self-managed platforms, combining Linux expertise, automation, incident response and engineering to keep critical systems reliable.
What You'll Do:
Operate and evolve large-scale data infrastructure powering demanding users and workloads. You’ll work across distributed systems, storage, processing platforms and in-house technology, with reliability at the heart of everything you do.
Your responsibilities will include:
1. Operate and monitor distributed data platforms, including Kafka, HDFS, Dremio and in-house pipelines.
2. Build automation and CI/CD solutions for rapid infrastructure and software deployment.
3. Respond to incidents, participate in on-call rotations and engineer lasting improvements that prevent recurring failures.
4. Diagnose complex Linux, networking, storage, performance and infrastructure issues.
5. Plan and execute upgrades, migrations, capacity improvements and technology rollouts.
6. Work directly with traders, researchers and developers to troubleshoot problems and recommend effective data solutions.
7. Develop infrastructure tooling and automation using Python and configuration-management technologies.
8. Evaluate emerging technologies and introduce improvements to a continuously evolving data platform.
9. Collaborate with systems and network engineers and colleagues across international offices.
Why Join Us:
1. Own technology at serious scale: Work with multi-petabyte data infrastructure and millions of queries each day.
2. Solve real engineering problems: Manage infrastructure where reliability, capacity and failure recovery genuinely matter.
3.
Build deep technical expertise: Develop hands-on experience with distributed systems, Linux, Kubernetes, Kafka, HDFS and automation.
4. Learn from experienced engineers: Join a highly skilled team with structured development and dedicated technical ramp-up.
5. Make a visible impact: Your engineering decisions directly influence the reliability and performance of critical data platforms.
6. Work with demanding users: Partner directly with technical teams who need fast, practical solutions.
About You:
You’re an engineer who thinks like an SRE: you understand that keeping systems running is only the beginning. You investigate failures, understand their causes and build solutions that make the next incident less likely.
You’ll ideally bring:
1. Hands-on production operations experience with self-managed infrastructure.
2. Solid Linux fundamentals, including processes, filesystems, networking, memory and disk behaviour.
3. Experience operating at least one complex infrastructure platform deeply, particularly Kafka, HDFS or Kubernetes.
4. Experience with monitoring, alerting, incident response and production troubleshooting.
5. A track record of infrastructure automation and CI/CD.
6. Python experience, particularly for infrastructure tooling, deployment automation or operational workflows.
7. Experience with configuration management or infrastructure-as-code such as Ansible, Puppet or Terraform.
8. An appetite for understanding how complex systems work beneath the surface.
9. Strong communication skills and the confidence to work directly with technical users.
10. Curiosity, ownership and a willingness to develop breadth across a sophisticated data ecosystem.
📌 Site Reliability Engineer – Data Infrastructure & Distributed Systems - HFT - Sydney
🏢 Westbury Partners
📍 Sydney