15 Sep
|
Tribus
|
New South Wales
15 Sep
Tribus
New South Wales
Job Description
Site Reliability Engineer – Data Infrastructure
n
Sydney | Quantitative Trading | Linux, Kafka, Python
n
We are working with a leading global quantitative trading firm that is growing its Data Engineering team in Sydney.
n
This is an SRE role focused on the infrastructure that powers large-scale data platforms. You will help operate a multi-petabyte workplace supporting millions of queries each day, working across Linux, distributed data systems, automation and production reliability.
n
This is not a traditional data engineering role focused on writing pipelines, DAGs or analytics workloads. The team owns and operates the underlying platforms themselves.
n
What you'll be doing
n
n
- Operate and improve large-scale distributed data infrastructure
n
- Work with technologies including Kafka, HDFS, Kubernetes and distributed query platforms
n
- Own monitoring, alerting, incident response and production reliability
n
- Automate infrastructure deployment, upgrades and operational processes using Python
n
- Troubleshoot Linux, networking, storage and distributed systems issues
n
- Improve capacity, resilience,
failure handling and deployment processes
n
- Work directly with traders, researchers and developers to solve data infrastructure problems
n
- Participate in an on-call rotation and engineer out recurring issues
n
n
What we're looking for
n
n
- Hands-on Linux systems administration and troubleshooting experience
n
- Experience owning production systems, including monitoring, incidents and on-call
n
- Operator-side experience with at least one of Kafka, HDFS or Kubernetes
n
- Python experience, ideally for infrastructure automation or operational tooling
n
- An understanding of networking, storage, processes, memory and system performance
n
- Experience with infrastructure automation, CI/CD or configuration management
n
n
Experience with Kafka or HDFS administration is particularly valuable, but you do not need to know the entire technology stack.
n
Engineers coming from SRE, infrastructure, platform engineering, systems engineering or production engineering backgrounds are encouraged to apply.
📌 Site Reliability Engineer - Data Infrastructure (New South Wales)
🏢 Tribus
📍 New South Wales