20 Sep
|
Talenza
|
New South Wales
20 Sep
Talenza
New South Wales
Job Description
Site Reliability Engineer - Data Platform
n
Sydney | Global Trading Firm
n
Join a global trading firm's high-performing Data Engineering team and help operate the large-scale platform that underpins trader research, simulation, reporting and decision-making.
n
This is a proper infrastructure role-not one for someone who has only built data pipelines on top of managed services. You'll get hands-on with the underlying platform: Kafka, HDFS, Dremio, Linux and in-house data tooling across a multi-petabyte workplace processing around two million queries each day.
n
You'll join a small, experienced Sydney team with international engineering counterparts, giving you genuine ownership, strong mentoring and exposure to complex distributed systems at meaningful scale.
n
The role
n
Run, monitor and improve large-scale data platforms including Kafka, HDFS, Dremio and internally built pipelines.
n
Troubleshoot real production issues across Linux, storage, networking and distributed infrastructure.
n
Build automation and CI/CD capability to make deployments faster, safer and more repeatable.
n
Support upgrades, capacity planning, incident response and long-term reliability improvements.
n
Work closely with systems and network engineers, developers, researchers and end users to solve complex data-platform problems.
n
Help evaluate and introduce new technology as the environment continues to evolve.
n
What we're looking for
n
Around 2-3 years' experience in an SRE, platform, infrastructure, systems or production engineering role.
n
Strong Linux troubleshooting skills across processes, filesystems, networking, disk and memory pressure.
n
Hands-on operator experience with at least one of Kafka, HDFS or Kubernetes-you have deployed, configured, upgraded, tuned or supported the platform itself.
n
Experience operating self-managed infrastructure, whether bare metal, datacentre, self-run VMs or self-managed Kubernetes.
n
Python experience for systems automation, operational tooling, health checks or deployment workflows.
n
Exposure to Docker, Kubernetes, Helm and infrastructure-focused CI/CD.
n
A curious, pragmatic mindset and a genuine interest in understanding how complex systems behave under pressure.
n
Experience with Dremio, Presto, Airflow, Prefect, Ansible, Puppet, Terraform or cloud platforms would be beneficial, but it is the operational mindset and underlying
Linux/infrastructure
depth that matter most.
n
This is a standout opportunity for an engineer who wants to move beyond managed services, get close to the underlying technology and build a career operating high-scale, business-critical systems.
n
Site Reliability Engineer - Data Engineering Sydney, NSW, AU
📌 Site Reliability Engineer - Data Engineering (New South Wales)
🏢 Talenza
📍 New South Wales