02 Sep
|
Important Business
|
Sydney
02 Sep
Important Business
Sydney
Production ML Systems | Data Pipelines | Deployment & Reliability
We are looking for a practical, production-minded ML Ops Engineer to own the data, infrastructure and operational pipelines that support our machine learning systems, with particular ownership of the production ML lifecycle for computer vision.
This is an end-to-end ownership role: from image, video and related data ingestion and dataset creation through training infrastructure, deployment, monitoring, retraining and production reliability.
The Opportunity
You will build and operate the platform that enables computer vision models to move from experimentation into reliable production use. You will:
- design and operate scalable data ingestion, transformation and processing pipelines
- build robust dataset creation, validation, labelling and versioning workflows
- create reproducible training, evaluation and experimentation pipelines
- deploy, version and roll back models safely across live environments
- establish monitoring, alerting, drift detection and automated retraining loops
You will work closely with ML and computer vision engineers, data engineers, software engineers and domain experts, while remaining accountable for the reliability and performance of the end-to-end ML lifecycle.
What You Will Own
This role sits at the intersection of ML engineering, data engineering, platform engineering and production operations.
Your ownership will include:
- data ingestion, storage, lineage, quality and lifecycle management
- workflow orchestration for labelling, training, evaluation and inference
- reproducible environments, experiment tracking and model registries
- deployment pipelines,
release controls and environment management
- system observability, cost, latency, throughput and reliability
- production feedback loops, incident response and continuous improvement
You will have the autonomy to choose the right tools, simplify brittle processes and build the operational foundations that allow the wider team to ship with confidence.
What We Are Looking For
Ideally, you will have 3-5 years of experience building and operating ML or data systems in commercial or other real-world production environments. We are looking for evidence of systems used by real operators or customers, rather than experience gained primarily through academic research.
You are comfortable working with:
- large, imperfect and continuously changing datasets
- batch, streaming or event-driven data processing workflows
- cloud infrastructure, containers and infrastructure as code
- ambiguous requirements and fast iteration cycles
- operationally critical systems where failures must be visible and recoverable
You move quickly, make pragmatic trade-offs and take responsibility for outcomes. You are likely stronger at building dependable systems than presenting elaborate architecture.
Technical Background
You will probably have robust experience across several of:
- Python,
SQL and production software engineering practices
- workflow orchestration and distributed data processing
- ML lifecycle tooling, experiment tracking and model registries
- containerisation and orchestration, including Docker and Kubernetes
- cloud platforms, infrastructure as code and CI/CD for ML systems
- object storage, dataset versioning, data validation and lineage
- monitoring, observability, alerting and model or data drift detection
- exposure to computer vision, video processing or GPU-based ML workloads
Experience with computer vision - particularly moving-object video, tracking, temporal context or trajectory analysis - is a bonus, as is a strong interest in developing in this area. The primary requirement is ownership of production ML and data pipelines.
We care more about judgement, execution speed and production ownership than academic prestige or a specific technology stack.
What Success Looks Like
Within the first few months, we would expect:
- stable, observable data pipelines supporting live ML workloads
- reproducible training and evaluation workflows with clear dataset and model lineage
- a dependable model release process with testing, versioning and rollback controls
- monitoring and alerting that identify data, model and system issues early
- clear ownership, runbooks and automated recovery or retraining processes
Success in this role is measured by the speed, reliability and repeatability of the full ML lifecycle - and by how confidently the team can move models into production.
Compensation
Compensation is flexible and designed to attract exceptional individuals.
📌 ML Ops Engineer (Sydney)
🏢 Important Business
📍 Sydney