21 Sep
|
seek.com.au
|
Queensland
21 Sep
seek.com.au
Queensland
Job Description
We are looking for a practical, production-minded ML Ops Engineer to own the data, infrastructure and operational pipelines that support our machine learning systems, with particular ownership of the production ML lifecycle for computer vision.
n
This is an end-to-end ownership role: from image, video and related data ingestion and dataset creation through training infrastructure, deployment, monitoring, retraining and production reliability.
n
The Opportunity
n
You will build and operate the platform that enables computer vision models to move from experimentation into reliable production use. You will:
n
n
- design and operate scalable data ingestion, transformation and processing pipelines
n
- build robust dataset creation, validation, labelling and versioning workflows
n
- create reproducible training, evaluation and experimentation pipelines
n
- deploy, version and roll back models safely across live environments
n
- establish monitoring, alerting, drift detection and automated retraining loops
n
n
You will work closely with ML and computer vision engineers, data engineers, software engineers and domain experts, while remaining accountable for the reliability and performance of the end-to-end ML lifecycle.
n
What You Will Own
n
This role sits at the intersection of ML engineering, data engineering, platform engineering and production operations.
n
Your ownership will include:
n
n
- data ingestion, storage, lineage, quality and lifecycle management
n
- workflow orchestration for labelling, training, evaluation and inference
n
- reproducible environments, experiment tracking and model registries
n
- deployment pipelines,
release controls and workplace management
n
- system observability, cost, latency, throughput and reliability
n
- production feedback loops, incident response and continuous improvement
n
n
You will have the autonomy to choose the right tools, simplify brittle processes and build the operational foundations that allow the wider team to ship with confidence.
n
What We Are Looking For
n
Ideally, you will have 3-5 years of experience building and operating ML or data systems in commercial or other real-world production environments. We are looking for evidence of systems used by real operators or customers, rather than experience gained primarily through academic research.
n
You are comfortable working with:
n
n
- large, imperfect and continuously changing datasets
n
- cloud infrastructure, containers and infrastructure as code
n
- ambiguous requirements and fast iteration cycles
n
- operationally critical systems where failures must be visible and recoverable
n
n
You move quickly, make pragmatic trade-offs and take responsibility for outcomes. You are likely stronger at building dependable systems than presenting elaborate architecture.
n
Technical Background
n
You will probably have strong experience across several of:
n
n
- Python,
SQL and production software engineering practices
n
- workflow orchestration and distributed data processing
n
- ML lifecycle tooling, experiment tracking and model registries
n
- containerisation and orchestration, including Docker and Kubernetes
n
- cloud platforms, infrastructure as code and CI/CD for ML systems
n
- object storage, dataset versioning, data validation and lineage
n
- monitoring, observability, alerting and model or data drift detection
n
- exposure to computer vision, video processing or GPU-based ML workloads
n
n
experience with computer vision - particularly moving-object video, tracking, temporal context or trajectory analysis - is a bonus, as is a strong interest in developing in this area. The primary requirement is ownership of production ML and data pipelines.
n
We care more about judgement, execution speed and production ownership than academic prestige or a specific technology stack.
n
What Success Looks Like
n
Within the first few months, we would expect:
n
n
- reproducible training and evaluation workflows with clear dataset and model lineage
n
- a dependable model release process with testing, versioning and rollback controls
n
- monitoring and alerting that identify data, model and system issues early
n
- clear ownership, runbooks and automated recovery or retraining processes
n
n
Success in this role is measured by the speed, reliability and repeatability of the full ML lifecycle - and by how confidently the team can move models into production.
n
Compensation
n
Compensation is flexible and designed to attract exceptional individuals.
📌 ML Ops Engineer (Queensland)
🏢 seek.com.au
📍 Queensland