06 Aug
|
Red Yellow Blue
|
Melbourne
06 Aug
Red Yellow Blue
Melbourne
Careers/Open role ContractMelbourne (mostly in office, some remote) $1,300–1,700/day AI-Ops Engineer Senior engineer who runs the AI in production. Evals, observability, cost caps, prompt evolution, and the unglamorous work that turns a working demo into a system a mid-market business can trust for years.
About The Role Most AI systems in the wild are one hallucination away from an incident and one runaway agent night away from a $40k cloud bill. Someone has to make them not do that. That's this role. You own the operational side of the AI systems RYB ships into mid-market clients, and the internal tooling (the brain) that makes those systems observable, cost-capped, and safe to leave running unattended.
You'll pair with the Fractional CIO Partner on strategy and with the Mid-level AI Engineer on build. A Typical Week Eval design and running — every workflow we ship has a real eval harness before it goes live and after every prompt change. You build them, you run them, you set the pass bar
Observability and cost governance — Langfuse, LangSmith, or similar tracing across every agent call, plus per-workflow token budgets and hard caps so nothing runs away
Prompt evolution — production-grade prompt versioning, A/B testing, and rollback discipline. Not "vibes-based" changes
Incident response — when an agent produces something wrong, you're the person who reads the trace, finds the failure mode, and ships the fix
Client hand-over — teaching a mid-market client's own team how to operate what we've built after we walk away
Internal R&D; on the brain — RYB's mission-control system that runs our own agents, integrations, and monitoring You're not the person selling the engagement. You're not the person deciding what to build. You're the person the client leans on when the system needs to work reliably for the next three years. The tech stack Three platforms cover the bulk of what you'll work with: Claude Code — how we ship.
Custom
Skills, MCP servers, agentic workflows, and internal build velocity. It's the primary agent runtime for both our own brain and most client engagements.
AWS — the operational hub. Lambda,
EventBridge, Secrets Manager, IAM, CDK. Every RYB brain agent and every client-facing integration runs here.
Microsoft — the client side. Most mid-market AU clients run M365 + SharePoint, and increasingly Copilot. You'll integrate with M365 admin, SharePoint as a RAG source, Azure OpenAI where Anthropic isn't the fit, and Power Automate where it beats Lambda. Observability sits alongside these — Langfuse, LangSmith, Arize Phoenix, Braintrust. Pick your poison, have a view.
What we're looking for 5–8 years in production engineering — TypeScript/Node, Python, or similar. You've been on-call. You've written the runbook.
Direct production LLM experience — at least one system you've shipped where the LLM was the failure surface you had to design around, not a nice-to-have
Claude Code fluency — or the confidence you'll be productive in it inside a fortnight.
Bonus: shipped custom Skills or MCP servers to production
AWS depth — Lambda, EventBridge, Secrets Manager, IAM. CDK a strong plus. The RYB brain runs on AWS and stays there
Microsoft 365 / Azure familiarity — M365 admin, SharePoint, Azure OpenAI, Copilot readiness. You don't have to be an MVP, but you've deployed into a Microsoft tenant more than once
Practical experience with eval design and observability — you know why "vibes-based" testing doesn't scale
Cost intuition — you cap workflows in token budgets, tier models by task, and don't ship a reasoning-model loop without knowing what it costs
Melbourne-based with right to work in Australia
Bonus: LangGraph, agentic prompt patterns, prior consulting experience How we work Client-first — when there's client work happening, the client's calendar wins
Boring reliability over cleverness — an agent that runs quietly for eleven months beats one that's brilliant on demo day
Markdown-first internally; runbooks are markdown; incident reports are markdown
Honest assessments over polished pitches — read the forms paradox and shadow AI for the operating principles
Pair on hard things; ship solo on the rest
Code review is non-optional and not a status game What we offer Day rate genuinely at the top of the AU senior AI engineering band
Real production responsibility — your systems run in client environments people bill from every day
Direct line to the fractional CIO on strategy and to the engineering team on build. No PowerPoint hierarchy in between.
Mix of clients across financial services, professional services, real estate, mid-market manufacturing
Pipeline currently supports 2–3 days/week ongoing, growing to 3–4 days as more discovery engagements land in FY26-27 What this is not Not a research seat. If your favourite question is "how does this model behave on a novel benchmark", RYB isn't the place.
Not a "prompt engineer" role. Prompt work is a slice of the job, not the whole thing.
Not a fractional CIO seat — if you want to be in the room owning the client relationship, look at the Fractional CIO Partner role.
How to apply Apply Through This Page.
Tell Us The most production-incident-driven LLM system you've shipped — what went wrong, how you found out, what you changed
Your current view on the observability stack you'd pick for a mid-market shop starting from zero, and why
One eval harness you've built and what it caught in the wild A short note on the highest per-request cost workflow you've been responsible for and how you brought it down CV / LinkedIn optional but appreciated. We'll always do a coffee before any commitment.
Apply
We read every application. Required fields are marked with *. First name *Last name *Email *PhoneLocation (city)LinkedIn URL Pitch *30–6,000 characters. Markdown OK — we'll see it as written.CV (optional PDF, DOC, or DOCX
5MB max) Website (leave empty) ← All open roles
📌 AI-OPS ENGINEER (Melbourne)
🏢 Red Yellow Blue
📍 Melbourne