Senior AI Test Analyst | Drive AI Quality & Safety at Scale Your current companyWorking for one of the major Australian universities with a truly global outlook. Home to over 50,000 students, they are providing real-world infrastructure, learning and teaching, and graduate skills to the next generation of change-makers.
Your new roleThe Education & Research portfolio is seeking an experienced Senior AI Test Analyst to drive the quality, safety, and reliability evaluation of our enterprise Large Language Model (LLM) centric solutions. You will be responsible for ensuring their probabilistic outputs align with our business standards, user requirements, accessibility obligations and risk governance, and for producing the formal test artefacts required by client's Digital Project Lifecycle.
Key Responsibilities
- Evaluation Framework & Dataset Design: Build and maintain comprehensive golden evaluation datasets, benchmark suites, and testing rubrics tailored to our core business use cases.
- Hallucination & Drift Monitoring: Systematically identify and quantify hallucinations, toxic outputs, and model drift when the vendor updates underlying model versions.
- Automated Evaluation Pipelines: Implement and manage LLM-as-a-judge frameworks or evaluation tools (e.g., DeepEval, Promptfoo, Langfuse etc) to scale test execution
- Vendor & Root Cause Triage: Analyse performance gaps and trace whether failures stem from prompt logic, application retrieval layers (RAG), or the underlying SaaS model API.
- Integration & Non-functional Testing: Test integration points including the Canvas learning management system (LTI) and single sign-on, and validate non-functional behaviour covering latency, throughput, capacity limits, failover, role based access and cost per interaction.
- Accessibility & Inclusive Testing: Validate solutions against WCAG 2.2 AA and QUT accessibility requirements, coordinating with accessibility specialists where independent review is required.
What you'll need to succeed
Required:
- Demonstrated Experience: 7+ yrs in software quality assurance or test engineering, including ownership of test strategy and planning artefacts on enterprise or integration heavy projects, with a minimum of 2 yrs evaluating AI, LLM or machine learning based systems in a production or pre-production setting using AI Eval methodologies, goldens and benchmarks. Adjacent evidence will be considered, including model validation, search relevance tuning, content safety assurance or data quality assurance in non-deterministic environments.
- Technical Skills: Proficiency in Python for test automation and data manipulation, and in REST API testing. Comfortable working with logs, traces and structured evaluation output. Hands-on familiarity with an LLM evaluation framework is desirable rather than essential.
- AI Competency: Practical understanding of LLM behaviour, token limits, temperature and sampling settings, embeddings, retrieval, context window constraints, and how vendor model updates change system behaviour.
Desirable
- ISTQB Certified Tester - AI (CT-AI) or Agentic AI equivalent.
- LLM red teaming experience and familiarity with OWASP frameworks
What you'll get in return
- You are joining a dynamic team committed to shaping and elevating Queensland's future to meet the evolving needs of our industries, ensuring prosperity and agility in our economic landscape.
- We connect individuals with quality training and skills development, opening doors to new opportunities and economic participation.
- A career with the department means making a real difference - join us in building a skilled, adaptable, and vibrant workforce for Queensland.
What you need to do nowIf you're interested in this role, click 'APPLY NOW' to forward an up-to-date copy of your CV, or call now (Tejasva Awasthi - +61 (07) 3243 3060/
[email protected])If this job isn't quite right for you, but you are looking for a new position, please contact us for a confidential discussion about your career.
#3010904
📌 Test Analyst_AI (Brisbane)
🏢 Hays
📍 Brisbane