Applied ML / LLM Engineer — STEM Grading & Evaluation (Melbourne)

Applied ML / LLM Engineer — STEM Grading & Evaluation (Melbourne)

06 Aug
|
Mark My Words
|
Melbourne

06 Aug

Mark My Words

Melbourne

Mark My Words gives teachers instant, rubric-aligned feedback on student work, with the teacher always in the loop. We're now extending beyond writing intomaths, science, and other STEM subjects. You'll own the grading and evaluation system for these new subjects end to end — pipeline, models, evaluation, and the self-learning loop that improves it.

The grading machinery is largely subject-agnostic, so you'll build on our existing platform rather than start from scratch. The new opportunity in STEM: unlike writing,maths and science answers are often verifiable— you can check them with a computation/tool step, which lets us push grading reliability higher than is possible in writing.
As an early hire on a small team, you'll have broad ownership and wear many hats.

About the role

Mark My Words gives teachers instant, rubric-aligned feedback on student work, with the teacher always in the loop. We're now extending beyond writing intomaths, science, and other STEM subjects. You'll own the grading and evaluation system for these new subjects end to end — pipeline, models, evaluation, and the self-learning loop that improves it.

The grading machinery is largely subject-agnostic, so you'll build on our existing platform rather than start from scratch. The new opportunity in STEM: unlike writing,maths and science answers are often verifiable— you can check them with a computation/tool step, which lets us push grading reliability higher than is possible in writing.
As an early hire on a small team, you'll have broad ownership and wear many hats.

What you'll do

- Build the STEM grading pipeline: ingest student work (handwritten or typed), compare against the school's rubric and a reference solution, assign scores with partial credit, and flag low-confidence cases for teacher review.

- Add averification layer — tool/code execution (e.g. symbolic math) so the system checks computation rather than guessing it.

- Own evaluation:



build eval sets from real teacher-graded work and track agreement-with-teacher metrics on every change.

- Maintain the teacher-correction loop and turn it into the self-learning data flywheel.

- Work with our handwriting-recognition/OCR pipeline so recognition errors don't get mistaken for student errors.

- Keep grading aligned to curriculum standards (incl. NAPLAN-style and school frameworks).

- Stand up efficient model serving and keep unit economics sane (route easy cases to small models, elevate hard ones).

- Fine-tune open models (LoRA/PEFT, RL where it helps) once we've collected enough labelled data — after prompting + pipeline hit a explicit ceiling.

- Monitor for drift, bias, and silent failures in production.

Tech stack

- Core: Python, PyTorch

- Models & training: Hugging Facetransformers,peft(LoRA/QLoRA),trl(SFT / RL fine-tuning),datasets

- Serving: vLLM

- Supporting: open-weight LLMs (e.g. Qwen family), RAG/retrieval tooling, OCR / handwriting recognition, evaluation harnesses, Docker, cloud

What we're looking for

- Strong Python and hands-on LLM experience beyond API calls — fine-tuning, serving, or structured pipelines.

- Familiarity with the Hugging Face ecosystem (transformers,peft,trl) and an inference engine like vLLM.

- A real understanding thatevaluation is the hard part — you've measured model quality against ground truth, not shipped on vibes.

- Pragmatism: comfortable starting with prompting + pipelines + tool use, reaching for fine-tuning only when data and need justify it.

- Comfort owning something end to end on a small, fast-moving team.

Nice to have

- RAG / retrieval and vector stores.

- OCR / handwritten-input processing.

- Background in education, assessment, or another high-stakes evaluation setting.

- Systematic, eval-driven prompt engineering.

Why join

Own a core part of the product as we expand into STEM, work with real schools, and build something that genuinely helps students and teachers. Real ownership, fast iteration.

#J-18808-Ljbffr

📌 Applied ML / LLM Engineer — STEM Grading & Evaluation (Melbourne)
🏢 Mark My Words
📍 Melbourne

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: applied ml / llm engineer — stem grading & evaluation (melbourne) / melbourne

Subscribe to this job alert:

Get the latest job offers by email for: applied ml / llm engineer — stem grading & evaluation (melbourne) / melbourne