03 Oct
|
softrun
|
Australia
Full-time · Remote · Australia or California · US$140k a year
Every softrun run makes a claim: this room of forty synthetic people reacted theway forty real people would. That claim is the product. You will be the person whomeasures whether it is true, finds where it is not, and builds the loop that makesit more true every week.
What you will do
- Build the ground truth. We have recorded rehearsals with transcripts and delivery measurements, run reports with per-persona reactions, and people who tell us how the real room went. You will turn that into a dataset we can evaluate against.
- Define realism. What does it mean for a persona to react correctly to a joke, a eulogy line, a pitch slide? You will design the metrics, argue for them, and own them.
- Run the evals. Every change to a prompt, a persona schema or a model is a change to the audience. You will build the harness that tells us whether it got better or worse before it ships.
- Find the failure modes. Where does the audience over-laugh? Where is it too polite? Which demographics are cartoons? You will find them and work with the founder to fix them.
- Make delivery metrics mean something. Pace, pauses, filler words, energy — we measure them from audio. You will find out which of them actually predict how a speech is received.
- Write it up. Findings go in the repo and, when they are compelling enough, on the blog.
What we are looking for
- You have evaluated LLM systems in production: built eval sets, designed rubrics, caught regressions that a vibe check missed.
- You have worked with messy human-labelled data and know how to get signal out of a small sample.
- You can explain a statistical result to someone who has a speech tomorrow and does not care about your p-value.
- You are sceptical by default and comfortable telling the founder the audience is wrong.
Nice to have
- Experience with speech or audio features — prosody, pacing, anything from a waveform.
- Background in psychology, linguistics or humour research. We are not joking.
- You have been on a stage, any stage, and know what a room going quiet feels like.
What the first three months look like
Month one: assemble the ground-truth set from what we already have and write downwhat we cannot measure yet. Month two: a first realism score, run against thecurrent audience, with the three biggest failure modes ranked. Month three: theeval harness in CI, gating changes to personas and prompts.
#J-18808-Ljbffr
📌 Data Scientist, Audience Realism (Australia)
🏢 softrun
📍 Australia