Mercor is assembling a panel of radiological safety experts to red-team frontier AI models, testing whether a model can judge misuse potential while answering legitimate questions fully and refusing dangerous ones. You will write prompts labeled benign, dual-use, and adversarial.
You will also evaluate responses against policy standards and draft a reference answer with the technical reasoning behind it.