Mercor is assembling a panel of radiological safety experts to red-team frontier AI models, testing whether a model can judge misuse potential while answering legitimate questions. You will write challenging prompts at benign, dual-use, and adversarial levels, evaluate responses to a policy standard, and draft the reference answer with the technical reasoning behind it.
This requires explicit, written justification for each judgment.
#J-*****-Ljbffr