A finished analysis tells you almost nothing anymore. A strong analyst and someone who pasted your task into an LLM hand in the same deliverable — the tell is whether they can explain it.
The candidate solves an actual problem with an in-app AI assistant — no lockdown, paste allowed. They work the way they really would.
Every prompt, every notebook cell, every revision is captured as they go — the work, not the person.
Directing vs delegating, whether they caught the planted error, and a clear Hire / Borderline / No — with the proof one click away.
We plant a plausible-but-false claim in the AI's output. Strong candidates notice it, check it, and remove it. Whether they caught it is the first thing on every report.
Every prompt is classified: did they direct the AI through their own process, or hand over the thinking? The ratio is the senior-vs-junior tell.
No proctoring, no full-screen jail, no keystroke surveillance.
Prompts, cells, and revisions — the process, not the person.
Reviewers can score the work with identifying details hidden.
Enterprise-grade data handling, audited and on the roadmap.