The AI Feature Reality Screen
Screen an AI feature before committing to build
We are considering building this AI feature: [THE FEATURE, WHO IT IS FOR, THE WORKFLOW IT TOUCHES]. Screen it before we commit: what data we actually have to power it (not what we wish we had), the eval harness - the 20 real cases we would test before shipping and what score counts as good, the failure mode that embarrasses us in front of a customer and how we catch it, the unit cost per use at our volume, and the honest answer to 'would a user trust this twice after one bad output.' End with build, buy, wait, or kill - and the cheapest test if the answer is wait.
Inputs needed
- ■Feature description, target user, workflow
What good output looks like
Data inventory is real, not aspirational; eval set is 20 named cases with a pass bar; worst failure mode named with its catch; a build/buy/wait/kill call with the cheapest test attached
The stream (0)
Reading the room…