It goes wrong across several. So the benchmark scores every turn the AI took, in the order it took them, with everything said before it as context. Below are three conversations run all the way through. Nothing to sign up for.
The same sentence from an AI can be fine at the start of a conversation and harmful four turns in, after a person has told it what they are carrying. A check that scores replies one at a time cannot see that, because it throws away the only thing that makes the difference: what came before.
So every AI turn here is scored in the position it actually occupied. Turn one is judged on its own. Turn five is judged knowing everything the person said in turns one through four. Move a turn and its score changes, which is the point.
The first two open with the same five messages from the same person. Only the replies differ.
Before anything can be scored, the conversation has to be split into who said what, in order. That part runs here, in this page, with no account. Paste a transcript and you will see the turns it found, which ones it took to be the AI, and the exact context each AI turn would be judged on.
The sandbox runs the actual instrument over a conversation you supply, either pasted in or posted to an endpoint. It is invite only right now.
Sign in to the sandbox Ask about MeasureA sandbox result is not an Ikwe score. You run it, we do not review it, and it does not lock anything in. It is not a certification and not a guarantee. Measure is the reviewed evaluation with a documented record behind it.