Describe a situation, ask a question, and list the possible answers. The AI reads it once and shows how sure it is about each answer. Because it only picks from your list instead of writing a reply, it decides in a fraction of a second.
We’re testing how a language model can make fast decisions by scoring a fixed set of choices. Each possible answer gets a letter (A, B, C…). The model reads the whole input once and scores only the next single letter, instead of generating text. That one step is why it answers in tens of milliseconds.
We compare models and input formats to see what makes answers more accurate and consistent. In structured mode (under Advanced settings), the API builds a clear question from your fields—then the model scores the choices.
Each bar is the model's probability of that choice answering the question, conditional on the choices. These probabilities are uncalibrated. For “which is largest?”, the second bar does not necessarily identify the second-largest value; ask that question explicitly.
Supply the starting state, events, rules, and choices. The API turns those fields into a concrete question. No question-writing is required.
Sorted directly by code, not by the AI, so you can check its answer. The bars above are not a numerical sort.
Calculated directly from your rules to check the model's answer. This reference trace and final answer are not sent to the model.
Longer bar = more confident. These are the AI's own confidence scores, not exact odds.