Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format
When you ask an AI to predict what a customer will do, how you ask changes the answer. Reading the score the model assigns internally turns out to rank outcomes more accurately than asking it to write out its reasoning and then answer.
Technical abstract & authors
Language models fine-tuned on customer behavior can predict outcomes and generate explanations, but these readouts are often treated as interchangeable. Holding model checkpoint and prompt content fixed, we compare probabilities obtained by scoring answer tokens with predictions generated after a written rationale. Across 13 model-domain cells covering four retail tasks in three markets, including two using fully public data and checkpoints, the scored readout ranks outcomes more accurately in 12 of 13 cells (two-sided sign test, p approximately 0.003), by 1.5 to 14.5 points in area under the receiver operating characteristic curve (AUC). We interpret these differences through the objectives matched by each readout, identify training choices that narrow the gap, and propose retaining generated rationales while sourcing ranking from the scored head.






