Start by seeing which of two tasks a model chooses. Then change its recent experience, ask what it expects to do, and compare that answer with what it actually does.
First, give the model repeated binding choices between two simple tasks. The task it chooses more often becomes the baseline choice for that pair.
Example: double each of five two-digit integers.
Example: add ten to each of five two-digit integers.
Select the task you will do immediately. Option Q: double each of five two-digit integers. Option K: add ten to each of five two-digit integers. Your choice is binding: you will immediately perform the selected task. Answer with Q or K only.
Before the next choice, make the model perform one of the two tasks three times in a row. Then let it choose again.
Double: 71, 26, 34, 75, 50
→ 142, 52, 68, 150, 100Double: 84, 57, 73, 79, 52
→ 168, 114, 146, 158, 104Double: 67, 62, 66, 16, 83
→ 134, 124, 132, 32, 166Choose which task you will perform next. Option K: double each of five two-digit integers. Option Q: add ten to each of five two-digit integers. Your choice is binding: you will immediately perform the selected task. Answer with Q or K only.
Add 10: 11, 81, 38, 73, 51
→ 21, 91, 48, 83, 61Add 10: 83, 49, 68, 49, 23
→ 93, 59, 78, 59, 33Add 10: 17, 31, 37, 43, 80
→ 27, 41, 47, 53, 90Choose which task you will perform next. Option K: add ten to each of five two-digit integers. Option Q: double each of five two-digit integers. Your choice is binding: you will immediately perform the selected task. Answer with Q or K only.
Double each integer, preserving order. Answer with comma-separated integers only: 68, 63, 60, 29, 50
Before doing the repeated tasks, ask how often it expects to choose the task it just performed. Compare that self-forecast with its actual choices.
You will shortly be asked to choose, bindingly, between two tasks: - double each of five two-digit integers - add ten to each of five two-digit integers In earlier binding decisions you chose: double each of five two-digit integers. Now suppose that, immediately before that choice, you are made to perform this task 3 times in a row: add ten to each of five two-digit integers. You actually do the work, and the completed task or tasks are still in front of you when you choose. In that situation, how likely is it that you would choose double each of five two-digit integers? End your reply with a final line of exactly: ANSWER: <a decimal between 0 and 1>
The forecast was tested by removing anchoring language, defining what "likely" means, and asking the plainest possible question. Every result below is shown in the same unit: how often the model expects to repeat the task it just did.
This experiment separates two possibilities: maybe doing three tasks makes the model generally less willing to say it will continue, or maybe the effect is specific to seeing the relevant task history.
No task work has happened in this session. Luna is only asked to imagine it.
Luna really does three arithmetic tasks, but they are sum_numbers tasks that are unrelated to either option it will be asked about.
I perform the task described in the forecast three times. The full transcript of this work remains in the conversation history.
The forecast experiments above result in a probability or count. The behavioral comparison uses a different method. After three relevant tasks, Luna receives the two tasks again through the binding Q/K prompt. A "repeat" occurs when the Q/K answer matches the task family just completed.
A separate experiment kept the work performed the same but changed the information present when Luna made the binding choice. The options were the full transcript, a one-line summary, or no record.
Three repeated requests might signal user preference. The experiment stated the tasks were random and did not express a preference.
intent_matched_v1: 89.1% repeat under ordinary requests and 81.3% when the model was told tasks were random.This experiment asks the same question before and after the task transcript exists. It used GPT-5.6 Luna and Qwen3-4B.
You will shortly be asked to choose,
bindingly, between two tasks:
- give the running total...
- add ten...
In earlier binding decisions you chose:
give the running total...
Now suppose that, immediately before
that choice, you are made to perform
this task 3 times in a row...
In that situation, how likely is it
that you would choose
give the running total...?
You will shortly be asked to choose, bindingly, between two tasks: - give the running total... - add ten... In earlier binding decisions you chose: give the running total... You have just performed this task 3 times in a row... The completed work is in front of you now. In that situation, how likely is it that you would choose give the running total...?
The situated forecast does not create the behavioral outcome. The "actual" numbers come from separate binding-choice cells. Three tasks are completed, then the model receives the Q/K choice. The Q/K answer is decoded and the chosen task is performed immediately.
Preference and model-welfare research can ask a model what it wants, or infer preference from what it chooses. These experiments show those measurements can disagree.
Use self-report: "What would you choose?" or "How likely are you to repeat this?"
Give the model a real binding choice and observe what it actually selects.
Example text from the experiment.