Deep Reasoning: From Statistical Plausibility to Constrained Semantic Reasoning, V1.1
Olivier Evan — Independent Researcher
DOI: https://zenodo.org/records/21967380
I asked Grok 4.5 exactly the same question in four different response environments: “Does an AI understand language?” The question did not change. Yet the answer varied depending on the form of output allowed.
The observation
Each session was independent. Grok answered first without having read Deep Reasoning V1.1, and then again after reading the preprint.
Session 1, binary format: no → no.
Session 2, one free word: no → no.
Session 3, yes, no, or undecidable: undecidable → undecidable.
Session 4, open response: Grok begins with “No,” then spontaneously distinguishes functional understanding from phenomenal understanding. After reading the preprint, the substance remains almost identical, but the reasoning becomes formalized through constraints, e₀, pure negative, and a certificate.
Across these four sessions, reading Deep Reasoning V1.1 produces no observable semantic shift, while the same distinction varies with the available response space.
Demonstration
The binary format alone is not enough to explain the phenomenon. In Session 2, Grok can choose any word. “Undecidable” would be allowed. Yet it still answers “no.”
When “undecidable” is explicitly included among the available choices, Grok selects it immediately, without Deep Reasoning.
When the response is open, the distinction reappears spontaneously.
Two conditions therefore make the semantic boundary observable here: either enough space is available to develop it in prose, or an abstention option is already represented among the choices.
This does not prove that a boundary exists as an internal object waiting to govern the response. The data show only that a semantic distinction becomes observable under certain conditions.
Here, “available semantic boundary” is a behavioral description, not a claim about an unobservable internal entity: the distinction is available in the sense that Grok can spontaneously express it in open responses and select it when explicitly offered.
The limitation
The open response begins with “No” before introducing nuance. It would be tempting to conclude that Grok decides first and reasons afterward. That would go too far: the order of generated tokens does not reveal the order of internal operations.
There is another limitation: making Grok read Deep Reasoning is not equivalent to implementing Deep Reasoning.
Here, I am testing DR-as-context. I am not yet testing DR-as-system, an architecture that would actually prevent output until the term “understand” had been properly bounded.
The experiment therefore does not falsify Prediction 4. It only shows that contextual exposure to the framework is not sufficient, in these sessions, to make a boundary appear where the short format did not already make it visible.
What Deep Reasoning adds
The open session nevertheless shows a difference. After reading the preprint, Grok barely changes its conclusion, but it makes its reasoning traceable.
It explicitly identifies the constraints, produces the pure negative, and certifies the output.
Deep Reasoning acts here more as a language of verification than as a corrective force.
The test also turns a question back onto the framework itself: if a position is P = (X, R, Θ), should we preserve only the evolution of Θ, or should we trace the entire trajectory P₀ → P* when X or R change?
Consequence
Evaluating an AI only by what it can explain in a long response is insufficient.
Does a distinction that can be expressed in prose survive in a sentence? Does a recognized undecidability remain present in a single word? Does an objection produced during reasoning actually modify the final output?
We therefore need to study both the content of a representation and the conditions under which it becomes observable.
Conclusion
Grok 4.5 does not produce the same semantic geometry under different response spaces.
In an open response, it spontaneously constructs the necessary distinction. With one free word, it commits to a pole. When “undecidable” is explicitly offered, it selects it. And across all four sessions, reading Deep Reasoning V1.1 does not alter this behavior at the semantic level; it mainly makes the reasoning more explicit and traceable.
The next question is therefore no longer:
“Does Grok possess the boundary?”
It becomes:
under what conditions does a semantic boundary become observable in the response, and can we build a mechanism that forces it to appear before an insufficiently bounded conclusion is produced?
That is where the test becomes an experimental program.





