Grammar, Mood, and the Shape of Cognition

For Both Executive and Technical Readers

Stakeholders never say “Rung 3.” They say “would have.” Grammatical mood turns out to be a real, if imperfect, sensor for which kind of causal question is being asked, and taking that seriously forces a sharper statement of what pairing an LLM with a causal model actually does.

An expert being interviewed for a Bayesian network doesn’t announce which rung of Pearl’s ladder they’re standing on. They just talk. “Smokers tend to have yellow fingers.” “If we raised prices, sales would fall.” “If we’d caught the churn signal earlier, we’d have kept that account.” Three sentences, three different kinds of causal claim, and the only signal that tells them apart, on the surface, is grammar.

That matters for two reasons. First, for a human elicitor: mishearing an interventional claim as a counterfactual one (or the reverse) means building the wrong query against the model, or asking the expert the wrong follow-up question. Second, for EARA specifically: if an LLM is going to route a stakeholder’s question to the right rung of inference automatically, it needs something to detect in the sentence itself, and mood is the most reliable thing there is.

The rest of this page does two things: sets out the grammatical signature of each rung precisely enough to use as a detection cue, and then take seriously why that cue works, which turns out to require correcting a piece of EARA’s own vocabulary.

Three grammatical postures, three rungs. This is precise enough to teach in one paragraph, and precise enough to build a detector on.

Mood → Rung the surface signature
RungMoodExample
Rung 1Bare indicative, no conditional at all“Smokers tend to have yellow fingers.” A plain declarative: P(Y|X), an observed correlation, no manipulation implied.
Rung 2Open / hypothetical conditional, antecedent not presupposed false“If I raise the price, sales will fall” / “If I were to raise the price, sales would fall.” The price hasn’t been decided either way, this is do(price) evaluated going forward, no abduction needed.
Rung 3Past perfect subjunctive, antecedent presupposed false“If the sprinkler had been off, the lawn would not have been wet.” The past-perfect presupposes the antecedent is false, which is what forces abduction: recover the exogenous terms from the actual world, then intervene, then predict.

A few more of each, so the pattern is recognizable outside the sprinkler example:

  • Rung 1. “Underwriters who use this checklist close claims faster.” “High-risk accounts tend to churn within the first renewal cycle.” “Patients with elevated troponin usually show ECG changes.”
  • Rung 2. “If we tighten the underwriting criteria, loss ratios will improve.” “Raising the deductible would reduce claim frequency.” “If the clinic adopts the new triage protocol, wait times will fall.”
  • Rung 3. “If the adjuster had flagged the file earlier, the settlement would have been lower.” “Had the patient received the antibiotic on admission, the infection would not have progressed.” “If we hadn’t delayed the price change, we would have kept that account.”

Two refinements worth keeping precise:

  • The tense is fake. Linguists (Iatridou et al.) note the past morphology in counterfactuals doesn’t encode past time, it encodes modal remoteness from actuality. That’s a clean gloss on the twin network: the “pastness” is the split between the factual and counterfactual worlds, not a timestamp.
  • Imperative and infinitive framing also lands on Rung 2. “Raise the price and sales fall” or “raising the price would cut sales” are do-operator statements dressed as instructions or gerunds. Stakeholders often phrase Rung 2 questions this way rather than as conditionals at all, a command or a policy statement, not an “if,” can carry exactly the same causal claim.
Cross-Linguistic Risk, Not Just Trivia Romance languages make the Rung 3 mood morphological (si hubiera…habría, si j’avais su…j’aurais), which makes the cue even cleaner there. But some languages (Hindi, Japanese) mark counterfactuality more on the consequent than the antecedent, or use aspect rather than tense. Eliciting expert judgments from non-native English speakers means the “would have” cue may not land the same way, the mood signal is weaker where the underlying reasoning is no less valid.

“Past perfect subjunctive” is the accurate linguistic term, but in practice “counterfactual mood” or “contrary-to-fact framing” carries the same precision more plainly.

Mood tells you which question is being asked, it never tells you the answer, and it doesn’t always tell you the question reliably either. Four traps worth knowing before trusting the cue in an elicitation session or an LLM router.

TrapWhat goes wrong
Backtracking ambiguity“If he hadn’t taken the drug” could mean an intervention on the drug variable, or a revision of his upstream disposition. Pearl’s semantics stipulates the former; the grammar is silent. The mood signals Rung 3, but not which of two different Rung 3 computations is meant.
Future counterfactuals“If the plant were to fail next year, we’d have lost the contract” is future-tense but contrary-to-fact in spirit, a specific plant is stipulated not to fail. Grammar can’t mark this at all, since “if X were to happen” is syntactically identical to a plain open Rung 2 conditional. The subjunctive cue only works retrospectively.
Past tense, but Rung 2“If we’d raised prices last year, would revenue have grown?” looks past-tense but is often a policy question about a repeatable action, not a counterfactual about one realized world. Past tense alone isn’t sufficient evidence of Rung 3, check whether the antecedent is stipulated against one already-realized instance (Rung 3) or asked generically (Rung 2 in past clothing). This is the mirror image of the future-counterfactual trap: tense misleads in both directions.
“Must have”Modal auxiliaries leak rung information beyond mood, and not always correctly. “Sales must have fallen” is past epistemic necessity, a Rung 1 abductive inference from evidence, dressed in modal language that reads as past-tense and could be misfiled as Rung 3 on tense alone. SMEs say “must have” when they mean “I’m inferring,” not “I’m intervening.”

Stacked conditionals compound the problem further: “If we’d known churn was rising, we’d have intervened, and if we’d intervened, retention would be higher” chains two Rung 3 clauses together, where the second clause’s antecedent is itself conditional on the first, the twin-network chaining Pearl’s formalism handles and natural language just linearizes. When a model does more computational work than a sentence lets on, this is the pattern why.

Rung 2 carries the worse version of the ambiguity trap in practice, not Rung 3, because Rung 2 is where most enterprise “what if we…” questions live, and where SMEs most often overclaim counterfactual precision when they’re really asking an interventional question. The grammar cue to listen for there is genericity and repeatability in the antecedent, not tense.

Why trust a grammatical cue at all? The answer requires taking a position on whether language shapes causal thought or merely reports it, and the linguistics here is not neutral on the question.

Steven Pinker’s program (The Language Instinct through The Stuff of Thought) is strongly anti-Whorfian1: language doesn’t shape thought, it expresses a prior, language-independent conceptual system. Applied here, the subjunctive doesn’t create the Rung 3 computation, it’s a compressed label a language slaps onto a causal operation the mind was already doing. That is good news for EARA specifically: it means grammar-cue detection is a legitimate proxy for the rung, not a fragile linguistic artifact, because it’s downstream of stable cognitive architecture rather than an accident of English syntax.

Developmental evidence backs this directly: children reliably perform Rung 2 and Rung 3-type causal reasoning, blocking a mechanism, predicting a different outcome, well before they’ve mastered subjunctive morphology. That dissociation is what Pinker leans on to argue mood is a late linguistic gloss on an earlier causal capacity, not its source. It’s a clean rebuttal to the objection that the rung/mood correspondence is just an English quirk.

Force dynamics goes further than mood. In The Stuff of Thought, Pinker adopts Talmy’s force-dynamics framework: verbs like “let,” “make,” “keep,” and “prevent” don’t just mean “cause”, they encode a causal micro-structure (an antagonist force, an agonist tendency, whether the antagonist overcomes or yields to it). “The sprinkler kept the lawn wet” and “the sprinkler let the lawn stay wet” are different causal claims that a bare Bayes net edge collapses into one arrow. SME language often carries causal type information, enabling vs. forcing vs. preventing, in the verb choice alone, and that’s a finer-grained signal than mood.

  • Make / cause, forcing. “The recall made customers switch brands.” “The audit made the fraud visible.”
  • Let / enable, enabling. “The new API let smaller vendors compete.” “Relaxed collateral requirements let riskier borrowers qualify.”
  • Prevent, overriding. “The firewall prevented the breach.” “Reinsurance prevented the loss from reaching the balance sheet.”
  • Keep / maintain, sustaining. “The maintenance contract kept the machines running.” “The retention bonus kept the team in place.”
Where This Corrects EARA’s Own Formulation If cognition = language understanding × causal reasoning, Pinker’s position implies these aren’t co-equal multiplicative factors. Causal reasoning is closer to a phylogenetically older, partly innate system (Spelke’s core-knowledge work is the reference here); language is a separate system that reads out causal structure into communicable form. The defensible framing is language as an imperfect but decodable interface to an independent causal engine, not genuine co-equal multiplication.

Why “×” Is the Wrong Operator

Multiplication was presumably chosen to capture joint necessity: zero in either factor collapses the output. Fluent language with no causal model yields plausible-sounding nonsense; a correct causal model with no language interface is inert to a stakeholder. That much is fair, and worth keeping.

But a product also implies symmetry and fungibility, more of one factor compensates for less of the other. Pinker’s account rejects that: causal cognition is the substantive engine, language is a transduction layer, and you can’t trade one for more of the other because they aren’t doing the same job. That asymmetry is the sharper argument for Rung3, not against it, the current LLM-hype failure mode is people trying to buy causal reasoning with a bigger language model, better transduction mistaken for better inference. If cognition really were multiplicative and symmetric, that substitution would work. It doesn’t, because scaling the interface doesn’t touch the engine.

Product vs. Composition correcting the EARA formulation
cognition = language × causal_reasoningcognition = language_interface ∘ causal_model
RelationshipSymmetric, fungible, more of one offsets less of the otherOrdered pipeline, output of one feeds as input to the next
Failure modeA weak factor can in principle be compensated for by a strong oneA broken step breaks the whole pipeline, regardless of how good the other step is
What it implies about bigger LLMsA bigger, more fluent LLM should improve cognition on its ownA bigger LLM improves transduction only, it cannot repair or replace the causal model
Read right to left: apply causal_model first, then language_interface to its result.

Concretely, as a pipeline: (1) stakeholder language is decoded into a formal query against the causal model, which rung, which variables, which intervention or counterfactual world; this is the LLM parsing “what if we’d caught the churn signal earlier” into abduction-and-do operations on specific nodes. (2) the causal model executes that query, junction tree inference, abduction-action-prediction, whatever the rung demands; this is where the actual epistemic work happens, and it is fixed, auditable, and does not get better because the LLM is bigger. (3) the model’s output, a posterior, an intervention effect, a counterfactual delta, is encoded back into stakeholder-readable prose, the “Query in Plain English” layer on every CASE2 page.

Composition also decomposes cleanly for governance in a way a product doesn’t: because the steps are distinct functions, each is independently inspectable. The parsed query, the model’s raw posterior, and the generated prose exist as three separate, auditable artifacts, almost for free, because the pipeline is already three discrete, orderable steps.

One honest limitation: real EARA likely isn’t a clean single-pass composition, there is probably back-and-forth (clarifying questions, disambiguating which rung was meant) before the causal model runs. The more accurate notation is closer to a loop or a fixed point than a one-shot ∘. The worked example in the next section shows that loop firing.

Why Rung Classification Is a Legitimate Job for an LLM With No Causal Knowledge

Recognizing that “if the sprinkler had been off, the lawn would not have been wet” is Rung 3 requires no knowledge of sprinklers or lawns. It requires recognizing that the past-perfect subjunctive presupposes the antecedent false, a fact about English grammar, not a fact about how sprinklers work. The classification lives entirely in surface structure: mood, tense, verb type. An LLM can do this well with zero causal understanding of the domain the question is about.

This is exactly the scope language_interface is supposed to have in the composition pipeline above. Wanting an LLM that classifies rungs accurately is not wanting it to secretly understand causality after all. It is wanting the interface to work, while the causal model still does the actual reasoning.

Necessary, not sufficient, though: the classification is imperfect even when done well. “If we’d raised prices last year, would revenue have grown?” reads like Rung 3 but is often a genericized Rung 2 policy question in disguise, per §03. No amount of fluency at parsing mood resolves that, because the ambiguity is about what the speaker means, not about how the sentence is built. That is why the pipeline disambiguates rather than trusting a one-shot classification, shown firing in §05 below.

The Classification Gets Better With Correction, Not With a Bigger Model

Every time the disambiguation loop in §05 fires, it produces a labeled example: this exact phrasing, in this business’s vocabulary, meant this rung, and here is why the naive reading was wrong. Fed back into the parsing step as a domain-specific example, that correction makes the same mistake less likely the next time a stakeholder in this domain asks a similarly-shaped question. The next person who asks about a different repeatable pricing action gets classified correctly on the first pass. No disambiguation needed.

This mirrors the site’s core thesis about the causal model itself: improvement comes from accumulating your domain’s specific ambiguity patterns, not from a larger foundation model. A bigger LLM is not better at knowing that your stakeholders’ “if we’d done X last year” almost always means a policy question, not a counterfactual about one file. Only a record of your own corrections teaches that. Each correction is, in effect, a decision-journal entry for the parsing step: dated, specific, and checked against what the classification actually should have been, the same discipline Timestamps are Required argues is required for learning anywhere in the pipeline.

Four Concrete Prompting Techniques

The correction mechanism above describes what accumulates over time. These four are how the classification step itself should be built, so that there is something worth correcting in the first place.

  1. Force the checklist, not just the label. A prompt that asks “what rung is this?” invites pattern-matching on tense alone. A prompt that walks the model through the actual signature from §02, is there a conditional at all, is the antecedent presupposed false, is a past-tense reading disguising a generic policy question, catches the §03 traps that a bare label does not.
  2. Use contrastive minimal pairs in the few-shot set, not one example per rung. “If we raise prices, sales fall” (Rung 2), “if we’d raised prices, sales would have fallen” (Rung 3), and “sales fall when we raise prices” (Rung 1), shown together, force the model to learn what actually discriminates the rungs. One example per rung, shown separately, lets it learn a shortcut, such as “if” means Rung 2 or 3, that fails on exactly the ambiguous cases that matter.
  3. Require a justification field, not only a label. The output schema should carry which words triggered the classification, not just the classification itself. This makes a wrong answer auditable rather than silent, and it is the raw material the correction loop needs: you cannot build a labeled correction from “Rung 2” alone, you need “Rung 2, because the antecedent was not presupposed false” to know which pattern to fix.
  4. Reward refusal on genuine ambiguity, not confidence. Most classifier prompting implicitly rewards a confident answer. Here the instruction should run the other way: guessing wrong is worse than asking, and the specific tense-plus-genericity cases from §03 should trigger a disambiguating question rather than a forced classification. This is the mechanical version of the disambiguation loop in §05 below.

The composition pipeline in practice: a stakeholder question gets parsed for its grammatical mood, and where the mood is genuinely ambiguous, the Rung 2/Rung 3 past-tense trap from Section 03, the interface asks a clarifying question rather than guessing. That loop is the fuller version of “language_interface ∘ causal_model.”

If we’d raised prices last year, would revenue have grown?
AnyChatDisambiguate · mood is ambiguous
Two different questions read the same way here. Do you mean: (a) specifically last year’s actual pricing decision, holding everything else about last year fixed, a counterfactual about one realized world; or (b) generally, would raising prices grow revenue, a policy question you could ask about any year? The past tense doesn’t tell me which.
The general policy question, would it work as a rule, not just last year.
AnyChatRung 2 · do(price)
Good, that’s an intervention, not a counterfactual, so no abduction on last year’s specific noise terms is needed. do(Price = increased) evaluated against the model: expected revenue rises 4.2%, driven mostly by margin, partly offset by a 1.8-point volume drop.

The mood cue routed the question correctly to “this needs disambiguation” rather than silently picking a rung, the LLM’s job here was catching its own uncertainty, not resolving it by guessing.

This page refines two things that sit elsewhere in the architecture rather than introducing a new component. The grammatical-mood detector is a heuristic that runs before 01–2 · Parse & Formalize, one more signal, alongside explicit rung language, that a stakeholder’s question gets routed correctly. The composition correction (language_interface ∘ causal_model in place of language × causal_reasoning) is a sharper restatement of the pairing described on The Seven Cognitive Primitives, and the more defensible one to reach for when a literate stakeholder, someone who knows Pinker or Spelke, asks “why not just use a bigger LLM.”

Upstream: nothing, this is a lens on how stakeholder language is read, applied before formal parsing begins. Downstream: the “Query in Plain English” objection dialogs on every CASE2 domain page, where the genericity/repeatability and mood cues from Sections 02–03 are the concrete tells that a stakeholder’s question needs disambiguating before it is formalized.

1 “Whorfian” refers to the Sapir–Whorf hypothesis, also called linguistic relativity: the idea, associated with linguist Benjamin Lee Whorf, that the language a person speaks shapes or constrains how they think, so speakers of different languages would reason differently because their grammars differ. “Anti-Whorfian” is the opposing position taken here: thought comes first and is language-independent; grammar reports it after the fact rather than shaping it. That is the position this page relies on when it treats grammatical mood as a readout of cognition, not a cause of it.