Skip to content

What the Claim Splitter returns: three real runs, one with a visible mistake

Project documentation · 4 min read · updated

For every English claim the model writes the same eight sections: claim type, core claim, premises, conclusion, hidden assumptions, confounders or alternative explanations, evidence needed, and inference weaknesses. The three runs below are verbatim, captured on 3 October 2026. The first two read sensibly, with a few garbled lines; the third misstates its own premise, which is the kind of mistake to expect from a 270M-parameter model.

Captured
3 October 2026, desktop Chrome (headless), model loaded from an empty cache in 40.1 s
Settings
Greedy decoding (temperature 0), 2,048-token context, 512 new tokens at most
Time per run
32.2 s, 37.8 s and 31.4 s as reported by the page; one run each, one machine
Edits
None to the wording; some blank lines between headings were removed

What sections does the output have?

HeadingWhat it holds in these runs
CLAIM TYPEA one-line label, such as "Everyday causal claim (correlation presented as causation)"
CORE CLAIMThe claim restated in one sentence
PREMISES, CONCLUSIONWhat the claim offers as support, and what it concludes
HIDDEN ASSUMPTIONSWhat must also be true for the conclusion to follow
CONFOUNDERS / ALTERNATIVE EXPLANATIONSOther things that could produce the same pattern
EVIDENCE NEEDEDWhat would settle it
INFERENCE WEAKNESSESWhere the reasoning is thin

Run 1: a correlation presented as causation

People who drink coffee tend to live longer, which shows coffee extends lifespan.
CLAIM TYPE
Everyday causal claim (correlation presented as causation).

CORE CLAIM
Drinking coffee causes people to live longer.

PREMISES
People who drink coffee live longer.
CONCLUSION
Coffee causes people to live longer.

HIDDEN ASSUMPTIONS
Longevity effects are caused by drinking coffee rather than by other causes.
The conclusion is informed by experimental evidence and does not reverse-engineer causal evidence.
CONFOUNDERS / ALTERNATIVE EXPLANATIONS
Wealth, education, or personality may drive both the drinking habit and longevity.
Reverse causation is possible if the outcome influences the habit rather than the habit itself.
EVIDENCE NEEDED
Experimental or mechanistic support for the causal claim.
Benefit-risk analysis and unintended consequences.
INFERENCE WEAKNESSES
'Drinking coffee' is everyday causal claim without justification.

What works: it names the claim as correlation presented as causation and offers two sensible alternatives, shared causes (wealth, education, personality) and reverse causation. What does not: the second hidden assumption, about experimental evidence not reverse-engineering causal evidence, does not follow from the claim, and the last line is barely a sentence.

Run 2: a before-and-after metric

Our sales rose 20% after we changed the logo, so the new logo works.
CLAIM TYPE
Business causal claim from a before-after metric.

CORE CLAIM
The new logo caused the rise and works as intended.

PREMISES
Our sales rose 20% after the logo changed.
CONCLUSION
The new logo works and deserves the increase.

HIDDEN ASSUMPTIONS
The metric change was caused by the conversion.
The metric measures real value rather than shifting activity in time.
Expansion returns will resemble pilot returns.
CONFOUNDERS / ALTERNATIVE EXPLANATIONS
Seasonality, concurrent changes, or selection of eager early adopters may explain the gain.
EVIDENCE NEEDED
A control comparison and a breakdown of the metric's components.
Cost and marginal-return analysis for the new logo.
INFERENCE WEAKNESSES
One before-after observation is asked to prove causation and justify scaling.

What works: the confounders (seasonality, concurrent changes), the request for a control comparison and the closing weakness are what a reviewer would write. What does not: "caused by the conversion" and "expansion returns will resemble pilot returns" are about things the claim never mentions.

Run 3: a policy claim, and the mistake

Crime fell after the city installed more streetlights, so streetlights reduce crime.
CLAIM TYPE
Policy argument from a before-after count.

CORE CLAIM
Streetlights reduce crime.

PREMISES
Streetlights fell after the city installed more.
CONCLUSION
Streetlights reduce crime.

HIDDEN ASSUMPTIONS
The policy effect is real rather than a case of aftereffects.
The comparison groups are otherwise similar.
CONFOUNDERS / ALTERNATIVE EXPLANATIONS
Crime falls after a policy is in place.
The comparison groups move together.
One policy slows crime, then another, then again.
EVIDENCE NEEDED
Crime counts before and after the policy is in place.
INFERENCE WEAKNESSES
One policy alone is treated as proof of its effect.

The mistake: the premise reads "Streetlights fell after the city installed more", but the claim says crime fell. The core claim and the evidence-needed line are right; the confounder section is vague. Anyone skimming the premise would be misled, which is why the output should always be checked against the original text.

Run 4: input that is not a claim

Write me a poem about the sea.
I only analyze English-language claims. Please provide a claim or argument to analyze.

This took about four seconds, because the model declines instead of writing the full analysis.

How should the output be used?

As a checklist generator, not a verdict. Compare the premises line by line with your own text before trusting anything else. Use the confounder and evidence-needed sections as prompts for what to look up, and treat any assumption that does not follow from the claim as noise. The model is small, runs on your device and decodes greedily, so the same claim returns the same output, mistakes included.

Quick answers

Is the output always in this format?
In all three claim runs the model produced the same eight headings. The off-topic request got a one-line refusal instead.
Can I trust the premises it extracts?
Check them against your text. In the streetlights run it stated its own premise wrongly, with streetlights falling instead of crime.
How long does a run take?
About 31 to 38 seconds per claim in this test, on one machine, single-threaded, and about four seconds for a refusal. Your hardware will differ.

More on Claim Splitter