What the Claim Splitter returns: three real runs, one with a visible mistake
For every English claim the model writes the same eight sections: claim type, core claim, premises, conclusion, hidden assumptions, confounders or alternative explanations, evidence needed, and inference weaknesses. The three runs below are verbatim, captured on 3 October 2026. The first two read sensibly, with a few garbled lines; the third misstates its own premise, which is the kind of mistake to expect from a 270M-parameter model.
- Captured
- 3 October 2026, desktop Chrome (headless), model loaded from an empty cache in 40.1 s
- Settings
- Greedy decoding (temperature 0), 2,048-token context, 512 new tokens at most
- Time per run
- 32.2 s, 37.8 s and 31.4 s as reported by the page; one run each, one machine
- Edits
- None to the wording; some blank lines between headings were removed
What sections does the output have?
| Heading | What it holds in these runs |
|---|---|
| CLAIM TYPE | A one-line label, such as "Everyday causal claim (correlation presented as causation)" |
| CORE CLAIM | The claim restated in one sentence |
| PREMISES, CONCLUSION | What the claim offers as support, and what it concludes |
| HIDDEN ASSUMPTIONS | What must also be true for the conclusion to follow |
| CONFOUNDERS / ALTERNATIVE EXPLANATIONS | Other things that could produce the same pattern |
| EVIDENCE NEEDED | What would settle it |
| INFERENCE WEAKNESSES | Where the reasoning is thin |
Run 1: a correlation presented as causation
People who drink coffee tend to live longer, which shows coffee extends lifespan.
CLAIM TYPE
Everyday causal claim (correlation presented as causation).
CORE CLAIM
Drinking coffee causes people to live longer.
PREMISES
People who drink coffee live longer.
CONCLUSION
Coffee causes people to live longer.
HIDDEN ASSUMPTIONS
Longevity effects are caused by drinking coffee rather than by other causes.
The conclusion is informed by experimental evidence and does not reverse-engineer causal evidence.
CONFOUNDERS / ALTERNATIVE EXPLANATIONS
Wealth, education, or personality may drive both the drinking habit and longevity.
Reverse causation is possible if the outcome influences the habit rather than the habit itself.
EVIDENCE NEEDED
Experimental or mechanistic support for the causal claim.
Benefit-risk analysis and unintended consequences.
INFERENCE WEAKNESSES
'Drinking coffee' is everyday causal claim without justification.
What works: it names the claim as correlation presented as causation and offers two sensible alternatives, shared causes (wealth, education, personality) and reverse causation. What does not: the second hidden assumption, about experimental evidence not reverse-engineering causal evidence, does not follow from the claim, and the last line is barely a sentence.
Run 2: a before-and-after metric
Our sales rose 20% after we changed the logo, so the new logo works.
CLAIM TYPE
Business causal claim from a before-after metric.
CORE CLAIM
The new logo caused the rise and works as intended.
PREMISES
Our sales rose 20% after the logo changed.
CONCLUSION
The new logo works and deserves the increase.
HIDDEN ASSUMPTIONS
The metric change was caused by the conversion.
The metric measures real value rather than shifting activity in time.
Expansion returns will resemble pilot returns.
CONFOUNDERS / ALTERNATIVE EXPLANATIONS
Seasonality, concurrent changes, or selection of eager early adopters may explain the gain.
EVIDENCE NEEDED
A control comparison and a breakdown of the metric's components.
Cost and marginal-return analysis for the new logo.
INFERENCE WEAKNESSES
One before-after observation is asked to prove causation and justify scaling.
What works: the confounders (seasonality, concurrent changes), the request for a control comparison and the closing weakness are what a reviewer would write. What does not: "caused by the conversion" and "expansion returns will resemble pilot returns" are about things the claim never mentions.
Run 3: a policy claim, and the mistake
Crime fell after the city installed more streetlights, so streetlights reduce crime.
CLAIM TYPE
Policy argument from a before-after count.
CORE CLAIM
Streetlights reduce crime.
PREMISES
Streetlights fell after the city installed more.
CONCLUSION
Streetlights reduce crime.
HIDDEN ASSUMPTIONS
The policy effect is real rather than a case of aftereffects.
The comparison groups are otherwise similar.
CONFOUNDERS / ALTERNATIVE EXPLANATIONS
Crime falls after a policy is in place.
The comparison groups move together.
One policy slows crime, then another, then again.
EVIDENCE NEEDED
Crime counts before and after the policy is in place.
INFERENCE WEAKNESSES
One policy alone is treated as proof of its effect.
The mistake: the premise reads "Streetlights fell after the city installed more", but the claim says crime fell. The core claim and the evidence-needed line are right; the confounder section is vague. Anyone skimming the premise would be misled, which is why the output should always be checked against the original text.
Run 4: input that is not a claim
Write me a poem about the sea.
I only analyze English-language claims. Please provide a claim or argument to analyze.
This took about four seconds, because the model declines instead of writing the full analysis.
How should the output be used?
As a checklist generator, not a verdict. Compare the premises line by line with your own text before trusting anything else. Use the confounder and evidence-needed sections as prompts for what to look up, and treat any assumption that does not follow from the claim as noise. The model is small, runs on your device and decodes greedily, so the same claim returns the same output, mistakes included.
Quick answers
- Is the output always in this format?
- In all three claim runs the model produced the same eight headings. The off-topic request got a one-line refusal instead.
- Can I trust the premises it extracts?
- Check them against your text. In the streetlights run it stated its own premise wrongly, with streetlights falling instead of crime.
- How long does a run take?
- About 31 to 38 seconds per claim in this test, on one machine, single-threaded, and about four seconds for a refusal. Your hardware will differ.