Skip to content

How a 270M-parameter claim analyzer runs entirely in your browser

Project documentation · 3 min read · updated

Claim Splitter runs GemmaClaim-270M, a fine-tune of Google's Gemma 3 270M, inside your browser tab. A 253 MB, 4-bit GGUF file is downloaded once from Hugging Face and cached, then run by wllama, a WebAssembly build of llama.cpp, with greedy decoding, a 2,048-token context and at most 512 new tokens. The page's own code never uploads your claim: the only network request it starts is the model download.

Base model
google/gemma-3-270m
Training
Full-parameter fine-tuning, then LoRA, on a synthetic claim-analysis dataset with refusal examples (model card)
Model file
gemmaclaim-270m-q4_k_m.gguf, 253 MB
Runtime
wllama (llama.cpp compiled to WebAssembly); runtime file 8,457,512 bytes
Limits
2,048-token context, 512 new tokens at most, input up to 1,200 characters
Decoding
Greedy (temperature 0)
Language
English only; a client-side check refuses other languages before the model runs

What is in the 253 MB download?

The quantised weights of the fine-tuned model: a 270-million-parameter model stored as a 4-bit (Q4_K_M) GGUF file hosted on Hugging Face. The page shows a progress bar while it downloads, asks the browser to cache it, and remembers in localStorage that the model was cached, so a returning visitor's model loads without a click. In a test on 3 October 2026 (desktop Chrome, headless, empty cache) the model downloaded and loaded in 40.1 seconds; your time depends on your connection.

How does a browser run a language model?

wllama compiles llama.cpp to WebAssembly. The page loads its 8,457,512-byte wllama.wasm from this site and creates the model with a 2,048-token context. A claim is sent as a streaming completion with temperature: 0 and at most 512 new tokens, and the reply is rendered as it arrives. In the test the page was not cross-origin isolated, so multi-threaded WebAssembly was unavailable and generation ran on a single thread; the page's status line says "WebGPU if available" when the browser exposes navigator.gpu. Timings from your own hardware will differ.

What prompt does the model get?

The fine-tuned model has absorbed the task, so the page sends only a short header around the claim, the compact format from the model card:

Analyze this claim.

INPUT:
<your claim>

ANALYSIS:

Decoding is greedy and stops at the end-of-sequence token, so the same claim should give the same analysis each time.

How does it refuse non-English or off-topic input?

In two layers. First, a small script in the page checks the text before the model sees it: it refuses when more than a quarter of the letters are non-Latin, or, for texts of six words or more, when there are at least four accented letters and at most one common English word, or when a list of common German, French, Spanish, Italian, Portuguese, Dutch or Turkish words matches at least twice and more than one time more often than common English words do. The reply is "I only analyze claims written in English." Second, the model was trained with refusal examples: for "Write me a poem about the sea." it answered "I only analyze English-language claims. Please provide a claim or argument to analyze." in about four seconds.

How long does an analysis take?

InputTime reported by the page
Coffee and lifespan claim32.2 s
Logo and sales claim37.8 s
Streetlights and crime claim31.4 s
Off-topic request (refused)about 4 s

One run each, single-threaded, on one machine. See the example output page for the text these produced.

What does the author report about quality?

The tool page reports a model trained on about 1,400 synthetic examples, including refusals for non-English and off-topic input, and a section-header F1 against reference answers that rose from 0.42 to 0.91 on 68 held-out prompts. That score measures whether the output has the right section headings, not whether the content is correct. The page also states plainly that the model is small and wrong sometimes, and the example output page shows a real mistake.

Quick answers

Does Claim Splitter send my text to a server?
The tool's own code makes no request with your text. Inference runs in the page with wllama, and the only network traffic it starts is the one-time model download from Hugging Face, plus the runtime file from this site. The site's analytics tags load on every page, as they do on all pages here.
Why is the first load slow?
The model is a 253 MB download. It is cached by the browser, and the page remembers that, so later visits load it without downloading again.
Can it analyze text that is not English?
No. A check in the page and the model's own refusal training both decline non-English input.
Which base model is it?
Google's Gemma 3 270M, fine-tuned for claim analysis. The model card lists the licence as "gemma".

More on Claim Splitter