Claimstone

Open source · Apache-2.0

A reading machine that refuses to overstate what it read.

Give it topics and a frozen list of questions. It finds the literature, obtains what it legally can, and reads every source. You get, per question, the evidence bound to verbatim quotes, with the coverage that evidence rests on.

  • every claim has a quote
  • no pooling
  • a person signs
chunk c-0412accepted
… in the pooled sample, the twelve-month return after a negative headline was 0.8 percentage points lower (p < 0.05) than after a neutral one, and the gap was not visible at shorter horizons …
Negative headlines are followed by 0.8 pp lower twelve-month returns (p < 0.05).
  • quote is an exact substring of the chunk
  • every number and inequality appears in the quote
chunk c-0412rejected → ledger
Negative headlines are followed by 1.2 pp lower twelve-month returns.
  • “1.2” does not appear in the quote

Illustrative example, not a result from a real source.

What it is for

Ask the literature. Check the answer.

When you read the literature there are two mistakes to avoid: claiming more than the sources say, and concluding that an effect doesn’t exist just because no evidence turned up. Claimstone is for when you can’t afford either.

  • There are hundreds of papers on your topic and no time to read them.

    It finds candidates through two independent routes, keyword search and citations, gets the legal copy of every paper it can, and reads them one by one.

  • AI summaries sound sure of themselves, but you can’t tell what the paper actually says.

    Every claim comes with a quote copied word for word from the paper. The code checks that the quote is really there and that every number in the claim appears in it. What fails is thrown out, and the list of rejections stays visible.

  • “I found nothing” could mean there is nothing, or that the papers couldn’t be obtained.

    It measures how much of what it found it actually obtained, against a threshold declared in advance. Below it, it stops and draws no conclusions. And it keeps three easily confused cases apart: the sources don’t answer, the sources say the opposite, the sources disagree with each other.

  • You need an answer you can defend, not a black box.

    For each question it prepares an evidence profile: every result, how many sources point each way, what was rejected and how much was read. A person reads it and signs. Every step writes plain text files you can open and search.

And it isn’t tied to one field. Topics, questions and kinds of source are input files; nothing in the engine is specific to finance or biology.

How it works

From your questions to an answer you can check, in six steps.

You provide the topics and the questions. Claimstone does the rest, one step at a time, and leaves a file at every step that you can open and check.

You provide the topics, a fixed list of numbered questions, and the kinds of source you accept. They are three plain files.

  1. 1discover

    Finds

    Searches for papers on your topics in two independent ways: by keywords and by following citations. You get the list of candidates, each marked with what kind of source it is, such as a peer-reviewed paper or a blog post.

    writes · candidates.jsonl
  2. 2acquire

    Obtains

    Gets the best legal copy of each paper, preferring open access, and never goes through pirate libraries. It records every attempt and why it failed, so you know how much it could really read.

    Papers behind a paywall can’t be fetched. You can add a copy you got yourself, from a library or by purchase. It goes through the same identity and full-text checks, and is reported on its own line, so it never quietly inflates how much was read.

    writes · acquisitions.jsonl · raw/
  3. 3normalize

    Prepares

    Turns PDFs and web pages into clean text and splits it into passages, so every quote can be traced back to an exact place.

    writes · documents.jsonl · chunks.jsonl
  4. 4extract

    Extracts

    A model proposes the claims it finds in each paper. The code then checks every claim against its quote. The ones that fail are rejected and listed.

    writes · claims.jsonl · rejections.jsonl
  5. 5review

    Rechecks

    A second, different model rereads each claim in its whole passage, to check it still holds in context.

    writes · reviews.jsonl
  6. 6synthesize

    Summarizes

    For each question it assembles an evidence profile from what was accepted: the results, how many sources point each way, what was rejected and how much was read. No model, no network, no statistics at this step: it only organizes and counts.

    writes · profiles.jsonl

You get an evidence profile for each question. A person reads it, decides, and signs.

The answers

Not just yes or no: five possible answers.

A question put to the studies doesn’t always have a yes or a no. Each verdict says what the evidence lets you claim, and no more. A person records it after reading the evidence profile.

The evidence points one way

SUPPORTED

The evidence says yes.

The profile is convincing, and whoever signs writes down why.

CONTRADICTED

The evidence says the opposite.

The profile is convincing in the other direction.

The evidence doesn’t decide, and that can happen in three ways

CONTESTED_IN_LITERATURE

The studies disagree.

The studies speak and contradict each other in a way that can’t be reconciled.

UNANSWERED_IN_LITERATURE

The studies don’t settle it.

They were read, and they aren’t enough to decide.

NEVER_ASKED

Nobody has studied it.

A person checked that it isn’t just a gap in the search.

These three look the same from outside, and they aren’t. Treating “the studies disagree” as “nothing found” is the mistake Claimstone exists to avoid.

The signature is tied to the evidence the person saw: if the evidence changes later, the verdict is marked out of date. Questions are numbered and frozen, and changing the list is a dated version change.

No verdict has been signed yet. These are the states the project recognises.

Where it stands

Early, and open about it.

Claimstone works from start to finish, but it is young. Here is what has been done and what hasn’t.

Works today

All six steps run, from the search to the evidence profile. One full round was run on open-access articles from PubMed Central about screen time: 37 of the 40 papers found were obtained, above the 80% threshold set in advance. It produced 1,721 accepted annotations and 271 rejected ones, all listed.

Not yet

No verdict has been signed. The first profile is ready for a person to read; the other seven are provisional. Two other collections of papers stayed below their threshold, so Claimstone correctly produced nothing for them. The portal for following the work is read-only for now.

Where help is wanted

Reports, measurements and new fields all help. The ways to take part are just below.

See the ways to take part →

Join in

Three ways to take part, from ten minutes to a whole project.

You don’t need to know the engine to help. Pick the size that fits your time.

Ten minutes

Point out what doesn’t add up

Read how the project describes itself and tell us where it says more than it can show, or where a check lets something through. An issue with the page and the sentence is enough.

Open an issue →
An afternoon

Put a decision to the test

Every design decision is recorded with the measurement that settled it. Pick one, repeat the measurement or bring a better one, and tell us what you find.

Read the decisions →
A project

Use it on your own field

Write the three input files for a topic you know, from public literature, and run it. Whatever the checks reject on your field is the most useful report we can get.

Follow the guide →

One rule is not negotiable: no claim enters without a verified quote. The rest is open to discussion. Read the rule