Skip to content
James Kirkman , home
At home Active

Trickle.news

A daily digest that measures the gap between what matters and what gets covered, and shows every step of the measurement.

Role
Creator (Claude agents implement)
When
Oct 2026 – present
Stack
Python · Claude · Federal Register API · Media Cloud · Crossref · PubMed · OpenAlex · Git

A public daily digest of new federal and state rules and research, written from primary sources only. A pipeline of code rules, a cheap judge model, and Claude decides what's notable, and an independent coverage count shows which stories the news missed. Its own failure rate is published, not hidden.

Problem

Every week, US agencies publish rules that change what people pay, what they’re allowed to do, and what treatment they can get. Some become front-page stories; many are never covered at all. Which ones fall through, and is there a pattern?

Answering that honestly needs two things that are easy to fake: a judgment of what matters made before anyone knows what the news did with it, and a coverage count you can trust, including its misses.

Approach

  • One pipeline, a second product. Trickle.news grew out of a private research pipeline that reads public signals. Labelling regulatory signals kept turning up rules that were real news but not business opportunities, so the same sources, ledger, and decision cascade now feed a public digest with its own question: would the people affected want to know this?
  • Models only where they earn it. Code rules decide about half the items, the cheap judge answers narrow questions about the rest, and Claude writes the summary. Code, not a model, supplies every date and every proposed-or-final status.
  • Check the checker. On October 9 a fact-checker held back 19 of 61 leads, nearly all false positives: “not yet peer-reviewed” read as “peer-reviewed,” a date phrase read as a name. A comparison of two writer models looked like a model problem until the checker was fixed; then the smaller, cheaper model held back two leads and wrote just as accurately.
  • Measure the gap, then measure the measurement. Each published rule gets a news-coverage count from Media Cloud’s open archive, counts only. Zero counts are tested, not trusted. The first independent recall check caught 7 of 12 covered rules, and the causes were specific: the window opened on the publication date though agencies announce days earlier, and official agency names aren’t how reporters write. The revised method opens a week early and searches names the way reporters do (“HHS”). The rerun on 58 rules lifted recall to 60% and precision from 58% to 71%. What’s left was mostly formal titles no reporter would type, so a third method searches the subject words reporters use, requires the agency, and lets a companion rule share its main rule’s coverage. Its recall check is next. Research coverage is now counted too, labelled experimental until a recall check measures it.
  • Show every step, including the misses. The public method page lists every step’s accuracy against its target, re-scored weekly on hand-labelled sets, with the metrics still below target marked as such; a weekly job opens an issue when one slips. A run dashboard shows every job’s counts from source to lead.
  • Protect people. Officials and authors can be named in their public role; private individuals, minors, and contact details never are, and a test caught a name leaking through one register’s scrub.

Architecture

From a primary source to a published lead and its coverage count

  1. Step 1: Sources

    • Source: Primary sources rules and research, never news outlets
  2. Step 2: Filter

    • Check or gate: Code rules decide about half with no model
    • Model: Cheap judge narrow yes/no questions
  3. Step 3: Write

    • Model: Claude writer
    • Check or gate: Fact-checker against dates and status from code
  4. Step 4: Publish

    • Data store: Daily edition
  5. Step 5: Measure

    • Tool or service: Coverage count Media Cloud, counts only

Recall checks against an independent search measure the coverage count itself, and the method changes when they fail.

  • Source
  • Check or gate
  • Model
  • Data store
  • Tool or service
New rules and research papers arrive from primary sources only. Plain code rules decide about half of them without a model. A cheap judge model answers narrow yes-or-no questions about the rest, and Claude writes a summary that a fact-checker verifies against the source. Leads that pass are published in a daily edition. A count of news coverage then measures which stories were reported, and periodic recall checks against an independent search measure the count itself.

Key decisions

  • Primary sources, politely

    Rules, registers, and research only, never news outlets. Robots files bind even documented APIs, and every rejected source is recorded with its reason.

  • Code before models

    Plain rules decided half of a month of federal rules with no model, and dropped nothing the labellers called news.

  • A cheap judge, asked narrow questions

    About 500× cheaper per decision than the large model, and accurate on questions about the text in front of it, though no better than a constant answer on questions that need world knowledge.

  • Set the threshold from the scores

    The judge's AUC was 0.83 but its median score 0.35, so the cut-off went to 0.3, not the textbook 0.5. At sign-off, precision 0.77 and recall 0.82 on 177 labelled rules.

  • A visible rule, not an agent acting in my name

    When an agent tried to record my delegation as a decision in my name, Claude Code's safety check stopped it. I changed the publishing rule myself, as a setting anyone can read.

Results

leads published in the first two days
246
precision / recall of the news gate, re-scored weekly
0.76 / 0.81
recall of the coverage count, after fixing what the first check found (precision 58% → 71%)
58% → 60%

Lead count as of the night of October 9, 2026; gate scores from the weekly re-score of October 10. Recall checks are for federal rules; state coverage was too thin to measure recall yet.