Live23 Sept 2026/8 verified items today/Claim check: Yes, a 744B model runs on a laptop with no GPU/Daily digest
Preview: items marked "(demo)" are sample data, not yet editor-verified. How we verify

Methodology

Trust is earned through a track record, so the system makes errors rare, visible and quickly fixed. These rules are binding.

Accuracy

How the pipeline works

Watchers (plain code, no LLM) poll lab blogs, GitHub releases, cloud changelogs, pricing pages, Hugging Face and arXiv every 5–15 minutes and only queue real changes. Five agents then help: a small classifier drops ~90% of raw items, an extractor pulls structured facts with the source line per fact, a verifier checks every fact against the source, a writer drafts pages for Tier 1 releases, and a proof runner executes the eval suites. Agents never publish. A human editor approves every fact, and only measured proof runs change a recommendation.

Evidence levels

Independence

Transparency

Editors

Trust metrics (public)

MetricTargetCurrent
Errors per 100 published itemsNear zero before public launch13.3
Time to correctionUnder 24 hours8h average
Release to verified coverageSame day for frontier releasesTracked from Phase 1
External citations per weekGrowing month over monthTracked from Phase 0

Current values are computed from the demo dataset in this preview build.