Methodology
Trust is earned through a track record, so the system makes errors rare, visible and quickly fixed. These rules are binding.
Accuracy
- Primary sources only: official release notes, docs, pricing pages, model cards. Aggregators are leads, never sources.
- Two-check rule: every price, region, limit and date is verified by a human against the source.
- Status labels on everything: Confirmed, Announced (not yet available), Rumor.
- AI drafts must cite the exact source line per claim, or the claim is dropped.
- Practitioner channels (Reddit, Hacker News) are early signals only, never sources of fact.
How the pipeline works
Watchers (plain code, no LLM) poll lab blogs, GitHub releases, cloud changelogs, pricing pages, Hugging Face and arXiv every 5–15 minutes and only queue real changes. Five agents then help: a small classifier drops ~90% of raw items, an extractor pulls structured facts with the source line per fact, a verifier checks every fact against the source, a writer drafts pages for Tier 1 releases, and a proof runner executes the eval suites. Agents never publish. A human editor approves every fact, and only measured proof runs change a recommendation.
Evidence levels
- Measured by us
- Vendor-documented
- Community-reproduced
- Community-reported
Independence
- No pay-for-ranking, reviews or placement. All revenue sources are disclosed.
- No affiliate links on release or proof pages.
- Vendor right of reply, shown on the page.
- Sponsored proofs: a vendor may pay for a run to happen, never for the result, and it is clearly labeled.
- We check each product's license before publishing benchmarks; some restrict publishing results.
Transparency
- Public corrections log with dates.
- A "Disagree" button on every page. Accepted fixes are credited (+15 points).
- Every proof states what was NOT tested.
Editors
- Editor seat 1 (to be named) — ML engineer, cloud AI platforms
- Editor seat 2 (to be named) — Research engineer, evals and RAG
Trust metrics (public)
| Metric | Target | Current |
|---|---|---|
| Errors per 100 published items | Near zero before public launch | 13.3 |
| Time to correction | Under 24 hours | 8h average |
| Release to verified coverage | Same day for frontier releases | Tracked from Phase 1 |
| External citations per week | Growing month over month | Tracked from Phase 0 |
Current values are computed from the demo dataset in this preview build.