Methodology
The hard part of this product is the rubric, so we publish it. Every grade, gap, hype position, and safety call on the site is produced by the rules below — and is reachable from the claim it produced.
Cite this methodology: 10.5281/zenodo.21364236 (concept DOI, resolves to the latest version) · current version 10.5281/zenodo.21364237 (v1.0.0, immutable) · source repository.
Anchored to established standards
We don't invent a grading system from scratch — we anchor to the frameworks a Cochrane methodologist, a journal editor, and a quality rater already recognize, and state exactly where we deviate and why.
| Layer | Anchored to | What we publish |
|---|---|---|
| Certainty of evidence | GRADE (used by Cochrane, WHO, NICE) | Our four certainty levels map to GRADE's high/moderate/low/very-low. |
| Risk of bias | RoB 2 (RCTs), ROBINS-I (non-randomised) | Which tool per study design; who applies it. |
| Evidence synthesis reporting | PRISMA (EQUATOR Network) | Search dates, databases, strings, inclusion/exclusion, study flow. |
| Effect interpretation | Minimal clinically important difference (MCID) where established | Our clinicalMeaning field's rules; explicit "no established threshold" where none exists. |
| Corrections & retractions | COPE (Committee on Publication Ethics) | Our corrections policy. |
| Disclosure | ICMJE disclosure format | Reviewer COI forms (see independence). |
Where we deviate, and why
- Grade and certainty are separate axes. GRADE bundles direction into recommendations; we split effect direction/magnitude from certainty. This is our best idea in the rubric — defended, not hidden.
clinicalMeaningis mandatory. GRADE permits "statistically significant" to stand alone; we don't.- Eight absence-of-evidence states, where most frameworks have one undifferentiated "insufficient evidence."
- Interaction severity carries its own
evidenceCertainty, separate from the interaction's severity tier.
We invite critique of this rubric, named and published unedited — see methodology critique.
Evidence grades
Grades describe the strength of the evidence, on an A–D scale:
Certainty is shown separately from the grade — a claim can be a confident read of weak evidence, or a tentative read of a larger literature. We never collapse the two.
Re-review cadence (the SLA table)
Freshness is not just reacting to news; it's a stated, kept cadence. Every evidence surface carries a next-review promise, visible on the page. The table below is what we hold ourselves to.
| Surface | Scheduled re-review | Also triggered by |
|---|---|---|
| Grade A/B pairs | every 12 months | new RCT or meta-analysis on the pair |
| Grade C/D pairs | every 6 months | new RCT or meta-analysis on the pair |
| Interaction pages (severity ≥ moderate) | every 6 months | safety signal, FDA/EMA communication |
| Interaction pages (none/low) | every 12 months | same |
| Hype pages | attention monthly, stage quarterly | virality spikes in the L4 pipeline |
| Methodology | annual, versioned | — |
Internal targets: ≥ 95% of pages within SLA at any time · median trigger→publish latency ≤ 14 days · ≤ 72 h for safety upgrades (see the safety fast-path in the freshness spec).
Corrections policy
We never silently edit. A correction is a first-class event — new date, visible label, and a permanent/changes/ URL. The first real correction we publish will do more for trust than any badge on the pricing page.
Claim states (including the honest absences)
- well-supported / mixed / weak — graded effect claims.
- studied, no effect (null-result) — tested and found not to work. This is a finding, and it is gradeless.
- not well studied (untested) — there isn't adequate evidence yet. Also gradeless. It is distinct from null-result, and we never blur them.
We do not fabricate to fill a gap. When we don't know, the answer says so explicitly rather than inventing a claim.
The belief-vs-evidence gap
For every popular belief we compute its distance from the evidence as a direction and an ordinal magnitude — never a fake decimal.
- belief-ahead — the claim runs ahead of the evidence (hype).
- belief-behind — the evidence is stronger than people assume (under-appreciated).
- belief-contradicts — the evidence tested it and found it null or opposite (a myth).
- belief-matches — belief and evidence are roughly aligned (settled).
Magnitude (negligible · modest · large · extreme) combines the grade-distance, a penalty when the evidence state is unflattering to the assertion, and a penalty when the claim is marketing or mechanistic speculation. Direction is high-confidence; magnitude is coarse and honest about being coarse. A belief never inherits the evidence's grade.
Hype-cycle position
Position is derived from two trajectories — attention momentum and the evidence-grade trajectory — not assigned by feel. Possible positions: emerging → surging → evidence-catching-up → settled → declining → debunked. A position only moves when the condition holds across windows or a shock forces re-evaluation, and debunked requires a confirming evidence decline, not merely falling attention.
Safety severity
Safety always comes first in the answer. Interactions are tiered:
People say · Law permits · Evidence shows
We keep three things visibly separate and never let one masquerade as another:
- People say — a popular belief, scored for its distance from the evidence (the gap above).
- The law permits — what a label may legally claim. Under regimes like US DSHEA, a supplement label can carry a structure/function claim (e.g. “supports immune health”) that does not have to be proven. A permitted claim is never a truth claim.
- The evidence shows — the graded claim with provenance.
So a product can legally say “supports a healthy immune system” even where the evidence for, say, preventing colds is a null-result. We surface the legal claim and the evidence, side by side, rather than repeating the label as fact.
Sources & authority
Every graded claim carries its sources with the signals that establish authority — authors, venue, year, peer-review status, institution — published in machine-readable form (schema.org) so answer engines can judge and attribute them.
How a claim is verified (and how you can check)
New evidence enters through a guardrailed pipeline, not a content form. Automation may propose a claim, but it publishes only after passing every gate below — and each published claim carries a public, tamper-evident receipt you can open.
- The gauntlet (automated gates). A candidate publishes only if it references real catalog entities (it can never invent a supplement or outcome), is backed by an authoritative source whose PMID/DOI actually resolves, passes the claim axioms, states certainty and a grade no higher than its study design supports (the deterministic grader's ceiling), avoids over-claiming or disease-treatment language, and isn't a duplicate. Rejections come back with reasons — nothing is silently dropped.
- Named, role-scoped human sign-off. A qualified reviewer signs off under their real name, attesting only to the part they're competent to judge (e.g. methodology, clinical-safety). No single person decides: the publish policy is a typed-role quorum — the strict mode requires methodology and clinical-safety sign-off from two distinct reviewers.
- Conflict-of-interest recusal. Reviewers disclose conflicts; a disqualifying conflict recuses that sign-off — it cannot count toward publication.
- A tamper-evident receipt. Every publish and every later regrade is written to a hash-chained transparency log. The public receipt for any claim shows who vouched for it, in which role, with what disclosed conflict, plus a chain-intact check — open
/receipt?claim=<id>from the trust badge on any claim. - Reader ratings are advisory only. A “Readers say” signal bridges across viewpoints (so a one-sided pile-on can't promote a note), but it is strictly non-gating — reader ratings never affect whether a claim is published.
Neutrality, review & governance
- We sell no supplements and take no product money or editorial affiliate revenue. Perceived neutrality is the product.
- Grades, gaps, and safety calls are human-confirmed; automation proposes, qualified reviewers sign off (see the verification flow above). Nothing auto-publishes.
- The rubric is versioned; re-scoring under a new version is itself a recorded change (see what changed) and is added to the claim's receipt.
- Not medical advice; not a medical device. This is decision-support and education — never diagnosis, prescription, or a substitute for a clinician.