How we grade evidence
Last updated: September 2026
In short
Every paper in the library is appraised on four things: how well it was done for its design, whether its funders had a stake in the result, how transparently it was reported, and how strongly its evidence supports its conclusions. A study is not marked down for being observational; it is marked down for funding conflicts. When we answer, the certainty we ask for scales with the stakes of the advice.
1. Two passes over every paper
First pass: fast screening and a baseline score
When a paper enters the library, a fast model reads the full text. It confirms the paper is peer-reviewed research on nutrition or health, extracts the bibliographic details, classifies the design (study type, sample size, duration, population), and gives a first score on the four dimensions below from the paper's own text, including its funding and competing-interest statements. This is what lets us cover tens of thousands of papers.
Second pass: a full investigation
A slower, more capable agentic review then re-reads the whole paper and goes outside it. With a search tool it looks up the principal authors' track record and past funding ties, the institutions and labs involved, and every funder named, asking whether any of them had a commercial stake in the outcome and whether the findings served that stake. It also writes a structured methodology critique: design, randomisation, blinding, control, confounders, statistics, strengths and weaknesses. Where this review exists it overrides the first pass, and its full reasoning is published on the paper's trust report. It runs continuously, oldest papers first, and revisits papers as what we know about their authors and funders grows.
2. The four dimensions
Each dimension is scored from 0 to 100%. The overall trust score is the mean of the dimensions that could be assessed; a dimension the reviewer could not judge is left blank, never guessed.
- Methodology: how well the study was done for its design. A randomised trial is judged on randomisation, concealment, blinding where feasible, its control, power, pre-registration and attrition. A prospective cohort is judged on how diet was measured, how thoroughly confounders were adjusted for, whether outcomes were objectively ascertained, and whether follow-up was long enough. A systematic review is judged on its search, its risk-of-bias assessment and how it handled heterogeneity. A large, well-run cohort can score as high as a well-run trial; what a design cannot establish is recorded separately, under evidence strength.
- Independence: funding and conflicts of interest. 100% means independent. The score falls when a party with a commercial stake funded, employed or supplied the authors and the findings served that party. Direction matters: an industry-funded study that found its funder's product ineffective is not treated as a conflict, while undisclosed ties are penalised whatever the finding.
- Reporting transparency. Pre-registration or a published protocol, data availability, complete methods, all pre-specified outcomes reported, disclosures, and an honest discussion of limitations.
- Evidence strength. What the design can establish (association or cause), effect size and precision, a dose-response gradient, hard versus surrogate endpoints, consistency with prior evidence, and whether the paper's causal language matches its design.
3. Why we do not use a drug-trial hierarchy
The most widely used grading system in medicine, GRADE, was designed around clinical trials of treatments and by construction starts observational evidence at “low certainty”. Nutrition researchers have long argued that this template fits most diet questions poorly and have proposed alternatives that grade a body of nutrition evidence on its own terms: NutriGrade, HEALM, and the World Cancer Research Fund's criteria (linked below). What they share is what we adopted: each study design is appraised against its own best practice, funding bias counts against a body of evidence, and consistent, well-conducted observational evidence backed by a plausible mechanism can carry a strong grade.
Dr. Michael Greger's post When Should We Rely on Observational Studies? is a readable case for this view and was the prompt for the September 2026 revision described on this page.
4. How answers use the scores
- Ranking, not rejection. Inside the app a low trust score sinks a paper in the ranking rather than hiding it, and a funding conflict shows as a badge. The public API and assistant connectors apply a trust floor and a conflict filter, because they promise a filtered answer.
- The answering model sees the design. Every source it reads comes with the paper's study type, sample size, duration, trust score and a one-line methodology note from the reviewer. It is instructed to weigh sources by how well their design answers the question at hand, to discount lower-trust sources even when they are trials, and to say what kind of evidence the answer rests on.
- Source quality, on Ask and Chat. The label on each answer describes the retrieved sources: high means at least three primary studies or syntheses of any design with high average trust; moderate, at least one with solid trust; limited, only narrative reviews or lower-trust sources. It is not a verdict on the whole body of evidence.
- Evidence grade, on Panorama. Panorama reads every relevant paper one at a time, so it can grade the body of evidence behind each question with a fixed, published rule set:
| Strong | At least three trials, cohorts or syntheses across more than one design (or including a synthesis); at least 75% of the trust-weighted findings point the same way; average trust at least 60%; no more than a third of the human studies with a funding conflict. A cohort-only body also needs a reported effect size or mechanistic support. |
| Moderate | At least two trials, cohorts or syntheses; at least 60% of the weighted findings agree; average trust at least 55%; no more than half with a funding conflict. |
| Limited | Everything else: suggestive, sparse, or lower quality. |
| Conflicting | Three or more human studies and no direction carries more than half of the weighted findings. |
| Insufficient | No human primary study or synthesis: only animal or lab work, case reports, or narrative reviews. |
Consistency is weighted by each finding's trust and confidence, so a few well-done studies outweigh many weak ones. Consistent null results grade the same way: strong evidence of no meaningful effect is still strong evidence. The numbers behind every grade are shown on the report.
5. Certainty in proportion to stakes
The level of proof an answer demands scales with what happens if the answer is wrong. We owe the framing to the post linked above; how we apply it is our own:
- A low-risk, reversible change with no plausible downside, such as eating more whole fruit, can be recommended on consistent observational evidence, and the answer says that is what it rests on.
- A supplement at a pharmacological dose, such as high-dose vitamin E, needs trial-level evidence; where trials exist, the answer rests on them rather than on early observational signals.
- Advice for pregnancy, children or a medical condition, and anything that could displace medical care, such as an elimination diet for a child, gets the highest bar: the answer stays conservative and points to a clinician.
6. Limits and changes
- Scores are assigned by models and can be wrong. Every trust report shows the full reasoning and has a report button; we read those reports and feed corrections into the next review.
- Every score records the rubric version that produced it. When the rubric changes, as it did in September 2026 to remove the drug-trial ceilings on observational designs, old scores are not silently rewritten; papers scored under an older rubric are revisited by the continuous review.
- Nothing here is medical advice. The grades describe the weight of published evidence, not what is right for you.
Further reading
Questions or corrections: use the report button on any trust report, or write to support@allnutrition.info.