1. What kind of study is it
Design determines what a study can tell you, and no amount of size compensates for the wrong design.
Laboratory and cell studies describe mechanism. They cannot tell you what happens in a person. A compound that kills cancer cells in a dish tells you almost nothing about whether eating it does anything.
Animal studies are a step closer and remain a long way off. Most interventions that work in rodents do not translate to humans.
Observational studies record what happens to people without intervening. They are essential for generating hypotheses and are structurally unable to establish cause, because the groups being compared differ in ways nobody measured.
Randomised controlled trials allocate participants to an intervention or a comparison by chance, which balances out both known and unknown differences. This is the design capable of demonstrating cause.
Systematic reviews and meta-analyses pool the underlying studies. They inherit the quality of what they pool. A careful synthesis of poor trials is a careful summary of poor evidence.
2. In whom, and does that include you
A trial in people with a diagnosed deficiency does not describe people with adequate levels. A trial in hospitalised patients does not describe healthy adults. A trial in young men does not necessarily describe older women. This single question dismantles a large share of supplement marketing, which routinely sells to everyone on the strength of research conducted in a narrow group who were not like the buyer.
3. How many, and for how long
Small studies produce unstable results. Chance alone throws up impressive looking effects in small samples, which is why replication matters more than any single finding.
Duration has to match the claim. Skin ageing happens over decades. A study running for a few weeks cannot speak to it, whatever it measures. When a supplement claims a benefit that would take years to matter and the trial ran for eight weeks, the mismatch is the finding.
4. Compared with what
An intervention is only ever better or worse than something else. Compared with nothing at all, almost anything looks active, because people who join a trial change other behaviours, and because expecting improvement produces reported improvement. A meaningful comparison group is a placebo or an existing standard treatment, and participants and assessors should ideally not know who received what.
In skincare, the comparison problem is acute. A trial of an active ingredient in a moisturising base, compared against no treatment, has demonstrated the effect of moisturiser.
5. Measuring what
This is the most useful question on the list and the least asked.
A surrogate outcome is a marker measured because it is convenient and assumed to track something that matters: an instrument reading of skin elasticity, a biochemical marker, a threshold for visible redness. A clinical outcome is the thing you actually care about: your skin looks better, your condition improves, you get ill less often.
Surrogates fail regularly. Interventions have improved markers while making no difference to outcomes, and occasionally while making outcomes worse. When a study reports only surrogate outcomes, it has produced a hypothesis, not a result.
Then ask whether a change would be perceptible. A change can be statistically detectable and far too small for anyone to notice. Statistical significance is a statement about how likely a result is to be chance, not about how big or how important it is. Large studies detect tiny effects. That a difference is real does not mean it matters.
6. Who paid, and who wrote it
Funding does not invalidate research, and it does shift probabilities. Industry funded studies are more likely to report results favourable to the funder, through mechanisms that are mostly not fraud: choice of comparison, choice of dose, choice of outcome, and the decision about whether to publish a disappointing result at all.
Look for the funding statement and the competing interests declaration, which reputable journals require. Then ask the structural question: is this a field where independent groups have replicated the finding, or does almost all the evidence come from parties selling the ingredient. That question does more work than assessing any individual paper.
7. Has anyone replicated it
A single study is a suggestion. Science is a process of repetition, and findings that do not survive it are common across every field. The first study of anything is the least reliable one, and it is also the one that gets the press coverage, because it is new.
8. Does the headline match the paper
By the time a finding reaches a reader it has passed through a university press office, a news desk and a headline writer, each of whom has an incentive to sharpen it. Common distortions include turning an association into a cause, dropping the population the study was conducted in, converting a relative change into an impressive sounding number without giving the underlying rate, and omitting that the study was in mice.
Relative and absolute change is worth internalising. A large sounding relative reduction in a risk that was tiny to begin with is still a tiny change. Both numbers are needed and usually only the flattering one is given.
9. Checking something yourself
PubMed indexes the biomedical literature and is free to search. Abstracts are free even where full texts are not, and an abstract usually reveals design, population, size, duration and outcome measures, which is most of this list. The Cochrane Library publishes systematic reviews with explicit methods and plain language summaries. NICE publishes UK clinical guidance with the evidence review attached. These are all public.
The most efficient move when you encounter a striking claim is to search for the topic plus the words systematic review, and read the plain language summary of what the whole body of evidence shows. It takes a few minutes and it is more informative than any individual study you will be shown by somebody selling something.
For how we apply all of this when assigning a grade, see how we grade evidence.