Human of One Labs

How H[1] evaluates scientific evidence

H[1] applies gold-standard evidence-evaluation frameworks, including GRADE—the method widely used in systematic reviews and clinical guidelines—alongside its own transparent and reproducibility-tested methodology. Together, these methods evaluate how much confidence a body of evidence can support for a specific outcome.

The system records every judgment and evidence claim, shows the evidence behind it, and returns the result in a form that both people and AI systems can inspect. Any rules or provisional thresholds developed by H[1] are clearly distinguished from GRADE itself.

To reach a certainty level, H[1] defines the research question and outcome, identifies the relevant studies, and evaluates the risk of bias within each one. It then examines the evidence as a whole for inconsistency, indirectness, imprecision, and possible missing evidence.

The AI records structured judgments and the reasons behind them. Explicit rules then derive the final certainty level: High, Moderate, Low, or Very Low.

You can see which studies were included, which were excluded, what limitations were identified, and why the certainty level changed. H[1] also repeats assessments to test whether the same evidence produces consistent judgments.

The purpose is not to eliminate scientific judgment. It is to make that judgment structured, traceable, testable, and open to correction.

Trust the trail, not merely the verdict.

How certainty is determined

Evidence begins at a starting level based on study design. The certainty level may then move down when the evidence has important limitations or, in appropriate cases, move up when observational evidence has particular strengths.

start from design   randomised     → High
                    observational  → Low

downgrade  −1 serious, −2 very serious, for each of
           risk of bias
           inconsistency
           indirectness
           imprecision
           publication bias

upgrade    +1 each, observational bodies only,
           and only if nothing was downgraded
           large effect
           dose-response gradient
           residual confounding

report     High · Moderate · Low · Very Low

A numerical score can appear objective while hiding the judgments that produced it. We instead use four certainty levels—High, Moderate, Low, and Very Low—and require every change in certainty to be supported by a stated reason.

H[1] preserves those reasons. The final level is derived from the recorded judgments, so a reader can inspect how the conclusion was reached rather than simply accepting a number.

How to read a certainty level

A certainty level answers a narrow but important question: how confident should we be in the estimated effect for this particular outcome?

It does not tell us how large the effect is, whether an intervention should be recommended, or whether it will work for a particular person.

Very Low is a common and legitimate result. It does not mean the finding is false — it means the evidence available cannot carry much weight yet. The domain judgments say which weakness is responsible, and that is usually the useful part.

A separate agreement measure — how many qualifying studies point the same way — is available through the graded_agreement tool for callers who want it. It is not GRADE, its thresholds are still provisional, and it is deliberately not shown beside a certainty level here.

What each level rests on

  • High — further research is very unlikely to change confidence in the estimate.
  • Moderate — further research is likely to have an important impact and may change the estimate.
  • Low — further research is very likely to change the estimate.
  • Very Low — any estimate is very uncertain.
  • No evidence found — not a level. The question was asked and nothing published answers it.

Reproducibility, measured and published.

A transparent method should produce reasonably stable judgments when the same evidence is assessed again. H[1] therefore repeats assessments and measures where the results agree and where they diverge.

Body assessment

The same body of evidence, assessed three times, independently.

7 / 7

domains reproduced exactly — every downgrade and upgrade judgment, unchanged across all three runs.

Study-level judgments

The risk-of-bias layer the body assessment is computed over. Five studies, re-judged from their stored extraction.

35 / 40

domain judgments reproduced (88%). One disagreement crossed the line that excludes a study from the body — the kind of difference that changes a rating, not just the wording of a reason.