How likely a study's design flaws are to have distorted its result.
Risk of bias is a structured judgement about whether flaws in how a study was designed or run could have systematically distorted its findings, as opposed to the findings simply reflecting normal statistical uncertainty. It is assessed at the domain level — for example randomisation, deviations from intended treatment, missing outcome data, outcome measurement, and selective reporting — and rolled up into an overall rating, typically low, some concerns, or high.
A study can be adequately powered and report a large, statistically significant effect yet still be misleading if, say, allocation wasn't concealed, outcome assessors weren't blinded, or only favourable outcomes were reported. Risk of bias assessment guards against taking effect sizes at face value when the internal validity of the comparison is compromised, which is a distinct problem from imprecision or lack of generalisability. It is what allows a reader to separate 'this is a true effect' from 'this is an artefact of how the study was conducted'.
In a paper it shows up in the methods (sequence generation, allocation concealment, blinding of participants/assessors, handling of dropouts) and in discrepancies between the registered protocol and reported outcomes. It is assessed using standardised tools matched to design — RoB 2 for randomised trials, ROBINS-I for non-randomised/observational studies — each working through fixed domains with signalling questions to reach a per-domain and overall judgement, usually done independently by two reviewers.
Domain-level judgements still involve reviewer interpretation, so ratings can vary between assessors and tool versions, and 'low risk of bias' reflects internal validity only, not relevance to a given patient or clinical importance of the effect size. It is commonly misused when an overall rating is taken as a single pass/fail label rather than examining which specific domains were flagged, or when a high-risk-of-bias study is dismissed outright rather than downweighted alongside other evidence.
This guide was auto-drafted and is pending editorial review.