DRIVER recommendations – Risk of bias
How to address potential sources of bias in the design of experiments and how to report this when publishing.
Item 2: Risk of bias
Bias occurs when errors in the experimental process lead to a systematic deviation between a study’s results or conclusions and the underlying 'truth'. The level of bias directly affects the reliability of a study – lower risk of bias generally means greater reliability. Experiments should be designed with potential sources of bias in mind, implementing measures to minimise them. These sources of bias and the steps taken to mitigate them should then be reported, enabling readers to assess the study’s reliability.
You can find helpful tips and additional context in the grey boxes in each section below.
Design recommendation
How to identify the key forms of bias affecting in vitro experiments.
How to apply bias mitigation methods to experiments including randomisation, blinding/masking and other design considerations.
What constitutes a proportionate approach to bias mitigation.
Identify potential sources of bias
Bias can enter experiments at any stage – from the initial experimental set-up through to data analysis and reporting. Even preconceived expectations about the outcomes of a study can influence how an experiment is designed or conducted. By identifying where and when bias may arise, researchers can choose methods or design features that reduce or ideally eliminate these sources of bias.
For in vitro experiments there are several key biases that should be identified and addressed if present:
- Allocation bias: Systematic differences introduced as a result of how experimental units are allocated to experimental groups – for example, applying treatment groups on one multi-well plate and controls on another, where plates may experience slightly different incubation conditions (e.g. CO₂ or humidity differences).
- Performance bias: Systematic differences introduced as a consequence of the way experimental groups are handled during an experiment – for example, inadvertently exposing samples from one group to longer periods outside the incubator during handling.
- Detection bias: Systematic differences between experimental groups that are introduced as a consequence of how outcomes are assessed – for example, analysing images from the control group first and, as fatigue sets in, applying less rigorous attention when assessing the later samples from the treated group.
Allocation bias typically arises during the set-up phase when units are assigned to groups; detection bias occurs when data are captured or analysed; and performance bias can occur at any point during the experiment. Because bias can emerge throughout the process, the entire experimental design – from set-up to final analysis – should be carefully examined for potential sources of bias. Once identified measures to address potential biases can be applied.
See the interactive content (i. Introducing the experiment) for a worked example on how these forms of bias can affect an experiment and influence the results.
Address bias where possible
Measures to address bias are known as bias mitigation. Bias mitigation methods vary depending on the potential type of bias and when during the experimental process they occur. In some cases, the same bias mitigations strategy can address different forms of bias when applied at different stages. It is important to first identify the source and type of bias before deciding on a mitigation method or strategy to apply.
See the interactive content (ii. Addressing bias) for more on bias mitigations and to see them applied in a worked example.
Bias mitigation measures
Using randomisation ensures that the experimental groups are, at the beginning of the experiment, as similar to each other as possible.
Randomisation is the process of assigning treatments to experimental units (e.g. wells, plates, or flasks) such that each unit has an equal probability of receiving any treatment. Without randomisation, treatment allocation may be influenced – consciously or unconsciously – by the researcher, introducing systematic differences between groups that are unrelated to the intervention. In in vitro experiments, this is particularly problematic, as factors such as plate position, evaporation, and handling conditions can affect cell behaviour and assay outcomes. Using randomisation to allocate experimental units to treatment groups minimises the risk of allocation bias by distributing these sources of variation more evenly across groups.
Performance bias can also be addressed using randomisation when multiple samples or experimental units are being treated, analysed or data is collected from them. In this scenario randomising the order the samples are processed prevents phenomena like pipetting drift or longer times outside the incubator to inadvertently affect one group more than another.
See interactive worked example (iii. Randomisation) for more information on how to apply randomisation using a worked example.
The use of inferential statistics (such as t-test or ANOVA) to analyse data is based on the assumption that experimental units are randomly allocated to treatment groups from a homogeneous background population, meaning that statistical analyses may not be valid on groups which have been allocated in a non-random manner.
Haphazard (or arbitrary) allocation is not true randomisation. Although assignments may seem random, the researcher can still influence them, introducing systematic differences between groups. For example, selecting multi-well plates “at random” from an incubator may consistently favour those at the front, which may have experienced slightly different conditions. This can introduce confounders and reduce the reliability of the results.
Masking or blinding is a strategy used to reduce bias by preventing researchers from knowing which experimental group each experimental unit belongs to.
Researchers may have expectations about the study's outcomes, which can unintentionally influence how experiments are conducted or how data are interpreted. For example, knowledge of treatment groups may affect how carefully samples are handled or how outcomes – such as cell counts or image analyses – are assessed.
Masking reduces this risk by keeping group allocation hidden, promoting consistent handling and assessment across all experimental groups. This is particularly important in in vitro systems, where small differences in handling or measurement can influence results.
Masking can address different sources of bias depending on the stage of the experiment at which it is used.
| Time point | Bias addressed | Why it matters |
|---|---|---|
During allocation | Allocation bias (when combined with randomisation) | Prevents systematic differences between groups before the experiment begins. |
During conduct of experiment | Performance bias | Prevents unequal handling of groups during the experiment. |
During outcome assessment | Detection bias | Prevents expectations influencing measurement or scoring. |
During data analysis | Outcome reporting bias (see Item 5: Experimental groups and exclusions) | Prevents selective analysis or reporting based on desired outcomes. |
Considering which forms of bias are most likely to occur or are likely to have the greatest impact on the experimental outcomes can help in determining where masking can be best applied.
For more on masking and how it can address different forms of bias see the interactive content (iv. Masking/blinding).
It may be difficult to mask group allocation throughout the entire study because of the experimental design or other practical factors (e.g. visible differences between experimental groups). However, masking can usually be implemented at some stage of the study. It is often easiest to apply during data analysis as it is less likely to introduce errors and there are fewer time pressures.
As well as the group allocation of each experimental unit, other important information to conceal includes any relevant properties of a sample, such as whether it was derived from a healthy or diseased source, along with any information that may provide clues about group allocation such as the date of sample collection (e.g. if disease samples are only collected on specific dates).
Many in vitro models are highly sensitive to small changes in their physical and chemical environment, such as temperature fluctuations or variation in media composition. Even minor differences can introduce bias if they create systematic differences between experimental groups – for example, a temperature gradient affecting plates on one incubator shelf more than another. Measures should be taken to minimise differences in conditions across groups, reducing the risk of introducing performance bias.
One strategy is to randomise or alternate samples to address the specific issue identified. For example, in cell culture experiments, wells on the periphery of a plate are often not as humidified as wells towards the centre, meaning that proliferation of cells in peripheral wells may be impaired relative to others. This can be addressed by randomising the position of samples from different groups (either simultaneously with random allocation to experimental groups or after allocation has taken place, depending on the experimental design) or alternating the position of the samples from each group.
Another strategy is to keep these factors constant for all groups. For example, in an experiment involving DNA extraction from tissue specimens, processing samples in more than one batch may introduce differences between batches, which could be misinterpreted as differences between groups and introduce bias (particularly if some experimental groups are overrepresented in some batches). This can be addressed by processing all the samples in a single batch, such that all groups are treated under the same conditions.
If logistical considerations mean that treatments or interventions must be carried out in multiple batches, samples from each experimental group can be split equally across each batch, so that each group is equally exposed to any differences caused by sample processing.
Detection bias can be introduced when outcome measures – particularly those requiring subjective assessment – are evaluated in a fixed order, by group (e.g. all control samples followed by all treated samples).
When assessing large numbers of samples, the assessor may become fatigued, reducing accuracy in later measurements. Alternatively, they may improve with practice, leading to more accurate assessments over time. In either case, measuring all samples from one group consecutively can introduce systematic differences between groups, biasing the results.
Additionally, researchers may be influenced by their expectations or a desire to observe a particular outcome. During sample assessment, this may lead to more time or care being devoted to certain experimental groups. This can introduce systematic differences between groups, resulting in detection bias.
This can be addressed by masking and either randomising or alternating samples from different groups so that the researcher assessing the outcome(s) cannot determine whether consecutive samples belong to the same group from the order of the samples alone.
Detection bias can be mitigated by using methods or equipment that are sensitive and appropriate to the range of measures expected in the experiment. Key considerations include:
- How accurate is the method?
- What is the range of measurements that provides reliable data?
- What is the variability associated with the measurements conducted?
- What is the sensitivity of the method or equipment used, relative to that range?
An in vitro method, including any associated equipment, should be developed and tested to ensure that the expected values for each outcome measure can be measured accurately within the defined lower and upper detection limits.
For a worked example demonstrating the effects of using uncalibrated equipment on data collection see the interactive content (iv. Using the right equipment).
When considering how to reduce detection bias, it is useful to evaluate whether automated techniques can be used. Automation can improve consistency by reducing subjective decisions-making and the potential for user errors during outcome assessment. For example, in a study assessing changes in cell size, automated imaging software can be used to measure cell area, reducing the potential for bias associated with manual measurements.
Balancing bias mitigation and practicality
It is important to balance the protection from bias offered by each technique with the practicalities of implementation. While measures to reduce bias represent best practice and some sources of bias can always be addressed, there may be valid reasons for not applying certain techniques in an in vitro study.
For example, completely randomising 96 samples across a microtitre plate may be impractical, as it could increase pipetting errors and significantly extend experiment time.
Similarly, masking may not be necessary when internal controls provide objective measurements. For instance, in a qRT-PCR experiment using absolute quantification, concealing sample group allocation may be unnecessary if reference standards are included on the same plate as experimental samples, as these allow for accurate, absolute measurement of target RNA concentration.
Designing an experiment that addresses every potential source of bias may be the aim in theory, but if it cannot be set-up and executed reliably, the effort is wasted.
Reporting recommendation
How to describe the types of bias considered in the design and conduct of each experiment.
How to report bias mitigation measures used in each experiment.
What constitutes appropriate justification for not addressing bias.
Specify all sources of bias identified
Researchers should clearly describe any potential sources of bias recognised during the design, conduct, analysis, or interpretation of the experiment. This transparency helps readers understand the context in which the experiment was performed and enables them to assess how reliable the findings are.
Key details to report include the type of bias identified, where in the experiment it is present and how it was addressed (see section below on reporting bias mitigation).
Potential sources of bias may include (but are not limited to):
- Selection bias (e.g. non-random assignment of experimental units).
- Performance bias (e.g. experimenters aware of treatment groups).
- Detection bias (e.g. unblinded outcome assessment).
Clearly identifying these issues does not undermine the credibility of the work; rather, it strengthens it by demonstrating careful consideration of methodological limitations.
Describe the measures used to mitigate or minimise bias
It is important to report a full and explicit description of every method used to reduce bias. Researchers should avoid relying on broad statements such as “samples were randomised” or “the experiment was blinded,” and instead describe how these procedures were implemented.
Examples of details to include:
- Randomisation: Method used (e.g. random number generator, stratified randomisation, blocking), what was randomised (e.g. plate positions, treatment order) and who performed it.
- Masking: Who was blinded, at what stage (data collection, outcome assessment, analysis) and how masking was maintained.
- Standardisation procedures: Whether protocols, equipment, timing or environmental conditions were controlled or harmonised.
- Automated or objective readouts: Use of instrumentation or software to reduce assessor subjectivity.
Providing these details prevents misinterpretation. For example, researchers often state that experimental units were randomly assigned to interventions without explaining how. Without this detail, what is described as randomisation may actually have been haphazard allocation, mistakenly assumed to be random.
Justify when bias mitigation was not implemented
It is equally important to state when measures were not used to address bias and to provide clear justification.
While bias should be minimised and addressed where possible, some experimental set-ups may make this difficult or impractical. Valid reasons may include:
- Safety considerations – for example, masking an operator to toxic or infectious treatments would pose unacceptable risk.
- Practical or technical constraints – for example, complete randomisation is not feasible across an entire 96-well plate due to the potential for pipetting error.
- Experimental design features that inherently reduce bias – for example, objective readouts using automated imaging, internal standards or on-chip controls making additional blinding unnecessary.
Clarifying these decisions ensures that readers, reviewers and future users of the data understand the practical context, the boundaries of the experimental system and any limitations introduced by the choices made.
Consider including a dedicated experimental design section within the methods. Important details about bias mitigation – both the measures that were applied and those that were not – may not naturally fit within traditional methodological reporting. A specific design section allows all key elements to be presented clearly, such as the experimental units for each experiment (see Item 1: Experimental unit), the experimental set-ups or layouts, potential sources of bias and the corresponding mitigation strategies. This approach signposts essential information for readers and reviewers, increasing transparency and supporting more informed interpretation of the findings.
- Yarborough M (2021). Moving towards less biased research. BMJ Open Science 5(1): e100116. doi: 10.1136/bmjos-2020-100116
- Pannucci CJ and Wilkins EG (2010). Identifying and avoiding bias in research. Plastic and Reconstructive Surgery 126(2): 619-625. doi: 10.1097/PRS.0b013e3181de24bc
- National Toxicology Program (2019). Handbook for Conducting a Literature-Based Health Assessment Using OHAT Approach for Systematic Review and Evidence Integration. National Toxicology Program.
- U.S. EPA Office of Research and Development (2022). ORD Staff Handbook for Developing IRIS Assessments.
- Schneider K et al. (2009). ToxRTool, a new tool to assess the reliability of toxicological data. Toxicology Letters 189(2): 138-144. doi: 10.1016/j.toxlet.2009.05.013
- Roth N et al. (2021). Development of the SciRAP approach for evaluating the reliability and relevance of in vitro toxicity data. Frontiers in Toxicology 3. doi: 10.3389/ftox.2021.746430
- Higgins JPT et al. (2019). Cochrane Handbook for Systematic Reviews of Interventions, 2nd edition. John Wiley & Sons.
- Altman DG and Bland JM (1999). Treatment allocation in controlled trials: why randomise? BMJ 318(7192): 1209. doi: 10.1136/bmj.318.7192.1209
- Nuzzo R (2015). How scientists fool themselves and how they can stop. Nature 526(7572): 182-185. doi: 10.1038/526182a
- Lazic SE (2016). Experimental Design for Laboratory Biologists: Maximising Information and Improving Reproducibility. Cambridge University Press.
- Maddox J et al. (1988). High-dilution experiments a delusion. Nature 334(6180): 287-290. doi: 10.1038/334287a0
- Begley CG and Ellis LM (2012). Drug development: raise standards for preclinical cancer research. Nature 483(7391): 531-533. doi: 10.1038/483531a
- Pamies D et al. (2022). Guidance document on good cell and tissue culture practice 2.0 (GCCP 2.0). ALTEX 39: 30-70. doi: 10.14573/altex.2111011
- Bal-Price A et al. (2018). Recommendation on test readiness criteria for new approach methods in toxicology: exemplified for developmental neurotoxicity. ALTEX 35(3): 306-352. doi: 10.14573/altex.1712081
- OECD (2018). Guidance Document on Good In Vitro Method Practices (GIVIMP).
- Reynolds PS (2019). Is it “random” or “haphazard”? Demonstrating effects of non-random allocation by simulation. In: JSM 2019 Proceedings. Joint Statistical Meetings, Denver, Colorado, 27 July to 1 August 2019. American Statistical Association.
A set of six items tailored to the design and reporting of in vitro experiments. Find out more on the landing page.
Go to Item 1: Experimental unit.
You are on this page.
Go to Item 3: Experimental model.
Go to Item 4: Experimental procedures.
Go to Item 5: Experimental groups and exclusions.