Are these data real? Statistical methods for the detection of data fabrication in clinical trials.

Sanaa Al-Marzouki, Stephen Evans, Tom Marshall, Ian Roberts

Journal: BMJ (Clinical research ed.) 2005;331(7511):267-70

PMID: 16052019

Abstract

OBJECTIVES

To test the application of statistical methods to detect data fabrication in a clinical trial.

SETTING

Data from two clinical trials: a trial of a dietary intervention for cardiovascular disease and a trial of a drug intervention for the same problem.

OUTCOME MEASURES

Baseline comparisons of means and variances of cardiovascular risk factors; digit preference overall and its pattern by group.

RESULTS

In the dietary intervention trial, variances for 16 of the 22 variables available at baseline were significantly different, and 10 significant differences were seen in means for these variables. Some of these P values were extraordinarily small. Distributions of the final recorded digit were significantly different between the intervention and the control group at baseline for 14/22 variables in the dietary trial. In the drug trial, only five variables were available, and no significant differences between the groups for baseline values in means or variances or digit preference were seen.

CONCLUSIONS

Several statistical features of the data from the dietary trial are so strongly suggestive of data fabrication that no other explanation is likely.

Address: Department of Epidemiology and Population Health, London School of Hygiene and Tropical Medicine, London WC1E 7HT.
Bant logo

© Copyright 2026, Nutrition Evidence

NED wishes to thank the following organisations for their support:

We use cookies to improve your experience and analyze site traffic with Google Analytics. By continuing to use our site, you agree to our use of cookies. Learn more.