Pragmatic clinical trials may be more fragile than they appear
(SewCreamStudio/Shutterstock)
A new study suggests that the harder a clinical trial works to be flexible and realistic, the more its conclusion leans on statistical assumptions.
A sophisticated trial design does not automatically make its result more reliable. Pragmatic cluster-randomised trials, adaptive platforms and other modern designs carry an air of methodological rigour, and a result from such a trial can feel more trustworthy for it.
Often the opposite is true: the harder a trial works to be flexible and realistic, the more its conclusion leans on statistical assumptions that usually go unchecked. In a paper published in Statistics in Medicine, my co-author Rachael Morton and I show this concretely for pragmatic cluster-randomised trials: the same data can give a significant or a non-significant result depending on the analysis method, and the published report rarely tells you how fragile the finding is.
Cluster-randomised trials randomise whole hospitals, schools or clinics rather than individual patients. Australian researchers run many of them, frequently through NHMRC-funded programs. They suit real-world conditions: messy data, variable patient populations, and sites of very different sizes. Those same features make them statistically hard to analyse honestly.
Different results depending on different analysis methods
We reanalysed four published trials. In three, the statistically significant headline finding did not survive a check against a simpler, more robust benchmark.
One was a school-based trial of strength exercises, originally reported as significantly improving body composition. Its nine schools ranged from 30 to 149 students, and omitting any single school could swing the treatment-effect estimate by more than twofold. A robust reanalysis returned a non-significant result. A placebo check, reassigning treatment to schools where none existed, showed the original method declaring a positive effect 62% of the time, when it should have done so only 5% of the time.
This is not a failing of the original investigators. It reflects how clinical biostatistics has evolved: toward flexible models, including mixed-effects, generalised estimating equations, and increasingly Bayesian and machine-learning approaches, that promise efficiency but quietly depend on assumptions about how patients within a site are correlated. When those assumptions hold, the estimates are precise. When they do not, an estimate can look confident while resting on very little.
How to interrogate the research
What does this mean for you as a reader rather than a statistician? You cannot re-run someone’s analysis, but you can interrogate it before you act on it. Three questions do most of the work.
- How many clusters were there, and how uneven were they? A handful of sites of very different sizes is a warning sign, because one or two large sites can drive the entire result.
- Was the headline analysis checked against a simpler, more robust method? If a single modelling choice carries the conclusion and no robust comparison is reported, treat the p-value with caution.
- Do the robust and the more elaborate analyses agree? When a paper reports both and they point the same way, the finding is solid. When only the elaborate model is shown, ask why.
These questions map onto a four-step framework we propose for trialists, called CARE:
- Clarify how unevenly the sites contribute;
- Apply a robust baseline analysis (a generalised linear model with a jackknife variance estimator, available in standard R and Stata, and suitable for continuous, binary and count outcomes);
- Refine with more elaborate models only when justified; and
- Evaluate by presenting the robust and refined results side by side. It is designed to sit inside a standard Statistical Analysis Plan, so as a reader you can reasonably expect to see it.
This matters more as the ground shifts. The US FDA has issued draft guidance opening regulatory trials to wider use of Bayesian methods, and Australian regulators will likely face the same questions. Those methods are valuable, but accepting their output without a robust benchmark risks treating a modelling assumption as evidence.
The bottom line for readers: treat a fashionable trial design as a reason for more scrutiny, not less. Before a pragmatic or other design-heavy trial changes what you do for your patients, check whether its headline result was tested against a simpler, more robust method. If it was not, hold the conclusion more loosely than the p-value suggests. Novel designs are not the problem. Trusting them on reputation is.
The full paper, including the replication code and reanalysed datasets, is open access in Statistics in Medicine: https://doi.org/10.1002/sim.70610
Dr Sergey Alexeev is a Senior Research Associate at Nura Gili: Centre for Indigenous Programs, UNSW Sydney, and an Adjunct Senior Lecturer at the University of Sydney. He uses large-scale Australian datasets to evaluate population health trends and policy.
The statements or opinions expressed in this article reflect the views of the authors and do not necessarily represent the official policy of the AMA, the MJA or InSight+ unless so stated.
Subscribe to the free InSight+ weekly newsletter here. It is available to all readers, not just registered medical practitioners.
If you would like to submit an article for consideration, send a Word version to mjainsight-editor@ampco.com.au.
You may also like
VIEW MORENewsletters
Subscribe to the InSight+ newsletter
Immediate and free access to the latest articles
No spam, you can unsubscribe anytime you want.
By providing your information, you agree to our Access Terms and our Privacy Policy. This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.