Need for Psychometric Theory in Neuroscience Research and Training: Reply to Kragel et al. (2021)
Abstract
We applaud the forward-looking nature of the Commentary provided by Kragel et al. (2021) on our article, “What Is the Test-Retest Reliability of Common Task-Functional MRI Measures? New Empirical Evidence and a MetaAnalysis” (Elliott et al., 2020). We fully agree with their emphasis on the importance of avoiding overgeneralization when considering measurement reliability in taskfunctional MRI (task-fMRI). Because no single reliability estimate can capture the multitude of possible task-fMRI measures, statements such as “every brain activity study you’ve ever read is wrong” (Cohen, 2020) are misleading and unnecessarily undermine our joint efforts to improve task-fMRI. In fact, that is why we addressed this very point in our article (see p. 801). Nevertheless, we take this opportunity to clarify three subtle but meaningful ways that our perspective diverges from that promoted by Kragel et al. in their Commentary. First, as we embrace the future, we must also account for and build on the past while being realistic about the state of the present. Kragel et al. point out the exciting potential of “multivariate measures optimized using machine learning” that they claim are “commonly used for biomarker discovery” (p. 622). While we agree that multivariate measures are becoming more widespread and should continue to be developed and explored (see p. 802 of our original article), such measures are still far from being universal in task-fMRI biomarker research. Because psychological science is a cumulative enterprise, criticism and honest assessment of the current state of the science are essential to the continued advancement of the field. In this vein, we surveyed the reliability of region-of-interest-based task-fMRI activation, which is one of the most commonly adopted measures reported in the literature over the past 2 decades. Our meta-analysis directly provided evidence for this continued use, as approximately half of the reliability estimates we found had been published in the previous 5 years. These measures are not relics of the past; they are in common use today and still frequently incorporated as primary measures in largescale, state-of-the-art imaging efforts focused on biomarker development and individual-differences research. For example, the Human Connectome Project, UK Biobank, and the Adolescent Brain Cognitive Development study all have incorporated fMRI tasks designed to activate particular brain areas and circuits (Casey et al., 2018; Miller et al., 2016; Van Essen et al., 2013). These are large-scale, expensive projects creating MRI data sets for future neuroscience research. Thus, the poor reliability reported in our article is critical for not only past but also present biomarker research using traditional task-fMRI activation. We hope that by reevaluating standard practices in light of the reliability limitations detailed in our article, we can guard against repeating and perpetuating these limitations in such future research. Second, we would like to highlight an important distinction between the main aim of our article and several of the examples offered by Kragel et al. in their Commentary. The central concern addressed in our article was whether commonly used measures of task-fMRI activation are reliable enough for individual-differences research and brain biomarkers. To answer this question, 996665 PSSXXX10.1177/0956797621996665Elliott et al.Psychometric Theory in Neuroscience research-article2021
- DOI
- 10.1177/0956797621996665
- Sources
- semanticscholar openalex