36 Comments
User's avatar
David W. Zoll's avatar

Love the idea of preregistration. Knowing that a study failed could be as important as knowing it succeeded!

Geeta Nadkarni's avatar

Severe kudos for how well researched and thought out this article is. I shared it with several friends who work in psychology. I’m a subbie for life!

Alex Mendelsohn's avatar

Perhaps the most surprising (and concerning) thing about reading medical literature is how few studies share any raw data. So many seem to just give averages. I am aware that the data would have to be anonymised. And perhaps there are other reasons I'm not aware of, like proprietary reasons?

While physics has more raw data sharing, I was still quite concerned about the paucity of studies that shared their raw data.

It is much more difficult to hide poor scientific practice with raw data. The Data Colada blog series on Francesca Gino's raw data is a good example.

Tommy Blanchard's avatar

Agreed. If medicine is anything like psychology, the reason for not sharing the raw data is pretty straightforward: there's no incentive to. Journals rarely ask researchers to, and people rarely look at it, so it's just more work to put the data in a presentable format, write up an explanation of the data, etc.

The incentives need to be aligned, which means, again, attaching prestige to studies that share data and having prestigious journals require it.

Alex Mendelsohn's avatar

That is a good point and makes a lot of sense Tommy - though I would add when I was submitting raw data as a physics researcher, it was as raw as it is possible to be raw. In fact a lot of the time it would literally be a .raw file! Not presentable, no explanation attached, just a data dump into a repository and the paper would have a link to it at the end. Unlikely to be looked at, but it was a more of a full transparency nothing to hide type gesture. Though I was surprised how frequently I would use something like a dataset for a paper no one had cited 15 years ago - plenty of PhD students out there putting off writing their theses!

I do like the idea that the incentives need to be aligned. Thinking out loud how would one attach prestige to studies that preregister and share data? Would it be something like a dual-pronged approach - holding journals to account through something like a retraction-watch (https://retractionwatch.com/) table, and then studies into what things produce papers more likely to be reproducible and higher quality? I mean, I'm making the assumption raw data would help - but I don't actually know...

Meandering thoughts aside, good article Tommy - it was a pleasure reading it!

Tommy Blanchard's avatar

The biggest area of leverage for aligning incentives IMO sits with editors. Make it journal policy that studies need to make data publicly available (some journals already do this, so there's precedent, and as you point out it doesn't need to be onerous). Preregistration is a bit harder since it is more of a structural change, but if the big journals and their editors at least put the pieces in place (making a big deal about submitting pre-registration to them, clearly mark papers that were pre-registered), then the community could align around pre-registered studies holding more weight, things will shift in that direction. There's already some movement in this direction, e.g. https://www.nature.com/nature/for-authors/registered-reports, we just have to push much harder in that direction

Alex Mendelsohn's avatar

You make a lot of sense, and based on your arguments, I would be inclined to agree that the onus lies with the editors. You can count me in on the push to make preregistration and raw data hold more weight. I'm not entirely sure what that would entail, but I'm in! Thanks for the conversation, Tommy. I really enjoyed it. Happy holidays!

Becoming Human's avatar

When we were doing landing page testing for marketing, which generally has a high participation, but also massive amounts of noise, we would find people stopping tests when the version they liked happened to be winning.

We referred to it as “statistical convenience”

Elizabeth Hamilton's avatar

Pregistration is a good idea but it's also true that sometimes the data can answer a question you didn't ask. In which case, I suppose that is worth another string if experiments.

Tommy Blanchard's avatar

I'm all for using the data you have to try to answer additional questions. This can be a great way to generate new theories and hypotheses and is more economical than running a new experiment for every new idea -- but these should be explicitly labeled as exploratory analyses and the results held as more tentative than pre-registered ones!

Lance S. Bush's avatar

You can report unexpected findings even with a preregistration; the thing to do would just be to make clear that this isn't something you preregistered.

Carl V Phillips, PhD's avatar

Very nice summary of some aspects of this problem. I first published about the problem (I focus on epidemiology, but it is much -- not entirely -- the same) over 20 years ago: https://link.springer.com/article/10.1186/1471-2288-4-20 I called it "publication bias _in situ_" because it is a form of publication bias and takes place entirely within a single study. (That is a more accurate cute terminology variation than "researcher degrees of freedom" which is not exactly a correct analogy to the concept, though I fully understand why the latter caught on better. Mostly I just use "model shopping" now, which captures the idea in an active voice without inviting complicated analysis of whether it is a valid term.) I am currently working on producing an engine to better demonstrate the magnitude of effects.

Epidemiology is a science of measurement, not hypothesis testing. So the "p hacking" version of model shopping that you emphasize is not the real problem. Psychology is rather behind epidemiology (which is a truly damning statement, given the state of epidemiology) in that it seldom recognizes that merely establishing the existence of something (valid in physics and some other sciences; wildly wrongheaded in social sciences) is not useful. I am absolutely sure that the effect of listening to the Wiggles on feelings of oldness is nonzero -- if the putative causal pathway has any chance of having an effect (as is the case) then the probability the effect nets to exactly zero is zero. So the question should be "how much effect does it have", but since psych research methods are incapable of answering that (for multiple reasons), the real question gets ignored. In epidemiology (and as would be the case in psychology if it were less primitive), it is the quantification of the effect that matters, not frequentist random error statistics, and that is even more biased by model shopping. Frankly it does not matter if we believe there is a "statistically significant" effect of Wiggles listening when really we should conclude there is a non-SS effect.

I am also currently working on a guide to writing study protocols (it will be policy for one journal, and a recommendation for others). The main motivation for that is one that means your partial solution of published protocols might not solve the problem as well as it could: Most every protocol you see fails to specify enough details about the model to be run to preclude model shopping. There are stupid levels of detail about how measurements will be done, which are obvious and/or inconsequential, but only vague statements of what the main statistical model will be. Pretty much the only solution is to actually provide the main equation that will be run as "the" result, with all accompanying details about functional forms to be used and such, which pretty much means "simulate a copy of the data and provide the code that runs on it to produce the main result". This is vanishingly rare, even in those "prestigious" protocoled studies.

Jonathan Tonkin's avatar

Nice one Tommy! In ecology, we had our own major drama over fabrication a few years ago. Search Jonathan Pruitt. He seemed equally unready to take blame despite dozens of retractions after detailed research into his papers and the fallout for all his coauthors was huge. It even got the name “Pruittgate”.

Preregistration has been called for for a while. I like the idea. Probably works best for clinical stuff or highly controlled experiments though.

The People Geek's avatar

Great article. Quantitative research is fraught with paradoxes. Likert scales, are they ordinal or scaled? What if I do put a number in front or not, horizontal or vertical displayed answer options? Negative or positively worded questions. The degrees of freedom are endless…

Scott Ko's avatar

Tommy, I frequently work in the field of management and leadership, and have been exposed on numerous occasions to academic research there. It feels like just an absolute wild west of questionable claims, research practices, and irreplicability. I've not gone to the depth of research you have on psych research but when I read the papers, so many of them just don't 'feel' right. There are so many contributing, subjective variables to how someone performs at work or the effects of 'leadership' that frequently aren't duly acknowledged or accounted for that often I don't understand how they've reached the conclusions they did.

The People Geek's avatar

Leadership is especially hard. As much of it is about influencing a dependant variable- performance, that there is no universal agreed definition of.

Lance S. Bush's avatar

By the way, this is worth checking out if you're interested in replication crisis stuff:

https://www.speakandregret.michaelinzlicht.com/p/revisiting-stereotype-threat

...As is abundantly clear: replication issues are ongoing, and major findings are still crumbling.

Tommy Blanchard's avatar

Et tu, stereotype threat?

Dora-Lynn Greene's avatar

Thank you for the depth you provided. I have read somewhere a similar article about this. I am not schooled in this area, but I have always gravitated towards the "scientific spectrum" and as a lay person I place more value on a scientific result than just an "opinion" or "belief" of someone. I also have noticed how science gets adjusted over time as the things that we can comprehend or access become "better" or different. For me personally I think we are all better off because of the people who work on things such as this. Also when a person is presented a scientific conclusion most people do not consider what goes into getting to this conclusion. I know that there are very strict and stringent frameworks for experiments, and if we can account for every instance of possibilities, to create a solid conclusion we must do so. If only because of the reach that comes from the presentation, we cannot be sure where or what bedrock the presentation will become part of. Meaning we cannot know in advance how the presentation will be used to create other conclusions. It is a very difficult field. Mind boggling in a way.

Lance S. Bush's avatar

I work in psychology so this article addresses issues in my field. I started my PhD in 2015, a few years into the maelstrom at a time when the replication crisis was the big topic. Things seemed to have quieted down since then. A few changes have been introduced, but I haven't seen the kinds of revolutionary shifts in the field I'd have hoped for.

If anything, I regard the situation as far more bleak. Even if we set aside fraud cases, which are significant but not as big of a contributor to the problem of low replication rates as p-hacking and other bad research practices, there are still deeper issues. Here are a few:

(1) There is a lack of any substantive, unifying theories grounding a great deal of research. One might say (and there are papers on this) that we have a Theory Crisis.

(2) Many studies have extremely poor generalizability. Call this the Generalizability Crisis (see: https://www.cambridge.org/core/journals/behavioral-and-brain-sciences/article/abs/generalizability-crisis/AD386115BA539A759ACB3093760F4824)

(3) There are problems of measurement and validity. So we may say we have a "validity crisis" (see: https://replicationindex.com/2019/02/16/the-validation-crisis-in-psychology/).

My own work focuses on problems of measurement and validity in the instruments used to assess how nonphilosophers think about philosophical questions, primarily moral realism and antirealism. I believe available evidence hints that the measures we use for these purposes are generally invalid. So, even if they replicated, it would be moot: they aren't measuring what people think they are, and cannot serve the empirical purposes they're put to.

This highlights a more insidious problem then low replication rates: studies can be uninformative or misleading even if they replicate.

Chris Mah's avatar

I'm not a psychologist nor a psych student but have long been sceptical of the use of the scientific method for psychology. Especially as I see extremely clomplex theories presented.

It seems to me that due to the inherent diversity and difference of people even within the same family, let alone those of entirely different cultures separated by distance and/or time. To many variables to account for. And if you do the observations made will have to be of very minor observations.

I'm not saying it's not possible but sometimes psychology and the related sociology just sound like science coated philosophy to me.

Michael Fuchs's avatar

I think the revolutionary improvement would be to publish null results. A map that shows known dead ends is much more useful. How to do that? A policy by major journals to reserve 25% of space to null result studies? Prestigious new journals in each field dedicated to null results? Scholarly conferences? Prizes for most important null results?

Mark Connolly's avatar

The XKCD cartoon was perfect, and I thought the term Hopium was a brilliant name for an addictive factor in research. It seems that the actually fraudulent research papers should be immortalized in some way as cautionary tales so when something is asserted people have a healthy dose of skepticism. Like a vaccine...

Ruv Draba's avatar

Well researched, synthesised and presented, Tommy.

There are plenty of ways to 'paint the target around the arrow', and some good ways to prevent it.