What we've learned so far at the Metascience Observatory
.. and why the heuristics you may be using to guess paper credibility are not actually useful
“Non-reproducible single occurrences are of no significance to science.” - Karl Popper
I founded the Metascience Observatory in October 2025 after receiving funding from an Astral Codex Ten Grant.
Previously I had heard about the replication crisis in psychology. That got started in the 2010s when people started finding that only about 50% of experiments published in leading psychology journals could be replicated.
The Metascience Observatory was founded to learn more about reproducibility across all of science. The initial idea was to search for papers that report the results from a replication experiment and collect those reports into a large dataset. Constructing this dataset remains a work in progress, but this post will summarize some of our major learnings so far.
Today the term “replication crisis” is largely associated with psychology in the popular imagination. However, evidence has been growing over the years that reproducibility is low in other areas of science like cancer biology, economics, preclinical biomedicine, and machine learning.
At the Metascience Observatory, we believe that the replication crisis is “science-wide” and that the association of the crisis with psychology is a “streetlight effect”.
When Nature surveyed 1,500 scientists across all fields of science in 2016, 70% reported that they had failed to reproduce a result. A similar survey of 869 scientists working in cell biology found a similar figure - 72% reported being unable to replicate a prior result.
The breakdown by discipline in the Nature survey is very interesting:
More scientists working in chemistry reported difficulty replicating experiments compared with medicine or “other”. As a former physics researcher, this does not surprise me that much. Materials science and the related field of condensed matter physics are well known to have reproducibility issues.
While 70% of respondents reported they had failed to replicate someone else’s result, only 30% reported publishing a replication result. Unfortunately, most replications are not published or publicly disclosed, and this is especially true for unsuccessful replications. We estimate that direct or “very close” replications are only present in about 0.5% of papers total.
The respondents had a wide range of opinions as to the overall reproducibility of papers in their field:
I imagine something similar holds for the public at large. Some people seem to be very trusting of scientific research, while others think it’s all crap.
One of the goals of the Metascience Observatory is to help people calibrate their views as to reproducibility. Having a good baseline degree of skepticism is important, especially when looking at findings in medicine.
Back in 2015 I wrote about Steve Novella’s concept of an “evidence threshold”. An evidence threshold is basically how much evidence you need before you start having confidence that a phenomenon is real. My hope is that the Metascience Observatory’s data on reproducibility will help people calibrate their evidence thresholds.
The Metascience Observatory’s replication database is also meant to be a resource for metascience research and understanding how reproducibility and rigor vary across different dimensions (time, discipline, country).
The replication database may also be useful as a research tool. We have a page where you can upload a bibliography file and see how many works have replications.
Defining replication

When people try to replicate a prior experiment as closely as possible (within reason), we call that a “direct replication”. To start, the Metascience Observatory was only interested in direct replications. Direct replications are very rare — it appears that less than 0.5% of published articles report about direct replication.
Due to the rarity of direct replications, eventually we started collecting three looser types of replication:
close experiment — Here scientists are testing for the same effect observed in a previous experiment, but they may make one or maybe two deliberate changes to the experimental procedure. Importantly, the thing being measured is the same.
close extension — Here scientists are testing whether an observed effect generalizes to a different setting, but only a small generalization. We found that when people talk about “replication” in genetics they often mean finding whether a genetic association generalizes to a different population. Close extensions are also very popular in psychology.
conceptual replications — The scientists are testing for the same effect observed in a previous experiment, but using a different experimental procedure.
By the way, here’s how replication rate in our dataset varies by replication type:
When researchers have more freedom regarding how they conduct the experiment, replication rates are higher. Please note that there is no hard boundary between these categories, so they are pretty loose, and an AI (Sonnet 4.5/4.6) did a lot of this categorization.
Currently, the Metascience Observatory does not include the following:
Technical replication - when raw data from an existing experiment is reanalyzed using the reported procedures to see if the same results can be obtained. If the authors provided their code, this might involve just running the code on the data. Ideally, this can be done with one click (a notion called “frictionless reproduction”). Technical replication is very important in areas like AI and machine learning research, but at least thus far we are not covering it.
Robustness checks / independent analysis - this is when the same data is analyzed by someone else. Somewhat confusingly, this sort of thing is often called a “replication” in Economics, and DARPA’s SCORE initiative called this “reproduction” (as opposed to “replication” where new data is collected).
“Internal replication” or “self replication” - this is where the same lab / group does a replication of a previous experiment they did.
How well does our method of assessing reproducibility work compared to other methods?
The best way to determine reproducibility is to take a random sample of papers, replicate a key experiment from each one, and then compare the results of the replication experiments to the original results. Immediately, however, one runs into an issue — you must have enough information about the original experiment to know how to replicate it. If you can’t get enough information on how the original experiment is done, you can’t assess reproducibility.
That problem turned out to be big in the Reproducibility Project: Cancer Biology (RP:CB). Initially papers were selected on citations — not exactly random sampling, but close. (As we show below, papers that are cited more are slightly less likely to replicate, but only slightly.) The researchers wanted to replicate all the experiments in 53 papers, but for about 75% of the experiments they couldn’t get enough information to do the replication, even after emailing the original authors. Their end result was still informative, but was based on a non-random sample.
The famous Reproducibility Project: Psychology (2012-2015) ran into similar issues. The project set out to replicate 100 experiments from three leading psychology journals that had all been published in 2008. To do this they had to match each experiment to a lab that was willing and able to replicate it. However, there were some experiments where they couldn’t find a lab that was willing and able to do the replication, often because the experiment required special equipment (like fMRI/ eye-tracking) or involved special populations (like infants, people with a particular disease). The authors of the project describe the resulting sample as a “quasi-random” sample.
The Metascience Observatory engages in a form of convenience sampling because we’re just agglomerating replication reports that are already out there in the literature. This sort of sampling is frowned upon in research because bias is very likely to creep in.
One can imagine many sources of bias that might creep in. Often one the first things new graduate students are asked to do is replicate a prior result. The goal is to prove that you’re able to get an existing “baseline” result. Once the existing result is replicated, then the idea is usually to move on to some sort of extension or related experiment. The baseline replication is often published alongside the new experiment to show that the lab is capable of doing things correctly. However, if the graduate student is not able to replicate the existing result, that often means the project is mothballed and the graduate student moves onto something else.
A unspoken assumption that is still pervasive through much of academia is that results coming out of prestigious labs are very likely to be real. So, if a graduate student can’t replicate the result from a prestigious lab, the assumption may be that the graduate student’s experimental setup is flawed in some way, not that the original result is wrong. So, showing you can replicate a result from a prestigious lab is considered favorable because it can boosts the credibility and image of your lab, while failure to replicate work from a prestigious lab may be considered a liability if published.
Another source of bias concerns how easy it is to replicate a finding. In recent years, publishing replications has become very popular in psychology, and they are growing in popularity in other fields. Easier experiments that can be run quickly are more likely to be replicated because the goal of academics is to get as many papers as possible per unit time. This potential source of bias is important since it’s intuitively plausible that easier experiments are more reproducible than complex experiments.
One might also theorize that people would be more interested in replicating more surprising and “big if true” findings, because for those findings it’s more important to know whether they are correct or not. However I’ve been told that’s not how things work, at least in psychology.
Note that we have several different forces potentially at work here, and cutting in different directions. If surprising findings are more likely to be replicated, then that biases the replications towards failing, if surprising findings are more likely to not replicate. If studies that come from prestigious labs are more likely to be replicated, that might bias those replications towards succeeding, if prestigious labs produce more reproducible work. And, if the studies that are easy to replicate are more likely to be replicated, that could bias those replications towards succeeding.
Because of all this I was very curious to do a comparison of replication rates obtained from initiatives that used some form of rough representative sampling with the “literature-harvested” direct replications in our database:
The comparison is crude, but encouraging (for details on which initiatives were included, etc, see this page). The numbers were closer than I was expecting.
We can also compare the “literature-harvested” rows in our database with not just direct replications but also “close” and “conceptual” replications as well. Then we have more data and can compare linguistics as well:
Here the numbers differ a bit more, but this is expected since “conceptual” replications have a slightly higher reproducibility rate than direct replications.
Replications and replication rate over time
According to the data we have so far, papers reporting replications were rare until around 2005.
The year 2005 is important because it’s when John Ioannidis published his paper “Why Most Published Research Findings Are False”, which is considered to have inaugurated the modern metascience movement.
All of the data we have so far indicates that the replication rate has remained remarkably stable over time, hovering around 50-60%:
However, we do find evidence of increases in replication rate in the last ten years or so. Recall that the replication crisis started in earnest in 2015 with the publication of the Reproducibility Project: Psychology. Since then there has been an uptick in replication rate in psychology:
This uptick may be due to reforms that have been pushed by the burgeoning open science movement. A 2023 paper in Nature Human Behaviour presented evidence that open science reforms increase reproducibility. The study examined 16 recent social-behavioral findings stemming from research that used open science best practices and found a higher replication rate among subsequent replications. Unfortunately, the paper was retracted since critics noticed issues with how their preregistration was done. Despite the retraction, I think that work is likely pointing at some real effects.
Effect size decline
For a lot of replications we were able to get both the original and replication effect sizes and associated p-values. Our AI pipeline still struggles with this a bit, so a lot of this data comes from the FORRT Replication Database (FReD), a large crowd-sourced initiative run by Framework for Open and Reproducible Research Training (FORRT). We gratefully thank them for making their dataset available for re-use under a CC-BY-4.0 license. Similar to what the FORRT team did, we try to convert different types of effect sizes (like Cohen’s d, etc) to a standard scale (details here).
Here “success”, “failure”, and “reversal” are determined by whether there was a statistically significant effect in the same direction as the original work.
As you can see, effect sizes in the replication are typically smaller than in the original (points are largely below the dashed line). What we’re looking at here is vaguely related to the well-known “decline effect”, where effect sizes decline over time in the scientific literature.
Reproducibility by discipline
It’s tempting to think that the low rate of reproducibility in psychology must be particular to that field, because human beings are uniquely hard to study and predict. Indeed, human behavior is deeply contextual and tied to culture, which varies over time and place. However, we suspect psychology may actually be one of the fields with the best reproducibility, owing to the simplicity of some many of the experiments. The aforementioned Nature survey seemed to suggest that. Here’s what we’ve found so far, looking at the paper level, not the effect level (details here):
Here’s a breakdown by subdiscipline:
Here’s a graph showing the decline in effect size by discipline, for disciplines where we have stats (see here):
Looking at potential correlates
People are very interested in correlates or “heuristics” or “credibility markers” for “are these findings actually true”, because most people do not have the time or expertise to thoroughly evaluate findings. People use things like journal, author, and university as heuristics. Unfortunately, work at the Metascience Observatory indicates that these heuristics are not very good. This is in line with previous work you can find scattered throughout the literature.
Citation count
This has been looked at before. A 2020 work tracked citations for the 100 papers studied in the Reproducibility Project: Psychology (2015). All of those papers were published in 2008. They looked at the average number of citations accrued over the next 10 years:
As you can see, papers that replicated received about as many citations as those that didn’t replicate, on average. Now we can repeat that analysis but with vastly more data (more info here):
The shaded regions are 95% confidence intervals. Findings that failed to replicate got cited a bit more, on average. Here’s another way of slicing the data:
Unfortunately, this anti-correlation may become more pronounced with the advent of AI-turbocharged papermills. Papermill papers (fake papers) are creeping more and more into prestigious journals and papermills are getting really good at citing their own prior works. People are increasingly paying for citations from papermills as well. The kind of people who pay for additional citations are probably not the most scrupulous researchers..
A recent preprint analyzed citation patterns in top cancer journals. The authors used a machine learning classifier to flag papers that might be papermill papers. Papers flagged with a “probability” > 0.9 (softmax, not a real probability) received twice as many citations in the first year and had 50% more citations on average:

By contrast, the flagged papers had fewer real readers as determined by online accesses and Mendeley reader count, which is very suspicious.
Journal impact factor and rank
Replication rate is not very well correlated with the “impact factor” of the journal (more here):
Note these are not the “official” Impact Factor (TM), which is a proprietary tool of Clarivate. However, it was calculated the same way using open access data from OpenAlex. The numbers are systematically compared to the official IF since OpenAlex also averages citations for editorials and letters to the editor, while tools like Clarivate Analytics and Scopus do not. Journals that contain a lot of letters to the editor, like Nature, thus have much lower impact factors here. We can also look at a scatter plot for journals where we have more than 15 papers:
We can also look at replication rate vs something called the “SCImago Journal Rank.” This is a ranking system similar to Google’s PageRank algorithm, where a citation coming from a high-prestige journal counts much more than a citation from a low-prestige journal:
We can also look at journal h-index (for the year the study was published):
Overall, these graphs all tell the same story - reproducibility doesn’t depend much on the prestige of the journal. If anything, reproducibility is slightly anti-correlated with prestige.
Our results are consistent with data showing higher rates of gene naming errors in more prestigious journals:
Other work looking at genetics publications showed more unusual results (“bias”) in journals with higher impact factors:

In neuroscience, researchers found no association between statistical power and impact factor:

I found these plots in Björn Brembs’ excellent 2018 article, “Prestigious science journals struggle to reach even average reliability”.
Author citations
What if a paper has prestigious authors?
The graph doesn’t change much if you plot replication rate vs the first author h-index, last author h-index, or max h-index across authors (see all those graphs here along with probit regression modeling).
These findings are consistent with prior research - in 2019 Mueller-Langer et al. found that the h-index of the top author had no effect on replication rate in a sample of 131 replications in Economics.
Here’s something else we found (more info here):
We try hard to avoid “self replications”, where a lab replicates their own work. However, author overlap turns out to be fairly common, we find it in about 25% of the rows in our dataset. Not surprisingly, when authors overlap replication rate goes up, but I was surprised by the size of the effect.
p-value
One of the things DARPA’s massive SCORE initiative looked at was using prediction markets to predict whether papers would replicate. They correlated the prediction market predictions with simple heuristics. Across twelve different types of heuristics, the p-value in the original work was most strongly correlated with the prediction market predictions:
(“Structured elicitations” were a type of group activity where they had a group of scientists get together to evaluate a work and pool their prediction into a single credence score. “A+” was a complicated machine learning system from Two Six Technologies that attempted to predict replication success.)
Inspired by this work, we looked at how reproducibility of effects correlated with the p-value reported in the original work. One of the issues we ran into is that p-value is often reported as something like “p < 0.05”. This practice is not very useful and many have called for it to be abolished — “p < 0.05” could mean p = 0.000001 or it could mean p = 0.045.
Because of this we ended up with two graphs - one for the explicit p-values and one for the p-values given with a “less than” symbol (read more about this here):
Here’s a summary table of correlates of reproducibility (more info here):
Where are we going from here?
Today, the Metascience Observatory has multiple workstreams. Here’s what we’re up to:
Replications database - what this post was about. Currently our system for searching for replications is pretty crude. We’ll be working on improving it and are optimistic we can get more replications from physics, chemistry, and AI/ML. We’ll also be doing more validation work on the AI pipeline, building off the validation work we did in February.
The Metascience Observatory Explorer - this is a dashboard aggregating open-source data on journals, publications, grants, and researchers. We currently aren’t working on it, but it remains accessible for users.
Bird’s Eye Reviews - this is a new form of literature review that presents high-level information on all papers published on a given topic, regardless of journal or publication language. While there are dozens of “AI for literature review” services, none do what Bird’s Eye Reviews do. Bird’s Eye Reviews are aimed towards metascientific understanding of the research landscape, and allow users to make filtering decisions. We currently have three reviews published: Long COVID, Antiviral Nasal Sprays, and Restless Legs Syndrome (RLS). For RLS we had some fun experimenting with AI-powered meta-analysis using Claude Code (Opus 4.5), building directly off the corpus of publications obtained during the Bird’s Eye Review. We created both a conventional meta-analytic pooling and a network meta-analysis. We haven’t yet done a detailed analysis of these artifacts but the overall ranking of drugs/treatments does match literature meta-analyses on RLS drugs pretty closely. We also build a “Bird’s Eye Review Studio” which is a web application that allows users to create their own Bird’s Eye Reviews and conduct AI-powered meta-analyses. If you’d like to be a beta tester, please reach out.
The Forensic Metascience Agent - This is actually a multi-agent AI pipeline where AI agents have access to 30+ tools for detecting statistical inconsistencies and data-integrity anomalies in scientific papers and their supplementary information. Some of the tools are discussed in James Heathers’ seminal work An Introduction To Forensic Metascience. The “FMA” is designed to compliment image-based detection tools like Proofig, Imagetwin, and ReviewerZero.ai and the copypaste detection system for scrutinizing open source datasets developed by Markus Englund at ScienceDetective.org. We’ve only run the agent on a handful of papers, but we’ve already found a lot of issues and mistakes in papers. The agent is a bit expensive to run, but we will be running it on a targeted selection of papers in biomedicine. We’re working in collaboration with Intellicat and will be using their API and their Content Credibility Index (CCI) and other metrics to help with targeting. We recently hired Greg Fitzgerald to help with this effort.
If you are excited by any of our initiatives and would like to volunteer to help us or see an avenue for collaboration, please reach out.
You can subscribe to our newsletter here.
Funding is critical for us to continue to be able to do this work and donations are greatly appreciated.






























