Skip to main content

Journal Accountability: When Retraction Counts Should Cost a Journal Its Standing

Picture

Member for

1 year 9 months
Real name
The EduTimes Editorial
Bio
The EduTimes Editorial

Modified

Retractions rose fourteen-fold since 2013 with almost no consequence
A few repeat journals drive most paper-mill retractions
Retraction thresholds, not embarrassment, should trigger public review

Publishers issued more than fourteen thousand retraction notices in 2023 alone, a global record that fell within a year. Nine thousand more retractions followed in 2024 and by late summer 2025 the yearly count had already crossed five thousand. A decade earlier, the annual total sat near one thousand. The machinery meant to catch bad work, journal review, editorial oversight, database indexing, kept waving weak papers through until public pressure eventually forced a correction and that climb traces back to that failure more than to any sudden drop in scientific care. The pattern shows up across fields, from cancer research to computer science and it looks like a structural gap, not a string of isolated scandals. Journals that publish poor work rarely lose much for it. Editors move on, publishers keep their subscription revenue and a new batch of rushed submissions tends to arrive before anyone has tallied the damage from the last one. Journal accountability, the idea that a publication's track record should carry real consequences, offers a way to close that gap.

[FIGURE 1 HERE] Caption: Retraction notices issued per year, 2013 to 2025. The 2023 spike traces to a single mass event at Wiley's Hindawi journals, not a one-year jump in misconduct.

A Flood That Outpaced the Gatekeepers

Submission numbers alone explain part of the strain. NeurIPS, one of the largest AI research conferences, received roughly 3,240 papers in 2017; by 2024 that figure had grown to 15,671 and it climbed again to 21,575 in 2025. A sister conference, ICLR, saw submissions rise nearly sixty percent in a single year between 2024 and 2025. Reviewer pools were not expanded at anything close to that pace, so conferences leaned more on younger, less experienced reviewers working under tighter deadlines. A widely cited 2021 experiment tested what happens when the same batch of accepted papers goes through review a second time: roughly half of the papers that had been accepted the first time were rejected on the second pass, a gap too large to write off as noise. Acceptance at a top venue has become something closer to a coin flip than a dependable quality signal and this problem is far from unique to machine learning.

Figure 1: Retractions by year. The 2023 spike was one Hindawi event, not a trend

The same overload shows up well outside computing, too. Cancer research has drawn its own scrutiny, with editors describing a wave of papers that follow templates common to so-called paper mills, operations that manufacture fake or recycled studies for a fee. Editors across disciplines have reported feeling overwhelmed by the sheer volume of submissions competing for a shrinking pool of qualified reviewers and that complaint has been echoed in surveys of working scientists over the past two years. None of this is confined to one field or region. Countries that saw the sharpest rise in publication output over five years also saw the sharpest rise in retractions, a pattern that has been linked to incentive systems rewarding paper counts over paper quality. Volume, without a matching investment in review capacity, produces more or less the outcome now unfolding across the literature.

Publishers Face Few Real Consequences

Wiley's Hindawi division shows just how thin that accountability really is. In 2023, the main academic indexing service pulled nineteen Hindawi journals from its master list after a paper-mill scandal came to light and the fallout eventually produced more than eight thousand Hindawi retractions that year. Wiley ultimately retired the Hindawi brand entirely instead of trying to repair it. Chemosphere, a journal with a respectable impact factor of 8.1, lost its place in that index in December 2024 after reviewers flagged a cluster of suspect papers tied to special issues. Cureus was put on hold in September 2024 and formally removed roughly a year later. Elsevier's Scopus database, a separate index, removed fifty-six journals across 2025 alone, on top of hundreds cut the year before.

Figure 2: Half of 142 journals had one retraction; four had 100+.

By the time anyone acted in most of these cases, the damage had already been done. Papers had circulated, been cited and in some instances shaped hiring and funding decisions long before anyone stepped in. Once a journal is placed on hold for review, about eighty-five percent are eventually delisted and most of those holds are resolved within roughly six weeks, which is a strikingly high confirmation rate for such a short process. Reviewers can evidently tell a troubled journal from a sound one fairly quickly once they look closely. That raises an uncomfortable question: why does that closer look tend to happen only after years of harm have already piled up? A functioning system would flag the warning signs while a retraction count is still climbing, not after a publisher has already built its reputation on compromised work.

Building Real Journal Accountability

Indexing databases already track retraction rates for every journal they cover, so the raw material for a fix already exists. What's missing is how visibly and how early that information gets used, instead of being treated as a bad year to manage quietly behind the scenes. A journal whose retraction rate crosses a defined threshold, relative to its output and sustained over a rolling period, not just a single rough year, could be made to face automatic public review rather than a discretionary one that depends on somebody noticing first. Springer Nature has already shown this kind of reporting is achievable at scale: the publisher disclosed that it retracted 2,923 of the roughly 482,000 articles it put out in 2024. Making that kind of disclosure mandatory across major publishers, not merely voluntary, would let indexing bodies and funders act on real-time signals rather than wait years for outside research-integrity sleuths to surface the same problems.

Some editors would object, reasonably, that retracting a paper is a sign the system is working rather than failing and that punishing journals for retractions could push them toward burying problems instead of fixing them. There is something to that. Indexing bodies have said as much themselves: journals are not penalized simply for retracting flawed work, since that reflects an editorial process doing its job. What should matter is what happens after a problem surfaces publicly and then goes unaddressed for months or years. A threshold built around that kind of unresolved pattern, not the act of retracting itself, protects honest correction while still catching journals that let obvious problems sit.

A version of the review mechanism needed here already exists, if only in rough form. The main index already places troubled titles on hold before a final delisting decision and Scopus runs its own monthly exclusion reviews through an advisory board. A journal typically sits on hold for a matter of weeks before a call gets made, so the review itself moves fast enough once it starts. The slow part is deciding to open that review in the first place, since that decision currently depends on outside pressure rather than any standing trigger tied to data publishers already report. Moving from a complaint-driven model to a threshold-driven one is less a matter of new technology than of commitment. Indexing bodies would mostly need to use numbers they already collect, on a fixed schedule, instead of waiting around for a scandal to force their hand.

There's also a fairness concern worth taking seriously: researchers in countries with fewer resources for review infrastructure might be hit hardest by tighter standards, since national assessment systems there often lean heavily on indexed publication counts. That risk deserves to be named, not brushed aside. But the current system already produces roughly that outcome, only later and more destructively. Ethiopia has recorded the highest retraction rate of any country in recent years, with Saudi Arabia, Pakistan, China and Egypt not far behind, largely in cases where publication output rose sharply without a matching rise in review capacity. Delayed accountability doesn't spare researchers working inside these systems. It lets flawed work accumulate under their names for years, until a delisting eventually invalidates credit they had already built careers on. Catching problems earlier and calibrating the response to the pattern rather than a single incident protects solid careers instead of punishing entire national systems after the fact.

University administrators and funding bodies have a direct stake in all of this. Publication counts from journals under active integrity holds shouldn't carry the same weight in tenure files, grant renewals or doctoral defenses that they do now. One review of universities with heavy reliance on since-delisted journals found that some institutions had drawn more than three-quarters of their output from venues later cut from major indexes, before those same institutions reformed their practices sharply once scrutiny arrived. Behavior clearly can change fast when the incentive changes. Policymakers overseeing national research assessment systems, several of which tie funding formulas directly to indexed publication counts, hold similar leverage and building a rolling retraction check into those formulas would push the incentive back toward care over sheer volume, using tools that already exist rather than inventing new bureaucracy.

The Standard the Record Demands

A fourteen-fold rise in annual retractions since 2013 is the story of modern research, not a footnote to it. Editors, indexing bodies and funders each hold a piece of the fix and each one has mostly treated it as somebody else's job to act first. Real journal accountability doesn't call for new institutions or elaborate enforcement regimes. It calls for publishers to report retraction data as routinely as they already report impact factors and for universities and funders to weight that data the way they already weight citation counts. The tools are already there. What's been missing is the will to use them before the damage compounds, not after. Every extra year of inaction adds another cohort of papers and another set of careers, built on work that won't hold up. The record is public already. It's time the incentives caught up with it.


This article reflects the analytical judgment of The EduTimes Editorial Board and does not constitute policy advice or the official position of any affiliated institution.


References

Beygelzimer, A., Dauphin, Y., Liang, P. and Wortman Vaughan, J. (2021) 'The NeurIPS 2021 consistency experiment', Neural Information Processing Systems Blog.
Casrai (2026) Why journals get delisted from Web of Science. Ottawa: CASRAI.
Chemistry World (2025) 'High profile chemistry journal removed from Web of Science index', Chemistry World, 28 January.
C&EN (2025) 'Ethiopia has highest rate of scientific paper retraction', Chemical & Engineering News, 23 January.
Ioannidis, J.P.A., Pezzullo, A.M., Cristiano, A., Boccia, S. and Baas, J. (2025) 'Linking citation and retraction data reveals the demographics of scientific retractions among highly cited authors', PLOS Biology, 23(1), e3002999.
Kim, J., Lee, Y. and Lee, S. (2025) 'Position: the AI conference peer review crisis demands author feedback and reviewer rewards', in Proceedings of the 42nd International Conference on Machine Learning. PMLR 267.
Retraction Watch (2024) Web of Science puts mega-journals Cureus and Heliyon on hold. New York: The Center for Scientific Integrity.
Retraction Watch (2025) Springer Nature retracted 2,923 papers last year. New York: The Center for Scientific Integrity.
Su, B., Zhang, J., Collina, N., Yan, Y., Li, D., Cho, K., Fan, J. and Su, W. (2024, revised 2025) 'The ICML 2023 ranking experiment: examining author self-assessment in ML/AI peer review', arXiv preprint arXiv:2408.13430.
Unnamed author(s) (2025) 'Gaming the metrics? Bibliometric anomalies and the integrity crisis in global university rankings', arXiv preprint arXiv:2505.06448.
Unnamed author(s) (2026) 'Recommending best paper awards for ML/AI conferences via the isotonic mechanism', arXiv preprint arXiv:2601.15249.
Van Noorden, R. (2023) 'More than 10,000 research papers were retracted in 2023 — a new record', Nature, 624(7992), pp. 479–481.

Picture

Member for

1 year 9 months
Real name
The EduTimes Editorial
Bio
The EduTimes Editorial