A key question for any evaluation process is whether it actually changes anything. Do researchers engage with the feedback they receive — and do they revise their work as a result?
We have begun tracking this systematically across our portfolio. The results so far are available here:
Did Authors Respond to Unjournal Evaluations? (interactive table and analysis)
Select text and open the Hypothes.is annotation sidebar from the upper-right edge of the page to leave an inline comment. We especially welcome corrections to document/version matching, missing public revisions, or our interpretation of whether a change reflects a suggestion. A free Hypothes.is account is required to post.
What we found so far
The April 2026 tracking file covers 57 evaluations, but adjustment status had been assessed for 22 rather than the full set. Within that partly assessed snapshot:
- 19 papers have a recorded formal or informal author response to the evaluation
- 16 papers have a recorded formal written response
- 5 papers are coded as showing clear evidence of substantive updating after the evaluation
- 7 authors are coded as stating an intention to update
- 15 papers are coded as showing at least some positive signal — evidence of updating, a stated intention, or minor revisions
The remaining 35 papers did not yet have an assessed adjustment status. Response status was also incomplete: 37 entries were blank in the April data. These categories describe recorded evidence; they do not establish that an evaluation caused a later revision. Completing and re-validating the full-set staff assessment remains an ongoing goal.
Improving the comparison checks
The earlier automated comparison was useful for finding candidate cases, but it was not enough to support causal claims. A newer PDF or a large text difference can reflect an unrelated revision, a different document, or a timing mismatch.
As of July 18, 2026, we have refreshed the public sources and are using a more conservative workflow:
- 32 current public paper versions have been captured;
- 13 have a preserved before-and-after comparison record;
- nine pass document, identity, evaluation-mapping, and timing checks and have received focused review; and
- four are excluded because the supposed later document predates the evaluation.
These are workflow-coverage figures, not findings about the effect of evaluations. For each candidate, we now preserve PDF snapshots and hashes, check that the before and after documents are versions of the same paper and are in the right order, and create page- and section-level change cards. Any claim that evaluator feedback contributed to a change still requires human review.
We have also added a no-model suggestion-to-change screen. It checks that each evaluator suggestion has a valid source quote and line anchor, looks for related new material in the later paper, checks whether the same concept was already present anywhere in the earlier paper, and compares the match with cross-paper placebo matches. It produces a review queue, not an influence label.
This stricter check changed the automated takeaway. A first card-level pass found three medium-priority lexical matches. After adding the whole-earlier-document check, all three were downgraded: the final pass found 17 low-priority leads and no medium- or high-priority link. One apparent Urban Forests match, for example, disappeared because the earlier appendix already tested alternative upwind-cone definitions.
The remaining low-priority leads are still useful for directing manual reading. The macro-climate revision adds an appendix called “Additional Robustness Checks,” which is consistent with a broad evaluator request for more checks but does not identify the evaluation as its source. The resilient-education revision adds questionnaire and protocol material about speakerphone use and other children. This documents the possible spillover issue, but it does not show that only one student benefited.
By July 18, we had completed focused reviews of nine candidates. These include one directly documented correction, two reasonably specific content alignments with attribution unconfirmed, three weaker correspondences, and three cases with no supported substantive match. Even the stronger content alignments do not establish attribution: the papers added relevant analyses, but we do not have author responses or revision memos saying that the evaluations prompted them. The interactive analysis gives the case-level evidence, including an apparent match that was rejected after manual checking showed the material was already present in the earlier paper.
We also repaired two stale endpoint selections. One produced the clearest result in this re-review: a public author response explicitly says that the evaluation caught a significant proof error, and the maintained document marks that error and points to the correction. For the other, The Pivot Penalty in Research, we replaced the coarse change card with a section-aligned manual comparison of the evaluated manuscript and supplement against the later main text and supplementary material. Seven of eight suggestions were not implemented; the eighth received at most a weak framing acknowledgment. Funding checks, field-based pivot measures, alternative citation windows, and sequential-dynamics discussion all looked potentially responsive at first but were already in the evaluated version.
We also audited all 17 cases where the initially selected input and public PDF were byte-identical. Thirteen now support the limited finding “no observed public document change as of July 18, 2026.” For each of these, we found either a pre-evaluation archived PDF or distinctive page, table, figure, or quotation evidence that binds the evaluation to the unchanged official document. Of the other four, two were repaired into the comparisons described above, one revised endpoint remains unavailable, and one record had no completed evaluation. We do not interpret an unchanged public PDF as evidence that authors ignored the evaluation.
The latest refresh did not change the 22-case reviewed denominator. It recovered two more distinct PDFs, but one was dated October 2024 against a June 2025 evaluation and the other November 2022 against an April 2023 evaluation. Counting either as a post-evaluation update would have biased the conditional response rate upward.
Overall rates — and the denominator that matters
Across 22 reviewed public endpoints, the mutually exclusive outcomes are:
- 13 (59%) with no observed public document change;
- 3 (14%) with a later version but no supported substantive implementation of the evaluated suggestions;
- 3 (14%) with weaker correspondences;
- 2 (9%) with reasonably specific alignments but no confirmed attribution; and
- 1 (5%) with an explicit public author link between the evaluation and a correction.
Thus six of 22 (27%) show at least some correspondence, three of 22 (14%) show a specific or explicit alignment, and one of 22 (5%) has direct public documentation. Conditional on the nine papers that genuinely had a later public version, the corresponding rates are six of nine (67%), three of nine (33%), and one of nine (11%). Three of nine updated papers had no supported substantive implementation in the versions reviewed.
These are small-sample descriptive rates, not portfolio-wide estimates. “No observed public document change” is different from “the paper was updated without implementing the suggestions.” The latter is still not evidence that authors ignored the feedback: they may have considered and rejected it, responded privately, or revised work we have not found. Percentages are rounded; the counts are primary.
We have also strengthened paper identity checks. Normalized DOIs anchor the record where available, and reviewed title aliases can be recorded explicitly. This helps follow a paper through retitling without treating every similar title as the same work.
Early patterns by field and type of suggestion
The field comparisons are too small for strong conclusions. In the 22 reviewed endpoints, all three environment papers had a later version, compared with three of six development-economics papers, one of two technology/AI papers, one of four innovation/meta-science papers, and one of six global-health papers. These differences are partly driven by which public versions could be recovered. Conditional response comparisons are even thinner: most fields contain only one or three updated papers.
The type of evaluation point looks more informative, though still only suggestive. The clearest case was a precise proof error that authors publicly acknowledged and corrected. Two closely specified requests for additional analyses aligned with later additions, without confirmed attribution. Three robustness, confounding, or measurement concerns had weaker correspondences. Broad requests about theory, mechanisms, welfare, framing, or interpretation were not clearly implemented in three deeply reviewed cases. Concrete requests may be easier to act on and easier to verify; this does not by itself show that specificity caused the response.
Ongoing work
This is an evolving project. The next priorities are to complete and re-validate staff assessment across all 57 tracked evaluations, resolve the 21 remaining timeline questions, recover the 15 current endpoints that still fail even with headless-browser fallback, recover the one known revised endpoint that is not publicly available, and code validated suggestions into a common aspect taxonomy. A weekly no-model job now refreshes public endpoints and rebuilds the conservative match queue. Automated and model-assisted checks organize the review; source provenance, aligned document comparison, and cautious human adjudication determine what we report. The table will be updated as new evidence is reviewed.
The methodology combines partly completed staff tracking of author responses and paper revisions with automated PDF comparison. The detailed evidence table and the historical exploratory LLM screen are available in the interactive analysis.