Interpret with care. These are model-as-judge semantic-overlap judgments, not ground truth and not a verdict on whether unmatched LLM concerns are invalid. The underlying human critique text and evaluator attributions are not republished here; examples are paraphrased.
14papers judged
74.9%mean human-issue coverage
46.9%mean LLM overlap rate
56matched issue groups
1. What was judged
For each paper, GPT-5.2 Pro received an enumerated set of human concerns and an enumerated set of model-generated concerns. It mapped related issues, scored their semantic overlap from 0–100, explained similarities and differences, and listed concerns that had no counterpart.
Coverage is the proportion of human concerns with at least one match at the original threshold of 30%. “LLM overlap rate” is the proportion of model concerns with such a match. It is deliberately not called precision: a human–LLM mismatch does not establish that a model concern is wrong.
2. Side-by-side examples
These representative examples are paraphrased from the historical judgment record.
Misperceptions and Demand for Democracy under Authoritarianism95% judged overlap
Human concern, paraphrased
Potential spillovers of campaign messages could contaminate comparison neighborhoods and attenuate estimated effects.
GPT-5.2 Pro concern, paraphrased
Plausible neighborhood interference violates the no-spillover assumption unless the study tests or bounds it.
Judge finding: Essentially the same identification concern, with the model adding possible diagnostics.
Does online fundraising increase charitable giving?75% judged overlap
Human concern, paraphrased
Handling of unusually large donations may be a consequential analytical choice, especially for profitability estimates.
GPT-5.2 Pro concern, paraphrased
Capping high donation values can materially affect inference in a heavy-tailed outcome and needs robustness checks.
Judge finding: Strong overlap: both focus on sensitivity of headline results to tail handling.
Observed regeneration patterns cannot straightforwardly identify a purely biophysical potential when land use and governance also matter.
GPT-5.2 Pro concern, paraphrased
A biophysical prediction may absorb socioeconomic and land-use effects, undermining a pure-potential interpretation.
Judge finding: Strong conceptual agreement, framed in different disciplinary language.
3. Browse all paper-level judgments
Select a paper to view the judge’s mapped issue groups, their overlap scores, explanations, and the unmatched-concern summaries.
4. Limits and next step
The run used GPT-5.2 Pro as the judge, so it is a useful diagnostic rather than independent validation. The next planned round should retain the stable issue IDs, have a stronger model make pair-level judgments, and escalate ambiguous cases to a second high-effort adjudicator. Threshold-based metrics should then be calculated in code, with sensitivity checks.