Correlation Is Not Causation, But Our Review Process Pretends Otherwise

In Aug 2026, Science published a study on China's solar policy and bird diversity. Its "significant" causal claims drew methodological criticism. As big data and ML lower barriers, are standards eroding? This article examines editors' duties in causal inference.

In August 2026, Science published a study asserting that China's solar expansion policy reduces bird diversity (doi:10.1126/science.aee0747). The authors reported that a one-standard-deviation increase in policy stringency lowers the Shannon diversity index by 2.10% (beta = -0.0125, SE = 0.0037, P < 0.01). Yet the paper's causal claims have since drawn intense methodological scrutiny.

Critics have identified endogeneity bias arising from non-random policy assignment, aggregation fallacies that mask heterogeneous treatment effects, and questionable exclusion restrictions in the instrumental-variable strategy. These are not isolated technical flaws; they signal a systemic erosion of methodological standards in policy-evaluation research.

The root cause is structural. As big data and machine-learning tools lower the barrier to empirical analysis, as sample sizes inflate to render minuscule effects statistically significant, and as policy-evaluation manuscripts flood premier journals, the gatekeeping function of editorial review risks being outpaced by the volume of submissions. The question is no longer whether flawed causal inference occasionally slips through, but whether editorial standards are systematically failing to catch it.

Here, we argue that social science editors must evolve from passive gatekeepers into active guardians of causal rigor. We distill five imperatives-rooted in the Science case but generalizable across economics, public policy, and development studies-that should be embedded in editorial guidelines and reviewer scorecards.

1. Insist on explicit identification strategies

The study's policy stringency index is not exogenous. High-PSI regions correlate with faster urbanization, greater energy demand, and heavier ecological stress-confounders that simultaneously drive both solar policy and biodiversity loss. The instrumental variable-historical sunshine hours interacted with climate-policy uncertainty-fails the exclusion restriction, because sunshine patterns plausibly influence vegetation productivity and hence habitat quality through channels other than solar policy.

In economics, core explanatory variables are rarely exogenous. Digital financial inclusion, for instance, is concentrated in regions with better business climates and higher human capital-omitted variables that jointly drive both fintech adoption and entrepreneurship. Editors must therefore require a standalone "identification paragraph" in which authors state their causal assumptions, justify exclusion restrictions, and discuss the direction of bias if assumptions fail.

2. Mandate heterogeneity analysis

The Science paper pools 2,344 counties into a single regression, reporting an "average" 2.10% decline. But the authors themselves acknowledge that effects are insignificant in desert regions and negative only in non-desert areas. The national average is a statistical artifact that describes no real place; disaggregated by vegetation type, altitude, or income, some subpopulations might show positive effects.

Editors must reject manuscripts that report only aggregate coefficients. Authors should be required to present interaction effects, subsample analyses, or causal forest estimates, and to specify explicitly for whom the treatment works, for whom it does not, and for whom it backfires.

3. Scrutinize data provenance

The study relies on citizen-science birdwatching records. Such data suffer from structural selection bias: observations cluster in wealthy, accessible regions, while poor, remote, and desert areas are underrepresented. Temporal trends in observer skill and reporting effort further contaminate the panel. In economics, e-commerce data skew toward young urban elites; administrative records reflect policy-responsive reporting behavior rather than objective truth.

Editors must require a data-provenance statement: How were the data generated? How does the sample compare with the target population? What is the attrition pattern? Triangulation across multiple sources and reweighting for coverage bias should be encouraged, not treated as optional robustness checks.

4. Separate statistical from practical significance

With 46,528 observations, even a 0.1% decline can achieve P < 0.01. The paper's 2.10% effect is self-described as "moderate," yet the abstract omits confidence intervals, leaving readers unable to assess whether the lower bound is economically meaningful. In policy research, a "significant" 0.05-percentage-point profit gain may be swamped by implementation costs; a "significant" test-score boost may fade out within three months.

Editors should enforce a strict rule: no P-values without confidence intervals, standardized coefficients, and cost-benefit context. Effect sizes must be benchmarked against competing explanations-business cycles, technological shocks, institutional change-so that readers can judge whether the finding matters.

5. Preserve temporal dynamics

The study collapses a decade of data into static estimates. Yet ecological responses lag; immediate effects may capture concurrent shocks rather than policy impacts. If negative effects were concentrated in 2014-2018 but attenuated after 2019 as ecological compensation improved, the ten-year average tells a misleading story.

Policy effects are rarely time-invariant. Editors must require event-study plots, dynamic treatment-effect estimates, and tests for structural breaks. Short-term costs and long-term gains-automation responses to minimum wages, for example-must not be averaged into a single coefficient that is useless for policy design.

6. A seven-dimensional framework for editorial action

These five imperatives can be operationalized through a seven-dimensional review framework that maps each editorial dimension onto a verifiable evidence requirement. The framework's feasibility rests on two pillars.

First, each dimension pairs a precise audit question with concrete deliverables. For identification, authors must supply an exclusion-restriction argument and overidentification tests. For heterogeneity, they must report subsample regressions or causal forests. For data quality, they must provide a provenance statement and coverage-bias diagnostics. For effect sizes, they must report confidence intervals and benchmarked coefficients. For temporal dynamics, they must supply event-study figures and persistence tests. For indicator validity, they must justify the theoretical link between measured variable and latent construct. For robustness, they must present Oster-style sensitivity analyses and placebo tests.

Second, the required tools are already standardized. Instrumental-variable overidentification tests, difference-in-differences parallel-trend diagnostics, causal forests, and Callaway-Sant'Anna dynamic treatment-effect estimators are available as off-the-shelf packages in Stata, R, and Python. Editors need not become econometricians; they need only insist that authors present the corresponding evidence, transforming the abstract ideal of "guarding credible inference" into an enforceable gatekeeping mechanism.

Conclusion

The Science controversy is not merely a flawed paper; it is a symptom of a deeper structural crisis. Under publish-or-perish pressure, as big data and machine learning lower empirical barriers, causal-inference quality control is retreating from the methodological frontier into an editorial blind spot. Authors are incentivized to hunt for significance; readers are primed to believe it; and editors-if they remain passive-risk letting correlation masquerade as causation.

We do not urge editors to master econometrics. We urge them to cultivate methodological vigilance: to ask, in every review, what is the identification strategy? Whom does the sample represent? What is the effect size in practical terms? Where does the result break down? How does it evolve over time? These questions demand no advanced mathematics-only respect for causal logic and fidelity to the editorial mandate of safeguarding scientific integrity.