Evidence Ledger · Supports existing guidance · reviewed 2026-09-05

Evidence of impact and interpretational limits of generative AI in STEM education: a systematic review and meta-analysis on cognitive learning outcomes

The reviewed source is consistent with an existing Brali boundary.

Brali decision: Supports existing guidance. This page summarizes a reviewed evidence boundary; it does not reproduce the source and it does not turn one paper into a universal prescription.

What the source supports

Brali should not treat generative AI as a learning intervention by itself. In this highly heterogeneous literature, the apparent positive pooled effect did not survive robust publication-bias correction, while more informative moderator evidence suggested that outcomes depend partly on what cognitive work the learner still performs. This supports the existing human-first AI collaboration principle in learning contexts: preserve a meaningful learner attempt, explanation, retrieval, reasoning or verification step and use AI to extend or refine that activity rather than silently replacing the activity being learned.

What it does not establish

Important limitations

Source context

Evidence of impact and interpretational limits of generative AI in STEM education: a systematic review and meta-analysis on cognitive learning outcomes

  • Source type: systematic-review-and-meta-analysis
  • Design: Preregistered systematic review and meta-analysis of peer-reviewed quantitative STEM-education studies involving learner interaction with generative AI and a comparison or control group. Searches of ERIC, PsycINFO and Web of Science were updated May 7, 2026 and supplemented by citation tracking. Eighty-five studies met review criteria; 49 studies contributed 59 effect sizes to the quantitative meta-analysis. Random-effects meta-analysis, risk-of-bias assessment, moderator analyses, coincidence analysis and Robust Bayesian Meta-Analysis were used to examine heterogeneity and publication bias.
  • Population: Learners in STEM education. Most reviewed work used text-based generative-AI systems and much of the literature was in higher education; among the 49 meta-analyzed articles, the detailed review classified 25 as university studies, 23 as K-12 and one with educational level unavailable.
  • Exposure / intervention: Learners used generative-AI systems during STEM learning activities. The review classified studies partly by whether AI substituted, augmented or redefined learners' cognitive activity relative to the comparison condition and by whether outcomes primarily assessed knowledge or skills.
  • Outcomes: Externally assessed cognitive learning outcomes; Knowledge versus skill outcomes; Heterogeneity across studies; Publication-bias-adjusted overall effect; Moderation by augmentation/substitution of cognitive activity; Reported learner challenges and instructional supports

Citation: Boolzen C, Kuhn J, Flegr S, et al. Artificial Intelligence Review. Published August 25, 2026.

DOI: 10.1007/s10462-026-11665-9

What this changes in Brali

Historical review target

human first ai collaboration

This identifier was used as an evidence-review target but was never published as a canonical Brali protocol. It is retained as provenance and deliberately does not link to a public route.

Editorial note: Evidence role: strengthen and teach the existing boundary rather than create another AI hack. For a learning task, ask what cognitive operation the user is trying to acquire. Require a meaningful unaided attempt or explicit reasoning step where that operation matters; then use AI for examples, critique, comparison, feedback or extension; finally verify or reproduce the target skill without leaning on the generated answer. The sarcastic version is accurate enough for editorial use: if the AI performs the exact thinking you are trying to learn, excellent news for the AI. The learner still needs a turn.

Use the boundary, not just the headline

When an AI agent or a person retrieves this decision, preserve both the supported claim and the unsupported or overstated claims. Dropping the boundary would turn a reviewed source into a stronger claim than Brali actually maintains.

All Evidence Decisions → · Research & trends → · Machine-readable review →