In this article
Today we talk about the study that has gotten the most media coverage in 2025 on the cognitive effects of using ChatGPT, and we're going to tell the whole thing. Including what the study does say, what it doesn't say, the honest limitations and the reasonable consequences. It's called Your Brain on ChatGPT, it's signed by a team at the MIT Media Lab led by Nataliya Kosmyna, and it was published in June 2025. The media coverage settled on the most simplified headline —«ChatGPT makes you dumb»— which the study explicitly doesn't support. My take is biased, because I've read the full paper, not just the headlines. Form your own with the detail.
I've had the paper on my desk since June 2025. It's the first rigorous academic study —with an adequate sample, careful methodology, neurophysiological measurement— on the effects of using AI assistants in writing. The general-press coverage simplified it in the worst possible way. Let's dismantle the simplification and tell what it really says.
What exactly the team did
The study was published on arXiv on 10 June 2025 under the identifier 2506.08872, with the full title Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. The main authors are Nataliya Kosmyna and collaborators from the MIT Media Lab. At the time of writing this article, the paper is in the peer-review process for an international academic conference; it's available as a preprint on arXiv and has been the object of analysis by specialists in cognitive neuroscience, language sciences and educational psychology.
The methodology consisted of the following. Fifty-four participants, divided randomly into three groups of eighteen people each. The first group, called Brain-only, wrote essays with no external help. The second group, called Search Engine, wrote with access to Google as a consultation tool. The third group, called LLM, wrote with access to ChatGPT as an assistant.
Each participant completed four writing sessions spaced over time, writing each time an opinion essay on a designated topic. During the writing they wore an electroencephalography cap that measured brain activity in different regions, with a focus on the areas associated with language processing, linguistic planning and working memory.
After the sessions, the produced essays were evaluated by judges blind to the experimental condition, according to criteria of quality, originality and argumentative coherence. Participants were also given later recall tests: what did they remember of the essay they themselves had produced?
In the last session, some participants from the LLM group were asked to write without assistance, to assess the residual effect of the earlier sessions. Some participants from the Brain-only group were allowed to use ChatGPT, also to assess the inverse dynamic.
The results, in brief
The study's main findings are the following. They're best read in their nuanced form, not in the journalistic simplification.
First, neural connectivity measured with EEG. The LLM group showed lower measured connectivity in some brain networks associated with linguistic planning and episodic memory during the task. The difference relative to the Brain-only group was statistically significant. The difference relative to the Search Engine group was also significant but smaller; the Search Engine group occupied an intermediate position.
Second, essay quality according to blind judges. The LLM group produced essays with technical quality comparable to or higher than the Brain-only group in terms of expository clarity, syntax and grammatical correctness. But the LLM group's essays showed less originality of approach and greater structural homogeneity across different participants. The LLM group's participants tended to produce essays that resembled each other more.
Third, later recall of one's own content. The LLM group's participants had greater difficulty specifically recalling which arguments they had included in the essay they themselves had handed in weeks earlier, compared to the Brain-only group's participants. The processing during production was shallower, which translated into worse consolidation in long-term memory.
Fourth, residual effect. When the LLM group's participants were asked to write without assistance in the fourth session, they showed a transient deficit relative to the Brain-only group in fluent production. That deficit was smaller than the one observed in earlier sessions and showed a tendency to recover with practice.
The results that weren't cited as much
The media coverage centered on the first and third findings and simplified the last two. It's worth rescuing what was lost in the simplification.
The first nuance is that the Search Engine group also showed significant differences relative to the Brain-only group. That is, intensive use of Google already produced a certain measurable cognitive delegation relative to writing with no help at all. The AI assistant amplified the effect, but the base effect already existed. This means the problem isn't exclusively ChatGPT; it's the wider category of cognitive offloading to external tools. And that, in turn, opens a more nuanced debate: any consultation tool produces some degree of externalization; the question is one of degree and proportion.
The second nuance is that the deficit observed in the LLM group was reversible with practice. When participants wrote without assistance, their performance improved session by session. The cognitive degradation wasn't a permanent trait; it was a reversible functional state. The study suggests the problem isn't brain damage —the brain doesn't break. The problem is operational skill that weakens with lack of practice and recovers with deliberate practice.
The third nuance is that the technical quality of the LLM group's essays was good. If the criterion was «produce a correct and clear text», the assistant users achieved it. The problem was in the dimension of originality and deep processing, not in the dimension of formal correctness. This has implications for how professional performance is measured in contexts where assistance has spread.
The honest limitations
The study has limitations worth recognizing, because any solid conclusion must start from them.
The sample is small: 54 participants in total, 18 per group. For robust statistical inference about general populations, that sample is modest. The findings are significant in the technical sense of p-value, but replication with larger samples is needed before generalizing.
The time is short. Four sessions spaced over a few weeks don't allow assessing long-term effects. The hypothesis of cognitive debt —cognitive debt accumulated through continued use over years— can only be tested with longitudinal studies that haven't yet been done.
The context is artificial. Writing essays in a lab with an EEG cap isn't exactly the same as writing an email or a report under real working conditions. The transfer of the findings to everyday contexts is plausible but not automatic.
The EEG metrics are still relatively coarse for fine cognitive conclusions. Electroencephalography measures aggregate electrical activity in superficial cortical areas; it doesn't allow the detail that techniques like functional magnetic resonance would offer. The inferences about «less processing» are based on global patterns, not on specific measurements of particular cognitive circuits.
And the task bias is important. Opinion essays are a specific form of writing; the results might not extrapolate cleanly to other forms —technical, creative, journalistic, legal.
Taken together, these limitations make it impossible to conclude from the study things like «ChatGPT makes you dumb», «the brain atrophies with AI» or «the damage is permanent». None of those claims is supported by the study, and the authors themselves are explicit in their limitations sections.
What the study does prove
Once the limitations are marked off, the core of what the study does establish remains, and it isn't little.
There's a measurable difference in brain activity between the person who articulates text without help and the person who produces text with AI assistance. That difference isn't trivial. It's exactly the difference qualitative observation had been describing since 2023, and for the first time it's measured with neurophysiological instruments.
That difference shows up as less deep processing during production, as greater homogenization across different producers, as worse later recall of one's own content, and as a transient deficit when production without assistance is requested after a period of continued use.
That difference is reversible with deliberate practice, which means the skill isn't lost as an ultimate capacity; it weakens as operational fluency.
And the difference is one of degree, not of nature. Use of a classic search engine already produces some cognitive offloading; the AI assistant amplifies the effect but doesn't introduce a completely new phenomenon.
For a person who designs their use of assistants with knowledge of the phenomenon, these findings are usable. Not for never using assistants —the aggregate cost is low if used with a head on your shoulders. For maintaining deliberate practice of writing without assistance with some regularity, especially in the contexts where one's own articulation matters.
The political question
This is personal opinion, but the study's methodology backs it. The appearance of rigorous empirical evidence on the cognitive effects of AI assistants changes something in the public conversation. What was once qualitative observation easily dismissed as nostalgia or technophobia is now a measurable neurophysiological finding.
That should translate into at least three reframings for the coming years.
In research, there must keep being replication with larger samples, with more diverse tasks, with longitudinal follow-up. The Kosmyna paper opens the line; the line needs dozens more studies to consolidate. Public and private funders of neurocognitive research should prioritize this area.
In education, there must be institutional discussion about when and how the use of assistants is allowed, not as moralizing prohibition nor as passive acceptance, but as an informed pedagogical decision. Exams, academic work, formative writing have specific functions that can be compromised if delegated without nuance.
In the public conversation, there must be active resistance to the two opposite simplifications: «ChatGPT makes you dumb» and «nothing's happening, all the same». Both are false. The reality is nuanced: there's a measurable cognitive effect, it's reversible with deliberate practice, it depends on how it's used. That reality doesn't fit in a headline, but it's the only honest one.
The hard fact to close on. According to the paper Your Brain on ChatGPT by Kosmyna et al. (arXiv:2506.08872, MIT Media Lab, June 2025), the LLM group's participants showed a mean reduction of up to 38% in the measured connectivity in brain networks associated with linguistic planning during the writing task, compared to the Brain-only group. That concrete figure is the quantitative translation of the phenomenon. It isn't the subjective perception of an anxious user; it's a neurophysiological effect measurable with available instruments. And, according to the study's own data, partially reversible with deliberate practice without assistance over a span of weeks.
Definitions
Electroencephalography (EEG): a non-invasive technique for measuring brain electrical activity through electrodes placed on the scalp. It allows recording activity patterns during tasks in real time, though with limited spatial resolution.
Functional connectivity: a measure of the degree to which different brain regions activate in a coordinated way during a task. Reduced connectivity suggests less integration between cognitive processes.
Cognitive debt: a concept proposed by Kosmyna et al. to describe the deficit accumulated in cognitive skills through routine use of AI assistants, visible when the assistance stops being available.
Deep processing: a form of cognitive processing involving extensive semantic analysis, connection with prior knowledge and personal elaboration. It's associated with better consolidation in long-term memory.
References
Nataliya Kosmyna et al., Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task (arXiv:2506.08872, MIT Media Lab, June 2025). The original paper of the study.
Lee et al., AI Writing Tools and Cognitive Offloading (Proceedings of CHI Conference, 2024). A complementary study on cognitive dependence in assisted writing.
Evan F. Risko & Sam J. Gilbert, Cognitive Offloading (Trends in Cognitive Sciences 20:676-688, September 2016). A theoretical review of the offloading of cognitive tasks to external tools.
Betsy Sparrow, Jenny Liu, Daniel M. Wegner, Google Effects on Memory: Cognitive Consequences of Having Information at Our Fingertips (Science 333:776-778, August 2011). A foundational antecedent on the change in memory due to the external availability of information.
Maryanne Wolf, Reader, Come Home: The Reading Brain in a Digital World (Harper, 2018). A neurocognitive framework on the changes in the reading and writing brain.
Lev Vygotsky, Thought and Language (MIT Press, 1962). A classic framework on the relation between linguistic articulation and the formation of thought.
MIT Media Lab, institutional site with publications and complementary research (media.mit.edu). The institutional context of the study.
To go deeper
Andy Clark, Supersizing the Mind: Embodiment, Action, and Cognitive Extension (Oxford University Press, 2008). A classic philosophical framework on distributed cognition that frames the current debate.
Susan Greenfield, Mind Change: How Digital Technologies Are Leaving Their Mark on Our Brains (Random House, 2014). A popular neuroscience framework on the brain effects of digital technologies.
Daniel Kahneman, Thinking, Fast and Slow (Farrar, Straus and Giroux, 2011). A framework on the two modes of thinking that frame the distinction between deep and shallow processing.
Stanislas Dehaene, Reading in the Brain: The New Science of How We Read (Viking, 2009). A neuroscientific framework on how the human brain was transformed by the sustained practice of reading, transferable to the practice of writing.
You might also like
- The price of letting AI write for you
- Augmented cognitive laziness
- Why rereading no longer works since you started using AI

Comments0
No comments yet.
Leave a comment