Krishn Bera

Cognitive Science PhD Student, Brown University

Proactive-Like Context Processing Emerges in Transformers Learning Simple Cognitive Tasks


Conference paper


Krishn Bera, Michael J. Frank
9th Annual Conference on Cognitive Computational Neuroscience, CCN, 2026

View PDF
Cite

Cite

APA   Click to copy
Bera, K., & Frank, M. J. (2026). Proactive-Like Context Processing Emerges in Transformers Learning Simple Cognitive Tasks. In 9th Annual Conference on Cognitive Computational Neuroscience. CCN.


Chicago/Turabian   Click to copy
Bera, Krishn, and Michael J. Frank. “Proactive-Like Context Processing Emerges in Transformers Learning Simple Cognitive Tasks.” In 9th Annual Conference on Cognitive Computational Neuroscience. CCN, 2026.


MLA   Click to copy
Bera, Krishn, and Michael J. Frank. “Proactive-Like Context Processing Emerges in Transformers Learning Simple Cognitive Tasks.” 9th Annual Conference on Cognitive Computational Neuroscience, CCN, 2026.


BibTeX   Click to copy

@inproceedings{krishn2026a,
  title = {Proactive-Like Context Processing Emerges in Transformers Learning Simple Cognitive Tasks},
  year = {2026},
  publisher = {CCN},
  author = {Bera, Krishn and Frank, Michael J.},
  booktitle = {9th Annual Conference on Cognitive Computational Neuroscience}
}

Abstract

Proactive cognitive control requires maintaining a holistic representation of task-relevant context in preparation for what to do next. Humans develop this capacity gradually, transitioning from reactive to proactive strategies over development, but the underlying computational dynamics of how this transition occurs remains an open question. In-context learning (ICL) in meta-learning neural networks, which has been likened to working memory and cognitive control, offers a system where these internal representations are directly accessible. As a model system, we use transformers meta-trained on a version of the AX-CPT paradigm requiring ICL, whereby internal representations are fully accessible to causal intervention. Using linear decoding, we find that the model builds a holistic abstract map of the task-set before it can reliably use this map to produce correct responses. Using activation patching, we show that this map is causally active: perturbing stimulus-response associations irrelevant to the current query nonetheless disrupts the model's output, and this effect grows stronger over training. These findings suggest that holistic context encoding, a hallmark of proactive control, can emerge spontaneously from learning to perform context-dependent tasks, without explicit optimization pressure. More broadly, our work demonstrates how ICL combined with mechanistic interpretability can reveal mechanisms underlying the emergence of cognitive control strategies.