Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants
Note: body below is original English text extracted from arXiv abs / HTML. Do not treat this file as a translation.
arXiv:2609.20143 · published 2026-09-17 · submitted 17 Sep 2026
Abstract
Cognitive offloading to AI can reduce opportunities to practice skills, creating risks of deskilling. However, it remains unclear how to prevent deskilling without restricting access to AI. Here, we design two interventions to reduce offloading decisions: (1) metacognitive feedback that makes the implications of offloading for users explicit, and (2) an effort-based reward that incentivizes less extensive LLM assistance. We test both in a preregistered online experiment ($N = 704$) with a 2$\times$2 design and a no-AI control. The task was to practice fraction arithmetic with an LLM-based assistant that provided solutions only on explicit request, followed by an unaided test. Metacognitive feedback reduced answer offloading (OR $= 0.47$) and improved test performance (OR $= 1.51$). We found no evidence that the reward affected either outcome. Our results identify metacognitive feedback as a promising design choice to reduce cognitive offloading.
Authors
Sebastian Maier, Kai Schwabe, Manuel Schneider, Stefan Feuerriegel
Key claims (verbatim-leaning English extract)
- Cognitive offloading to AI can reduce opportunities to practice skills, creating risks of deskilling. However, it remains unclear how to prevent deskilling without restricting access to AI. Here, we design two interventions to reduce offloading decisions: (1) metacognitive feedback that makes the implications of offloading for users explicit, and (2) an effort-based reward that incentivizes less extensive LLM assistance. We test both in a preregistered online experiment ($N = 704$) with a 2$\times$2 design and a no-AI control. The task was to practice fraction arithmetic with an LLM-based assistant that provided solutions only on explicit request, followed by an unaided test. Metacognitive feedback reduced answer offloading (OR $= 0.47$) and improved test performance (OR $= 1.51$). We found no evidence that the reward affected either outcome. Our results identify metacognitive feedback as a promising design choice to reduce cognitive offloading.
- Preregistered N=704, 2×2 + no-AI control; fraction arithmetic practice then unaided test.
- Metacognitive feedback reduced answer offloading (OR=0.47) and improved test performance (OR=1.51).
- No evidence that effort-based reward affected either outcome.
- Keep LLM unrestricted; intervene on offloading decisions via explicit metacognitive implications rather than guardrails that withhold answers.
Structure (section headings from HTML)
(section headings from HTML)
- Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants
- Introduction
- Related Work
- 2.1. AI assistance in learning
- 2.2. Cognitive offloading
- 2.3. Metacognition
- 2.4. Research gap
- Hypotheses
- 3.1. Access to AI
- 3.2. Metacognitive feedback
- 3.3. Effort-based reward
- Method
- 4.1. Study procedure
- 4.2. Learning platform
- 4.2.1. System architecture
- 4.2.2. LLM assistant
- 4.2.3. Classifying extent of LLM support for reward computation
- 4.2.4. Task materials
- 4.3. Interventions
- 4.3.1. Metacognitive feedback
- 4.3.2. Effort-based reward
- 4.4. Measures
- 4.5. Participants
- 4.6. Ethics
- 4.7. Statistical analysis
- Results
- 5.1. Main results
- 5.1.1. Unaided performance was not significantly affected by LLM access (H1)
- 5.1.2. Metacognitive feedback reduced answer offloading and improved unaided performance (H2)
- 5.1.3. Effort-based reward did not significantly affect answer offloading and unaided performance (H3)
- 5.2. Answer offloading was associated with lower unaided performance
- 5.3. Metacognitive feedback reduced repeated answer offloading
- 5.4. Patterns of LLM use across conditions
- 5.5. Individual differences in offloading and unaided performance
- Discussion
- 6.1. Answer offloading as a source of deskilling risk
- 6.2. Interpreting the effects of metacognitive feedback and the effort-based reward
- 6.3. Theoretical implications
- 6.4. Design recommendations
- 6.5. Limitations and potential for future work
- Conclusion
- References
- Appendix A Learning Items
- Appendix B System Prompt of the LLM Assistant
- Appendix C Offloading Classifier
- C.1. Prompt templates
- C.2. Coding rubric
- Appendix D Prompt for Generating Metacognitive Feedback
- Appendix E Implementation Details for the Interventions
- E.1. Metacognitive feedback
- E.2. Effort-based reward
- Appendix F Overview of Measures
- Appendix G Conversation Transcripts
- Appendix H Robustness Checks
- Appendix I Exploratory Analyses
- I.1. Overall assistant use and the levels of AI support
- I.2. Perceived mental effort
- I.3. Metacognitive accuracy and metacognitive sensitivity
- I.4. Moderators of metacognitive feedback
Body excerpts (original English)
Cognitive offloading to AI can reduce opportunities to practice skills, creating risks of deskilling. However, it remains unclear how to prevent deskilling without restricting access to AI. Here, we design two interventions to reduce offloading decisions: (1) metacognitive feedback that makes the implications of offloading for users explicit, and (2) an effort-based reward that incentivizes less extensive LLM assistance. We test both in a preregistered online experiment ( N = 704 N=704 ) with a 2 × \times 2 design and a no-AI control. The task was to practice fraction arithmetic with an LLM-based assistant that provided solutions only on explicit request, followed by an unaided test. Metacognitive feedback reduced answer offloading (OR = 0.47 =0.47 ) and improved test performance (OR = 1.51 =1.51 ). We found no evidence that the reward affected either outcome. Our results identify metacognitive feedback as a promising design choice to reduce cognitive offloading. LLM assistants offer a new way to support learning and skill development through natural language interaction. In educational contexts, such LLM assistants can explain unfamiliar concepts, provide feedback on reasoning, and generate individualized problems to practice, for instance by scaffolding programming practice ( Kazemitabaar et al., 2024 ) or supporting mathematics learning ( Liu et al., 2026a ) . However, access to LLM support does not necessarily translate into better skills when the assistance is no longer available. There is growing concern that LLM assistance may even undermine skill development or erode existing skills, commonly discussed as deskilling ( Lenharo, 2026 ) . For example, high-school students who practiced mathematics with unrestricted ChatGPT scored higher during practice yet lower on the subsequent exam without LLM assistance ( Bastani et al., 2025 ) . Similarly, experienced endoscopists detected fewer adenomas when performing colonoscopies without AI after months of AI-assisted practice ( Budzyń et al., 2025 ) , and LLM assistance during creative tasks lowered subsequent independent creativity ( Kumar et al., 2025b ) . One possible explanation is how people use the LLM assistance, particularly whether they work through a problem with support or delegate the problem-solving to the LLM. How people use a technology can shape whether it supports or harms skill development ( Salomon et al., 1991 ) . A useful lens for understanding these differences is cognitive offloading , which involves using actions or external tools to change a task’s processing requirements and reduce cognitive demand ( Risko and Gilbert, 2016 ) . For example, a student can request the complete answer (which we call answer offloading throughout the rest of the paper), thereby removing exactly the effortful processing on which durable learning depends ( Bjork and Bjork, 2011 ; Bastani et al., 2025 ; Gajos and Mamykina, 2022 ) . The same student could instead ask how to approach the problem when stuck or request a deeper explanation, which are forms of instrumental help that support continued engagement with the task and thus may facilitate learning ( Nelson-Le Gall, 1981 ; Lehmann et al., 2024 ; Kumar et al., 2025a ; Varone et al., 2026 ) . These different ways of engaging with LLM assistance may have different consequences for learning, yet little is known about how interaction design can influence a user’s offloading decisions without restricting the assistance available. To reduce offloading, existing interventions have mainly changed what output the assistant provides. A common approach is to use guardrails that constrain the assistant toward tutoring behavior by offering explanations and guiding questions while withholding direct solutions. Such approaches can preserve later performance without LLM support ( Bastani et al., 2025 ; Bassner et al., 2026 ; Kazemitabaar et al., 2023 ) . However, restricting assistance also has drawbacks: withholding answers can demotivate lower-skilled learners and increase their dependence on the system ( Myung et al., 2026 ) , while users may circumvent guardrails when unrestricted assistance is available ( Kapoor et al., 2026 ) . An alternative is to keep the LLM assistance unrestricted and instead support users in deciding how much of the problem to solve themselves and how much to hand over to the assistant. However, design interventions that directly target these decisions are largely unexplored. Here, we target two psychological mechanisms underlying users’ decisions: (1) how they judge what their use of AI assistance means for their own learning and (2) how rewards shape the incentive to do the work themselves. The first mechanism is based on metacognition, which refers to how people assess and regulate their own thinking ( Flavell, 1979 ) . Learners may delegate work without fully recognizing what this means for their own practice, especially when a fluent AI explanation creates a sense of understanding even though they have not applied the method themselves ( Fan et al., 2025 ) . The second mechanism is motivational incentives, as rewards can change the incentive to rely on external assistance rather than perform the work oneself. Prior research shows that making external assistance less rewarding or independent performance more rewarding can reduce offloading ( Gilbert et al., 2020 ; Sachdeva and Gilbert, 2020 ) . These two mechanisms later motivate our interventions based on metacognitive feedback that makes the implications of users’ AI use visible, and an effort-based reward that encourages less extensive use of LLM assistance. We therefore ask the following research question: Research Question. Can (a) metacognitive feedback or (b) an effort-based reward reduce answer offloading and improve subsequent unaided performance? To answer this question, we design two interventions that target the offloading decision: (1) metacognitive feedback that makes the implications of the learner’s LLM use visible, and (2) an effort-based reward that awards points for less extensive assistance. The metacognitive feedback is shown between learning items and describes how extensively a learner has requested assistance so far, what this implies for skill development, and prompts learners to reflect on their understanding in upcoming items. The reward changes the incentive for requesting assistance; i.e., a correct answer earns 10 points when the learner used no AI or only general guidance, 5 points when using an intermediate result, and 1 point when requesting the complete answer. We test both interventions in a preregistered online experiment ( N = 704 N=704 ) with a 2 × \times 2 design and a no-AI control (Figure 1 ). Participants practiced fraction arithmetic with an LLM assistant. By default, the LLM assistant provided only guidance on how to approach the task and gave complete answers only on explicit request, thus leaving learners to decide how much of each problem to solve themselves or hand over to the assistant. The participants then completed a test without LLM assistance. We find that metacognitive feedback reduced answer offloading and improved test performance. We found no evidence that the reward affected either outcome. Our research makes the following contributions: Empirically, we provide evidence that metacognitive feedback on a learner’s AI requests reduces answer offloading and improves subsequent unaided performance. We found no evidence that a reward affected either outcome. Theoretically, we extend the cognitive offloading paradigm to LLM-assisted learning by focusing on users’ decisions about how much of a task to solve themselves and how much to hand over to the assistant. Practically, we derive design recommendations for reducing the risk of deskilling without restricting LLM use. Importantly, our findings suggest that designers should make the implications of AI use visible, make complete-answer requests more deliberate, and support learning-oriented help-seeking. LLM assistants promise individualized support for learning at scale, for example by adapting explanations to the learner’s needs, responding to solution attempts, and providing help when difficulties arise. A growing body of research demonstrates this potential. For example, AI tutors prompted to follow teaching best practices can outperform in-class active learning ( Kestin et al., 2025 ) , AI-generated mathematics hints can produce learning gains comparable to human-tutor hints ( Pardos and Bhandari, 2024 ) , and pedagogically designed assistants such as CodeAid already support entire programming courses ( Kazemitabaar et al., 2024 ; Liffiton et al., 2023 ) . Beyond the assistant itself, AI can also personalize the learning path, for example by tracing a learner’s evolving knowledge to recommend the next exercise ( Ozyurt et al., 2025 ) . However, there are concerns that LLM assistance during practice can also undermine skill acquisition ( Kasneci et al., 2023 ) . Bastani et al. ( Bastani et al., 2025 ) found that high-school students practicing mathematics with unrestricted LLM access performed better during practice but worse on a subsequent exam without LLM support than students in a no-AI control. Similar patterns have been observed, for example, in divergent and convergent creativity ( Kumar et al., 2025b ) , coding ( Shen and Tamkin, 2026 ) , and reading comprehension and fraction problems ( Liu et al., 2026b ) . Observational data from a large-scale learning platform points in a similar direction, namely that retention declined as students shifted work to generative AI ( Rismanchian et al., 2026 ) . Learning science provides one explanation for the above-mentioned pattern. Durable learning is known to depend on effortful processing, including working through solution steps oneself (often referred to as “desirable difficulties”) ( Bjork and Bjork, 2011 ) . When AI provides a worked solution, it removes effortful processing and leaves the learner with fewer opportunities to practice problem-solving independently ( Stadler et al., 2024 ) . The learner receives a worked solution without having constructed any part of it. Similar patterns were also observed in earlier tutoring systems; for example, learners who repeatedly requested hints until the answer was revealed tended to show lower learning gains ( Aleven et al., 2016 ; Baker et al., 2004 ) . Together, these findings suggest that assistance can hinder learning when it replaces rather than supports the reasoning and problem solving the learner needs to practice. Consequently, whether LLM assistance supports skill development may depend less on whether it is available than on how learners use it and how much of the problem solving they perform themselves. In line with this, the mere availability of LLM support shows no average effect, while solution-generating use is associated with less understanding and explanation-seeking use with more understanding ( Lehmann et al., 2024 ) . Similarly, interactions that prompt a learner’s own thinking are known to preserve cognitive engagement better than those that deliver finished solutions ( Xu et al., 2025a ) , and explanations from an interactive LLM assistant can support learning more than having access to the answers alone ( Kumar et al., 2025a ) . AI-supported practice can also leave later performance without assistance intact ( Bassner et al., 2026 ; Kazemitabaar et al., 2023 ) , whereas answer-giving AI can improve performance during practice while producing weaker subsequent learning ( Bastani et al., 2025 ) . Consistent with this, a meta-analysis of coding studies finds that generative AI raises productivity during use but has no significant effect on learning, with in-session gains not transferring to unaided assessment ( Maier et al., 2026a ) . This pattern echoes classical help-seeking research, which distinguishes instrumental requests for enough support to continue solving a problem oneself from executive requests that delegate the solution ( Nelson-Le Gall, 1981 ) . As such, we expect that LLM assistance can have effects in both ways; i.e., either by scaffolding learners’ practice or by replacing the cognitive work needed to develop the skill. Existing interventions have mainly addressed these learning risks by changing what the assistant provides. For example, the use of guardrails in LLM assistants can restrict full answers and instead encourage the LLM to provide hints or guiding questions. Such guardrails were effective in a field experiment where they prevented the performance loss observed for students with unrestricted LLM access ( Bastani et al., 2025 ) . Related designs steer assistants toward Socratic, pedagogical behavior or scaffold learners through problems step by step ( Liu et al., 2024 ; Jurenka et al., 2024 ; Ma et al., 2025 ) . Similar approaches appear in commercial learning modes, including ChatGPT Study Mode 1 1 1 https://openai.com/index/chatgpt-study-mode/ , Gemini Guided Learning 2 2 2 https://blog.google/products-and-platforms/products/education/guided-learning/ , and Claude Learning Mode 3 3 3 https://www.anthropic.com/news/introducing-claude-for-education , which emphasize guiding questions and support for learners’ reasoning. However, restricting what the assistant can provide also has drawbacks. Withholding direct answers can demotivate lower-proficiency students and increase their reliance on the system ( Myung et al., 2026 ) , while users may also circumvent guardrails to access the complete answer ( Kapoor et al., 2026 ) . However, these approaches all constrain the LLM assistant itself. Hence, whether careful interaction design can help learners decide how much help to request without restricting the LLM assistance available to them is unknown. Cognitive offloading refers to reducing cognitive effort by shifting a part of a task onto an external aid, for example by writing something down or setting a reminder instead of keeping an intention in mind ( Risko and Gilbert, 2016 ) . A specific example is the decision of how much help to request from an LLM assistant, which is the focus of our paper. HCI has long examined how cognitive work is distributed across people and external artifacts ( Hollan et al., 2000 ) ; however, LLMs change the scope of what can be offloaded. In particular, a learner can ask for a hint and continue working through the task, or request the complete solution and hand much of the cognitive work to the assistant. The amount of help a learner requests can therefore be understood as an offloading decision. Offloading is not harmful per se. It can free cognitive capacity by shifting work that is not central to the current goal onto an external aid and can even improve memory for new materials ( Risko and Gilbert, 2016 ; Storm and Stone, 2015 ) . Whether the reduced cognitive effort is helpful or harmful for skill development therefore depends on what is offloaded by the learner ( Kalyuga and Plass, 2025 ) . A recent synthesis in education makes this distinction explicit by separating beneficial offloading from offloading cognitive work that a task is intended to build ( Lodge and Loble, 2026 ) . For example, asking how to approach a problem leaves the actual practice intact, whereas requesting the complete solution removes much of the practice itself; at the extreme, users may even adopt AI output without engaging with the underlying task ( Shaw and Nave, 2026 ) . We refer to this negative form as answer offloading . Whether people offload depends on the perceived benefit relative to the effort of using an external aid ( Gilbert, 2024 ) , as well as on their confidence in their own ability ( Hu et al., 2019 ) . For learners, this can favor offloading: when the immediate goal is to complete the task, delegating the problem-solving saves effort, while doing the work oneself offers little immediate reward. At the same time, learners may not fully recognize what they give up by doing so, especially because fluent assistance can inflate their perceived learning success even when they have done less of the cognitive work themselves ( Koriat and Bjork, 2005 ) , while reliance on the assistant can reduce the self-regulation involved in learning ( Fan et al., 2025 ) (see the discussion on metacognition in the next section). Together, these mechanisms point to two ways of reducing detrimental offloading: (1) changing the incentives for doing the work oneself and (2) helping learners recognize the potentially negative implications of AI use. In human–AI interaction, reliance describes the extent to which users adopt AI output in their own judgment ( Schemmer et al., 2023 ; Raees et al., 2026 ) . Much of this literature asks how users should rely on AI when its advice may be wrong. In learning, however, even correct assistance can create a different challenge: users may rely on the assistant in ways that reduce their own cognitive engagement. The HCI community has begun to frame generative AI through the lens of cognitive offloading ( Tankelevitch et al., 2025 ) , even though work on offloading to AI assistants is largely theoretical ( León-Domínguez, 2024 ; Grinschgl and Neubauer, 2022 ) or correlational ( Gerlich, 2025 ; Lee et al., 2025a ) . Controlled studies are largely the exception. For example, grading the extent of AI assistance can preserve more cognitive engagement than full automation ( Chen et al., 2025 ) , while, in a creativity task, an LLM that asks questions to elicit a user’s own ideation process rather than automating the idea generation can preserve idea quality and perceived ownership at the price of higher perceived effort ( Maier et al., 2026b ) . However, both exceptions redesign how the assistant responds; in contrast, whether users can instead be supported in regulating how much cognitive work they hand over remains unknown. Metacognition refers to thinking about one’s own thinking, which is classically divided into monitoring one’s cognitive states and controlling one’s cognitive activities ( Flavell, 1979 ; Nelson and Narens, 1990 ) . Deciding how much work to hand over to an external aid is one such control decision, and it depends on how learners assess their own understanding and need for assistance. Using LLMs already requires metacognitive monitoring and control. Because users largely steer the interaction themselves, they must decide what to ask, monitor whether the responses serve their goals, and adjust their interaction accordingly ( Tankelevitch et al., 2024 ) . These strategies can thus influence performance with AI. In a field experiment with consultants, LLM access improved creative performance particularly for consultants who used such metacognitive strategies ( Sun et al., 2025 ) . Learning adds another challenge for metacognitive demands, as learners must judge not only whether the assistance helps them complete a task, but also whether they learn to perform the task independently as a result. Importantly, learners tend to overestimate their understanding when they practice a task with support but where such support is unavailable during the test ( Koriat and Bjork, 2005 ) . One reason is that such LLM support may make learners feel more fluent than they actually are. For instance, students who practiced with unrestricted LLM access did not recognize the later performance decline during an unaided exam ( Bastani et al., 2025 ) ; this finding is consistent with broader concerns that AI assistance may reduce self-regulation during learning ( Fan et al., 2025 ) . A learner working with an LLM assistant needs to monitor their own understanding and regulate how they use the assistant in response. Metacognitive feedback offers a principled way to help users in both monitoring their understanding and regulating their use of AI assistance. In AI-assisted decision making, tailored interventions (e.g., cognitive forcing functions, prompts to encourage a user’s own reasoning, and explanations that facilitate verification) can help users evaluate AI advice more critically ( Buçinca et al., 2021 ; Danry et al., 2023 ; Vasconcelos et al., 2023 ) . In computer-based learning environments, metacognitive prompts help learners monitor their understanding ( Bannert et al., 2015 ) , while metacognitive feedback can improve learning outcomes, with meta-analyses reporting moderate effects ( Zheng, 2016 ; Guo, 2022 ) . More recently, work has extended metacognitive feedback to generative AI interactions. For example, support agents can encourage users to reflect on their goals and reasoning while working with generative AI ( Gmeiner et al., 2025 ) . In educational contexts, interventions teach prompting and help-seeking strategies through separate training activities ( Xiao et al., 2025 ; Valle Torre et al., 2026 ) , or provide general metacognitive feedback while students work with the tool ( Xu et al., 2025b ; Singh et al., 2025 ) . To the best of our knowledge, no prior work has provided learners with feedback based on their own AI use during learning and whether such feedback can improve self-regulation is thus unclear. In summary, prior research (Table 1 ) shows that LLM assistance can reduce cognitive engagement and thus hinder skill development. However, no study has proposed and evaluated interaction design interventions that support users in regulating their own AI use to improve subsequent unaided performance. Hence, reducing answer offloading through interaction design remains an open challenge. We develop and experimentally test different interventions to reduce answer offloading and improve skill acquisition. We derive three preregistered hypotheses with two behavioral outcomes, namely, answer offloading during learning and unaided performance during subsequent tests (Section 4.4 ). The hypotheses and analysis plan were preregistered on AsPredicted ( https://aspredicted.org/gx2fa8.pdf ). Access to AI AI support can oftentimes improve immediate task performance ( Gajos and Mamykina, 2022 , e.g.,) , but it may also reduce critical thinking and the processing depth that people invest in solving tasks ( Stadler et al., 2024 ; Lee et al., 2025a ; Chen et al., 2025 ) . Given that learning requires such effortful processing ( Bjork and Bjork, 2011 ; Soderstrom and Bjork, 2015 ) , better performance during practice with LLM assistance does not necessarily translate into actual skill acquisition. Consistent with this, studies across various domains have found lower unaided test performance after practice with unrestricted LLM assistance compared to an unassisted control ( Bastani et al., 2025 ; Liu et al., 2026b ; Shen and Tamkin, 2026 ) . We thus expect: Participants with LLM assistance during learning perform worse on the subsequent unaided test than participants who learn without LLM assistance. Metacognitive feedback Offloading decisions depend in part on how learners assess their own understanding and need for assistance. LLM support can make this assessment more difficult, as fluent answers may increase perceived understanding while reducing self-regulatory engagement ( Koriat and Bjork, 2005 ; Fan et al., 2025 ) . Interventions based on metacognition have been found to improve learning behavior and student performance ( Zheng, 2016 ; Guo, 2022 ) . Recent work further suggests a similar benefit for self-regulated learning with AI ( Xu et al., 2025b ; Lee et al., 2025b ) . For example, in the context of intention offloading, metacognitive feedback (e.g., via strategic advice ( Gilbert et al., 2020 ) and brief training with feedback ( Ngai and Gilbert, 2026 ) ) can help people regulate when to rely on external aids. Transferring these results to our context of LLM-assisted learning, we expect: Participants who receive metacognitive feedback are less likely to offload the answer on a learning item than participants without metacognitive feedback. Participants who receive metacognitive feedback perform better on the subsequent unaided performance test than participants without metacognitive feedback. Effort-based reward Cognitive effort depends partly on the incentives attached to doing the work oneself versus relying on an external aid ( Gilbert, 2024 ; Kool and Botvinick, 2018 ) . Hence, reducing the incentives for cognitive offloading should shift learners toward solving more of the task themselves. For example, in reminder-setting paradigms, people offload less when they earn fewer points for using an external reminder than recalling the item unaided ( Gilbert et al., 2020 ) . Extending this idea to LLM assistance, a reward that awards fewer points after more extensive assistance should make doing more of the work oneself more attractive and should reduce answer offloading. We therefore expect rewards to reduce answer offloading and thus improve learning. Consequently, we hypothesize: Participants who receive an effort-based reward are less likely to offload the answer on a learning item than participants without an effort-based reward. Participants who receive an effort-based reward perform better on the subsequent unaided performance test than participants without an effort-based reward. We conducted a preregistered online experiment in August 2026 in which participants practiced fraction arithmetic with an LLM assistant and afterwards completed a test without LLM assistance (which we refer to as “unaided performance”). We preregistered the hypotheses, design, sample size, exclusion rules, and analysis plan on AsPredicted before data collection. Analysis code, anonymized data, and LLM reporting following best-practice guidelines ( Feuerriegel et al., 2026 ) are available in our anonymized repository. 4 4 4 https://anonymous.4open.science/r/cognitive-offloading-CHI2027 Our experiment comprised six steps. (1) Participants were redirected from Prolific to our custom study platform and were randomly assigned to one of the five conditions. (2) A baseline block collected demographics and other controls such as AI use frequency (Section 4.4 ). (3) Participants received a brief review of the basic rules needed to solve fraction arithmetic problems, and participants in conditions with assistant access were additionally told that the LLM assistant would be available during the learning phase, but that the subsequent test measures unaided performance. (4) Participants then completed 10 learning items, with the correct solution shown after each item. The first exposure to the intervention is after the practice item; accordingly, we measure answer offloading not on the first item but only on the subsequent nine items. Depending on condition, participants completed the learning phase without LLM assistance (no-AI control), with LLM assistance but no intervention (AI-only), with metacognitive feedback, with reward, or with both interventions (Section 4.3 ). (5) After the learning block, participants rated their perceived mental effort and predicted how many of the six test items they would solve correctly. (6) Participants completed the six test items without LLM assistance (i.e., unaided performance) and rated their confidence after each one as well as their perceived mental effort during the test. Finally, participants completed an honesty check and were briefly debriefed about the study background. We built a custom web platform so that every part of the participant experience (item presentation, chat, scoring, intervention delivery) was under experimental control and fully logged. For this we used a React-based frontend, which communicates with a Python backend that manages participants’ sessions, with data stored in a PostgreSQL database. We built the LLM assistant around GPT-OSS-120B ( OpenAI, 2025 ) , which is an open-weight, frontier LLM and which we accessed through Amazon Bedrock ( openai.gpt-oss-120b-1:0 ). We set the temperature to 0 for more reproducible behavior. During the learning phase, the assistant appeared in a chat panel shown next to the current item. Each item has its own conversation. Finally, all chat transcripts, answers, timestamps, and intervention exposures are stored server-side. Figure 2 shows the architecture.