1. Introduction
There has been an increasing emphasis on fostering students’ Self-Regulated Learning (SRL) skills to improve academic success and lifelong learning (Panadero, 2017; Zimmerman, 2002). However, research has shown that students struggle to develop effective SRL-based learning strategies and they usually encounter challenges planning, monitoring, and reflecting on their learning processes without guidance; therefore, educators have attempted to develop ways to support learners’ SRL (Heikkinen et al., 2023). In tradition, SRL is grounded in constructivist learning theory, which views learners as active agents who construct knowledge through experience and reflection. Effective learning means that learners develop strong awareness of how they learn, what they learn, and how they learn (Allen, 2022; Lin & Chang, 2023).
Advances in educational technology and Learning Analytics (LA) allow instructors to gain insights into students’ engagement and SRL by capturing their digital learning traces (Alhazbi et al., 2024; Ifenthaler & Yau, 2020). Generating data or deploying an Artificial Intelligence (AI) tutor is not a panacea. Studies have shown that LA interventions have achieved mixed outcomes (Heikkinen et al., 2023). Few interventions addressed all phases of the SRL cycle, which is recommended. Effectiveness is defined by improved learning outcomes, and by the improvement of the self-regulation process itself, including increased engagement with materials and students’ long-term mastery of SRL skills.
Therefore, the goal of this study is to evaluate a multi-agent AI assessment tool that supports students’ SRL while leveraging LA to provide instructors with insights into learning progress. Students take a Multiple Choice Question (MCQ) assessment where they can use limited hints to seek help from the Teaching Assistant AI (TA-AI) that can provide hints, but not answers. The Analytics-AI gathers all student performance and their interactions with TA-AI and provides feedback to the instructors on topics on which students struggle the most. Lastly, AI-Improver provides the software developer student feedback on the tool to improve the prompt engineering. The approach is guided by a design-based research (DBR) methodology, which facilitates iterative refinement of the intervention through cycles of design, implementation, and evaluation in classroom settings (Tinoca et al., 2022). By situating this work in the context of constructivist learning principles, SRL, and LA, we seek to advance knowledge on effectively supporting SRL in Computer Science Education (CSE).
1.1. Research Questions (RQ)
Based on the objectives of this study, the following research questions were proposed:
-
RQ1. How is students’ refinement of TA-AI prompts associated with adaptive help-seeking during quizzes?
-
RQ2. How do students use the TA-AI hints across questions of different difficulty?
-
RQ3. Is the use of the TA-AI hints associated with quiz performance?
RQ1 aims to inform the theoretical understanding of adaptive help-seeking and prompt refinement behaviors. RQ2 and RQ3 evaluate the TA-AI system’s effectiveness from a functional perspective by examining how prompts are used across questions of varying difficulty and whether prompt usage is associated with quiz correctness. Together, these questions allow both behavioral and system-level examination of the proposed TA-AI tool.
Among the various SRL behaviors, this study focuses specifically on help-seeking because it represents the most directly observable and measurable SRL behavior within the AI-assisted quiz context. In the SRL cycle, help-seeking operates at the intersection of the performance and self-reflection phases: students must first monitor their understanding (self-reflection), recognize a knowledge gap, and then strategically decide to seek external assistance (performance) (Newman, 2012; Zimmerman, 2002). Unlike other SRL behaviors such as goal-setting or time management, help-seeking in this system generates concrete, traceable artifacts in the form of student-generated prompts, enabling systematic analysis of both the frequency and quality of self-regulatory engagement.
2. Literature Review
The research is rooted in SRL theory, emphasizing the importance of goal setting, self-monitoring, and reflective practice for optimal learning outcomes.
SRL: SRL refers to the active management of learners’ learning processes by setting goals, applying strategies, monitoring progress, and adjusting tactics to achieve desired results (Zimmerman, 2002). Multiple theoretical models of SRL share the idea that SRL involves a cyclical process that generally encompasses three phases: a preparatory or planned phase, a performance phase, and a self-reflection phase (Panadero, 2017). Research consistently shows that students with strong SRL skills tend to achieve higher academic performance and exhibit greater persistence in learning, both in traditional classrooms and online environments (Broadbent & Fuller-Tyszkiewicz, 2018; Radović et al., 2024). Importantly, SRL is not viewed as a fixed trait or innate ability, but as a set of learnable skills and habits that can be cultivated through targeted instruction and sustained practice (Panadero, 2017). This view aligns with constructivist principles, which posit that learners when supported under appropriate conditions, can become active agents in directing and managing their own learning (Ballantyne et al., 2025; Lin et al., 2026; Lin & Chang, 2023).
Constructivism: In constructivism, knowledge is not passively received, but actively constructed by learners through interaction with content, tasks, and social contexts. Within this paradigm, learners are central agents in the learning process, constructing meaning from experiences rather than merely absorbing information. This perspective emphasizes learner autonomy, metacognitive reflection, and active engagement, all of which are integral to SRL. Prior research has noted that students’ willingness and ability to self-regulate is influenced by whether the learning environment is student-centered and autonomy-supportive or teacher-directed and passive (Allen, 2022).
LA for SRL Support: In recent years, LA in education has emerged as powerful avenues to enhance learning and teaching. These digital traces can serve as indicators of students’ learning strategies and self-regulation behaviors (Winne et al., 2019). For example, the frequency of log-ins, the regularity of study sessions, time management in meeting deadlines, and patterns of resource access may all signal different levels of self-regulation (Alhazbi et al., 2024). A central promise of LA is to augment instructors’ awareness of student learning processes that are otherwise not directly observable, especially in online or large-class settings. Studies have shown that providing teachers with analytics-driven early warning signals can improve student retention and success, as teachers are able to intervene earlier with at-risk students (Herodotou et al., 2019).
AI for SRL Support: Alongside analytics, AI techniques are increasingly being integrated into educational tools to personalize learning and scaffold complex skills like self-regulation. Chang et al. (2023) outlined design principles for AI chatbots aimed at supporting SRL, emphasizing features such as goal-setting assistance, timely feedback on tasks, and personalized reminders or hints. More recently, Hajian et al. (2025) turned to the motivational side of self-regulation, proposing a Motivation Construction Model that maps three motivational theories onto a sequence of prompts a generative AI tool can deliver. Their model shows how AI can help learners set goals, recognize what makes a task worth their effort, and reframe early setbacks. These are the kinds of supports that often decide whether a student acts on the strategies that analytics make visible. Moreover, AI techniques like machine learning can complement LA by identifying complex patterns in student data that may be overlooked by human observers. These insights can help predict which students are likely to benefit from targeted interventions (Zawacki-Richter et al., 2019). However, even advanced analytics systems often fall short of providing support that genuinely enhances students’ self-regulatory skills; simply alerting a student or teacher to a problem does not ensure the student will know how to adjust their learning strategies (Schumacher & Ifenthaler, 2021). Therefore, current research is increasingly focused on “closing the loop” by not only diagnosing SRL issues via data, but also delivering formative feedback or prompts grounded in SRL theory to help students adjust their approach (Heikkinen et al., 2023).
Zone of Proximal Development (ZPD) and Scaffolding Theory: ZPD is defined as the distance between what a learner can accomplish independently and what they can achieve with guidance from a more knowledgeable other (Vygotsky, 1978). Scaffolding theory, introduced by Wood, Bruner, and Ross (1976), refers to the structured support provided by an instructor or tool that enables learners to complete tasks within their ZPD. Following Wood’s work on adaptive tutoring, the “region of sensitivity to instruction” describes the range of task difficulty where scaffolding is most effective, and the “success boundary” marks the threshold at which learners transition from independent success to requiring external support (Wood & Middleton, 1975). These constructs are central to this study, as they frame the conditions under which AI-assisted hints are expected to be most beneficial for activating metacognitive engagement.
AI-Assisted Assessment: In parallel with advances in LA, recent work has begun to examine the use of AI-driven support during assessment contexts, particularly in systems that provide scaffolding without compromising assessment integrity (Francis et al., 2025). Prior research on intelligent tutoring systems has demonstrated that constrained hints can support adaptive help-seeking while discouraging answer copying or help abuse (Aleven et al., 2006). However, most of this work predates large language models and does not examine how students interact with conversational AI tools during graded assessments.
Conversational AI and Metacognitive Prompting: More recent studies on educational chatbots and large language models have shown that students often use AI tools not only to obtain information, but also to externalize their reasoning, verify understanding, and regulate uncertainty during problem solving (Chang et al., 2023). From a SRL perspective, student-generated prompts can be viewed as observable metacognitive artifacts that reflect monitoring, evaluation, and strategic help-seeking behaviors rather than mere answer requests (Chang et al., 2023; Newman, 2012).
Despite growing interest in AI-supported learning, there remains limited empirical work examining how students refine prompts, allocate AI assistance under constraints, and adapt help-seeking strategies across tasks of varying difficulty during formal assessments, particularly in computer science education contexts. This gap motivates the present study’s focus on prompt refinement, adaptive help-seeking, and the role of difficulty calibration in AI-assisted quiz environments.
3. Methods and Context
To address the RQs, this study employed a mixed-methods design. For RQ1, a data conversion approach (Teddlie & Tashakkori, 2008) was used; prompts were analyzed using deductive content analysis to characterize adaptive help-seeking behaviors and transformed into frequencies for statistical analysis. For RQ2 and 3, quantitative measures were used to examine the relationships between quiz question difficulty, prompt usage, and performance. The following sections describe the definitions and analysis procedures for each component of this study.
3.1. Participants
38 graduate students in the Master of Applied Computer Science program in Western Canada participated in the study as part of coursework. The course focused on assembly language programming and was a required course in the program. No other AI tools were used in the course besides the TA-AI quiz system described in this study. Quiz content was aligned with weekly lecture topics, and students were not grouped by prior experience or other criteria. The course structure and activities encouraged learners to take initiative and regulate their own learning process, providing context to observe behaviors. All participants had access to the integrated AI tools, which were embedded into the course’s learning management system. This context created a constructivist learning environment where students could actively construct knowledge with the AI tools serving as scaffolding. All data were collected in compliance with Institutional Review Board guidelines, and informed consent was obtained from all participants prior to data collection. TA-AI offered on-demand guidance (mirroring a human teaching assistant). The participants were thus empowered to engage in SRL strategies such as help-seeking, self-monitoring, and iterative refinement of their work throughout the course.
3.2. System Architecture
The system is a web-based assessment platform that integrates multiple AI agents in a client-server architecture to support quizzes. The architecture consists of a frontend application, a backend service layer, persistent data storage, and external AI services, as shown in Fig. 1. Students interact with the system through a React and TypeScript frontend deployed on Vercel. The frontend provides the Multiple Choice Question (MCQ) quiz interface, displays the remaining number of available AI hints, and allows students to submit responses and request AI-generated hints. All frontend interactions are transmitted to the backend using RESTful APIs. The backend is implemented using Node.js and Express.js in TypeScript and is deployed on Render. It manages user authentication, quiz logic, hint usage constraints, and communication with external services. All quiz responses, hint requests, system logs, and user feedback are stored in MongoDB Atlas to support secure data management and scalable access.
When a student requests a hint, the backend forwards a prompt-engineered query to the OpenAI server. The generated response is constrained to provide guidance without revealing direct answers. This functionality constitutes the Teaching Assistant AI (TA-AI), which operates synchronously within the quiz workflow. To enforce non-disclosure of answers while still providing useful scaffolding, the system uses a guardrail prompt for TA-AI interactions. Algorithm 1 shows the prompt template sent to the OpenAI API to generate hint-only guidance for a specific MCQ. Algorithm 1 operationalizes the core assessment constraint of TA-AI: providing scaffolding without disclosing answers. By explicitly prohibiting option identification, direct solutions, and option elimination, the prompt reduces the likelihood that TA-AI outputs compromise quiz integrity while still supporting student reasoning through targeted conceptual cues and self-check questions.
In addition to TA-AI, the system includes an Analytics-AI component that processes stored interaction data, including quiz performance and hint usage records. The Analytics-AI generates aggregated summaries identifying commonly missed concepts and usage patterns, which are provided to instructors to support instructional planning and review. The system also incorporates an AI-Improver component that collects student feedback and satisfaction ratings related to AI-generated hints. This feedback is used by developers to iteratively refine prompt engineering strategies and improve system behavior across assessment iterations. Figure 2 shows the quiz interface used in the study. Students may request TA-AI assistance. The remaining number of available hints is displayed within the interface, and each hint request is processed through the backend-mediated interaction with the OpenAI server described above. Overall, the architecture supports controlled AI-assisted help-seeking during assessment while enabling data collection and analysis to inform instructional decisions and system refinement.
3.3. Data Collection
We collected both quantitative and qualitative data to investigate the association between prompt refinement and adaptive help-seeking, as well as the students’ use of TA-AI prompts and the quiz questions for which the hints were requested. Collected data captured evidence of SRL behaviors across its forethought, performance, and self-reflection phases. Quantitative data included system logs, academic performance metrics, and usage analytics. The platform logs captured every interaction with the TA-AI: questions posed to TA-AI, dashboard views in Analytics-AI, and instances of AI-Improver feedback being accessed. These logs enabled an objective measure of how frequently and how students utilized each tool, providing insight into their self-regulatory behaviors (e.g., frequent TA-AI queries indicating active help-seeking). Academic performance data includes weekly quiz scores for the duration of two weeks to measure any performance changes correlating with the use of AI support.
3.4. Prompt Refinement and Adaptive Help-Seeking
For answering RQ1, prompt refinement was operationally defined with two indicators within this study’s context: temporal progression across the instructional period and the interactional structure with the TA-AI usage. Prompts from Week 1 were treated as representing the exploratory phase, during which students were familiarizing themselves with the TA-AI tool under minimal guidance. In contrast, prompts from Week 2 were considered demonstrating more informed use following explicit instruction on the intended role of the TA-AI tool. Additionally, prompts were further classified as single-prompt or multi-prompt interactions based on the system logs. Multi-prompt interactions were defined as when a student used two or more consecutive prompts within the same quiz question, creating a back-and-forth conversation instead of a single one-off request. Adaptive help-seeking was examined through the nature of the prompts generated by students, specifically, the presence of the metacognitive prompts. Drawing on prior literature (Chang et al., 2023), metacognitive prompts are defined as inquiries intended to foster a learner’s own learning judgment and metacognitive growth in contrast to the outcome-based cognitive prompts. Since multiple prompts were allowed to be spent on one quiz question, the student prompt log was further arranged by attempt-level, categorized as follows:
-
Single cognitive (S-C) prompt: Student used one single cognitive prompt in requesting a hint from the TA-AI.
-
Single metacognitive (S-M) prompt (Adaptive): Student used one single metacognitive prompt in requesting a hint from the TA-AI.
-
Multi-cognitive (M-C) prompts: Student used more than one prompt in requesting hints from the TA-AI, in which all prompts are cognitive within the attempt episode.
-
Multi-metacognitive (M-M) prompts (Adaptive): Student used more than one prompt in requesting hints from the TA-AI, in which one or more prompts are metacognitive within the attempt episode.
This decision aims to examine whether any of the conditions could potentially foster or lead to adaptive help-seeking.
3.4.1. Qualitative Coding Scheme and Procedure
The classification of students’ TA-AI prompts into cognitive and metacognitive prompts was conducted using a qualitative coding scheme following a directed content analysis approach (Hsieh & Shannon, 2005). The initial codebook category and definitions was derived from the theoretical framing based on the conceptual paper from Chang et al. (2023) that classify student prompts into cognitive and metacognitive prompts, with adaptation to the code definitions in the context of computer science quiz-taking conditions instead of the general writing/learning tasks.
In this study, cognitive prompts were defined as executive and operational inquiries aiming at obtaining information or solutions needed to produce the quiz answer. In contrast, metacognitive prompts were defined as adaptive, process-oriented inquiries focused on understanding the quiz question, monitoring the student’s own thinking, or evaluating student-generated answers rather than directly requesting the answer itself. The coding procedure followed these steps: (1) The initial codebook developed based on theoretical framing from Chang et al. (2023). (2) Two coders independently coded all 164 prompts using the codebook. (3) Initial inter-rater agreement was calculated, and discrepancies were identified as systematic errors, mainly with prompts related to quiz understanding. (4) The codebook was refined to resolve ambiguous category boundaries. (5) All prompts were then re-coded independently using the refined codebook.
Final inter-rater agreement was 87%, and remaining discrepancies were resolved through consensus discussion. Validation was achieved through inter-rater reliability and consensus-based resolution of disagreements, ensuring consistent interpretation of the coding scheme.
The finalized codebook, including category definitions with inclusion and exclusion criteria, is provided in Table 1. Codes were aggregated at the quiz attempt-level to profile students’ prompting approach. Four profiles were derived: Single-Cognitive, Single-Metacognitive, Multi-Cognitive, and Multi-Metacognitive. Single-Cognitive (one prompt coded as cognitive), Single-Metacognitive (one prompt coded as metacognitive), Multi-Cognitive (multiple prompts, all coded as cognitive), and Multi-Metacognitive (multiple prompts, with at least one coded as metacognitive). To illustrate each prompting profile, Table 2 presents anonymized examples from the dataset.
3.5. Prompt Usage and Performance
For RQs 2 and 3, the frequency of prompt use at the item level and students’ quiz performance were examined. As quiz questions were categorized by difficulty in the system (12 low, 8 mid, and 1 high difficulty), we are able to explore whether students adapt to different prompting strategies when facing different quiz questions, as well as whether certain difficulty questions are optimal for fostering students to engage in adaptive SRL. Difficulty levels were determined by the course instructor based on the cognitive demand of each question: low-difficulty questions assessed recall or direct application of a single concept, medium-difficulty questions required integration of multiple concepts or multi-step reasoning, and high-difficulty questions involved complex problem-solving with non-obvious solution paths.
3.6. Data Analysis
3.6.1. RQ1: Prompt Refinement and Adaptive Help-Seeking
Fisher’s Exact Test was used to examine if the attempt-level prompting approach (adaptive help-seeking) is associated with operationalized prompt refinement: (a) the Instructional Phase (Week 1 vs. Week 2) and (b) Behavioral Structure (Single vs. Multi-prompt).
3.6.2. RQ2: Prompt Usage
Fisher’s Exact Test was used to examine the association between prompting approaches and Task Difficulty. Additionally, Spearman’s rank-order correlations were used to test whether ordinal quiz difficulties and the raw frequency of prompts are associated.
3.6.3. RQ3: Quiz Performance
To evaluate the performance outcome of help-seeking through TA-AI, we employed two binomial logistic regression models. This analysis was based on the binary nature of the dependent variable (Correct/Incorrect). Model 1 (Overall Impact): This model tested the entire corpus of quiz attempts to assess the general relationship between help-seeking and the success on quiz questions. The predictor was a binary variable as hint usage (Prompted vs. Non-prompted), with quiz difficulty included as a covariate. This model tested the baseline association between the presence of TA-AI interaction and quiz correctness, as well as providing behavioral context of help-seeking, specifically, whether students primarily engaged with the TA-AI when encountering difficulty. Model 2 (Scaffolding Efficacy): This model focused specifically on the subset of prompted attempts Aiming to test which specific prompting approaches (S-C, S-M, M-C, M-M) supported students’ quiz solving once a help-seeking attempt was initiated. For both models, results are reported as Odds Ratios (OR) with 95% confidence intervals. Single-Cognitive prompts and the Low Difficulty condition were used as the baseline reference groups.
4. Results
The results section is organized into four parts. We start with descriptive statistics to outline the collected data, followed by analysis findings for each of the three RQs.
4.1. Descriptive Statistics
Table 3 summarizes the descriptive statistics of this study, including participant count, performance, help-seeking attempt (hint usage) and prompt categories frequencies across the two-week period. Participants achieved performance, with mean quiz scores slightly increased from Week 1 to Week 2
4.2. RQ1: How is students’ refinement of TA-AI prompts associated with adaptive help-seeking during quizzes?
As shown in Table 4. Initial analyses examined whether the instructional phase (Week) or the behavioral feature (Single vs. Multi-prompt), which were representative of prompt refinement in the study, influenced the likelihood of adaptive help-seeking. The data is analyzed based on help-seeking attempts (hint usage) rather than raw prompt count. Fisher’s Exact Tests were not significant across weekly and single/multi-prompting conditions. This suggests there are no significant relationships between operational prompt refinement conditions and the manifestation of adaptive help-seeking, as the presence of metacognitive prompts in this study.
4.3. RQ2. How do students use the TA-AI hints across questions of different difficulty?
To examine the relationship between quiz difficulty and student inquiry, we first analyzed whether task difficulty influenced the depth of help-seeking. A Spearman’s rank-order correlation returned no significant relationship between quiz difficulty and the raw frequency of prompts per hint usage This shows that students did not simply ask more questions or invest more prompts to the TA-AI as the quiz difficulty became higher. However, a Fisher’s Exact Test revealed a significant association between task difficulty and the type of prompting approach used Table 5 shows the raw count and percentage of prompting approach distribution across quiz difficulties. The post-hoc analysis of standardized residuals revealed a significant and symmetrical shift in prompting strategy associated with quiz difficulty. In low difficulty questions, students are more frequently using shallow, transactional interaction with the TA-AI: Single-Cognitive (S-C) prompts were significantly more likely to emerge, while instances of Single-Metacognitive (S-M) prompting were significantly less likely In medium difficulty, student behaviors were completely opposite. Single-Metacognitive prompts became the dominant and students were less likely to engage in Single-Cognitive (S-C) prompting However, no significant differences were found across multi-prompt instances and high-difficulty quizzes. These results suggest that in this study, medium difficulty acts as a threshold that discourages shallow answer-seeking and fosters higher-level metacognitive inquiry.
4.4. RQ3. Is the use of the TA-AI hints associated with quiz performance?
To evaluate the impact of TA-AI interaction on quiz performance, two conditions were examined: an overall assessment of the help-seeking condition and a specific analysis of prompted items only.
4.4.1. Help-Seeking Context
Logistic regression was used to establish the baseline relationship between help-seeking and correctness. Results showed a significant negative association between prompting (using hints) and quiz success suggesting that students were 65% less likely to be correct when they asked TA-AI for hints compared to when no hints were requested (see Table 6). This association should not be interpreted as a causal effect of TA-AI on performance, but rather as evidence of students’ calibrated allocation of hint usage when encountering obstacles.
This finding indicates that prompting for TA-AI hints functioned as a strategic help-seeking behavior decision by students: Students allocated their limited hints to the TA-AI specifically when they encountered challenges, rather than using them unnecessarily or randomly.
4.4.2. Scaffolding Efficacy within Prompted Attempts
The model for the prompted subset did not reach conventional significance though descriptive trends were observed. Quiz difficulty emerged as a significant predictor in the model for determining quiz success. Specifically, medium difficulty was significantly associated with a higher likelihood of correctness with odds ratios showing that students were 3.79 times more likely to succeed in these questions compared to the baseline condition of prompting for hints in low-difficulty questions. Additionally, the Multi-Metacognitive approach was associated with nearly 3 times the odds of success compared to the Single-Cognitive prompting approach. These results provide preliminary evidence that the efficacy of TA-AI scaffolding could be sensitive to task challenge, and TA-AI support is most effective when students are working within a specific “sweet spot” of difficulty.
5. Discussion
5.1. Metacognitive Monitoring in Help-Seeking Intent
The overall logistic regression predicting quiz correctness showed a significant negative association between hint usage and correctness While it might suggest that TA-AI hinders performance: quiz questions where students prompted the TA-AI are 65% less likely to be correct compared to non-prompted quiz questions. However, a deeper interpretation of the prompt use contexts might suggest this reflected students’ metacognitive calibration.
In SRL theory, learners are commonly distinguished between proactive and reactive learners, and help-seeking behaviors are seen as volitional, reactive learning strategies in response to obstacles or failures during learning tasks (Newman, 2012; Zimmerman, 2002). In this study, students meaningfully used the TA-AI as a reactive safety net. As each student has limited hints during the quiz, correctly identifying when their internal knowledge was insufficient and external help is needed would be an important factor determining the timing of the student’s use of prompts. This pattern may suggest that students engaged in metacognitive monitoring and decision-making, carefully managing their limited resources based on their perceived ability to benefit from assistance (Ryoo et al., 2025).
This explanation also aligns with the concept within the fundamental definition for adaptive help-seeking as students’ awareness of a lack of understanding (Newman, 1991). It is also documented in related work that “the students who experience greater difficulties during learning (and high hint use is a clear sign of difficulty) tend to come away with lower learning outcome” (Aleven et al., 2006, p. 116). Rather than using the TA-AI as a “cheat code” for easy marks, students reserved assistance for attempts they were more likely to fail independently (in contrast to Park et al., 2025, where unconstrained AI tools like GitHub Copilot induced severe learner dependency in low-level programming tasks). This interpretation is further supported by the data: the significant odds ratio of directly quantifies that prompted attempts were associated with a 65% reduction in the likelihood of correctness, indicating that students selectively allocated their limited hints to questions where they recognized their knowledge was insufficient rather than distributing them arbitrarily.
5.2. Difficulty as a Scaffolding Cue
A central finding of this study is that quiz difficulty was found to serve as a “trigger” for higher-order engagement with the TA-AI. The Fisher’s Exact Tests showed a clear student behavioral shift: low difficulty questions are significantly associated with Single-Cognitive prompts (shallow answer-seeking). In contrast, medium difficulty questions showed a strong and significant association with Single-Metacognitive queries, indicating a transition toward reflective problem-solving.
This behavior shift confirms the presence of a Zone of Proximal Development (ZPD, Vygotsky, 1978) within the TA-AI construct. In low-difficulty scenarios, the AI assistance is superficial and fundamental; in high-difficulty scenarios, the gap may be too wide for AI hints or thinking strategies to bridge. However, in the “Medium” zone, the challenge is sufficient to encourage the student to reflect and adopt higher-order thinking. The prompting performance model, while only marginally significant displayed a descriptive trend that the TA-AI is most successful in medium questions providing the context of exact cognitive nudge required to push a struggling student toward a correct solution.
This difficulty threshold aligns with the “success boundary” or the “region of sensitivity to instruction” identified in foundational scaffolding theory, where external support is most likely to result in meaningful learning gains (Wood & Middleton, 1975). To illustrate, a “sweet spot” problem would be a medium-difficulty question requiring multi-step reasoning (e.g., tracing register values through a sequence of assembly instructions), where the student possesses foundational knowledge but needs a conceptual nudge to integrate the steps. In contrast, a non-sweet-spot problem at low difficulty (e.g., identifying the purpose of a single MOV instruction) elicits shallow, single-cognitive prompts because the gap is minimal, while a high-difficulty problem (e.g., debugging a complex recursive subroutine) may exceed the scaffolding capacity of hints alone, as the knowledge gap is too wide for targeted guidance to bridge effectively.
5.3. Behavioral Trends vs. Performance Outcomes
The marginal significance of the prompting performance model and the lack of a significant correlation between raw prompt volume and correctness suggest that the TA-AI’s value may lie more in the learning process than in the immediate correctness during quiz taking. Although the TA-AI did not guarantee correct answers for the student, it appeared to foster strategic persistence. Students adopting the Multi-Metacognitive prompting approach showed higher odds of success compared with those using Single-Cognitive prompts, while this difference did not reach statistical significance (Multi-Metacognitive predictor This pattern suggests that by shifting the focus from “What is the answer?” to “Am I thinking in the right direction?”, even when quiz scores remain unchanged, the quality of student prompting might represent a meaningful step toward developing the persistent and consistent engagement that drives long-term academic success (Goh, 2026).
5.4. Implications
5.4.1. Design Principles: the “Sweet Spot”
An immediate key takeaway from this study is that TA-AI is most effective when the quiz question difficulty is calibrated. Targeting medium difficulty quiz questions could be a good approach for students’ metacognitive activation. The data suggest that instructors should aim for a “sweet spot” of difficulty to foster better learning: Problems that are challenging enough to require reflective guidance but accessible enough that scaffolding hints (without providing any direct answers) can still lead to a breakthrough. By targeting this ZPD, educators can ensure the AI acts as a thinking partner and not just an obscure source of answers.
5.4.2. Design Principles: the Two Major Use Cases
Two major use cases were documented in the study: Cognitive Prompt is fundamental for basic problem solving to fill the essential knowledge gaps if students are lacking it; Metacognitive Prompt showed deeper thinking to move beyond just seeking quiz answers. Future design can center around these two common use cases by having more tailored components corresponding to them for different prompting conditions. Instead of designing “one-size-fits-all” interfaces, future TA-AI systems could offer some more tailored components for these two specific modes, such as specific scaffolding for “Knowledge Gap” for cognitive prompting requests and “Logic Check” for metacognitive verification. Designing around these confirmed behaviors rather than guessing how students might interact with the system could help create tools that more effectively meet students where they are in the problem-solving process.
5.4.3. Pedagogical/Research Design: Tiered Quota and Dialogue Depth
The current design that limits students to a fixed prompt quota can also influence how students allocate their help-seeking chances. This forced students to be strategic (also avoid help-abuse, e.g., Aleven et al., 2006), and it is reflected in how they would be more likely to prompt in struggling situations. We suggest that educators consider a tiered quota system for future applications. For example, more forgiving limitations for multi-prompts within the same help-seeking attempt, while still keeping the surface-level, help-seeking attempt numbers limited. This could potentially add depth to student-AI interactions that were more scarce in this study. This design encourages students to pivot to deep reflection over quick answers, as it makes “delve deeper” a rewarding return for persistence, or a more forgiving process than “just getting the answer”.
5.4.4. Pedagogical/Research Design: Training on Prompting
A lesson learned from this study is that students did not naturally refine their prompting strategies as they progressed through the two-week period. The ratio of cognitive to metacognitive prompts remained almost identical throughout the sessions, suggesting that effective prompting is a skill that might need to be explicitly taught instead of being developed through intuition or exploration with minimal/limited guidance. To address this, we recommend that educators incorporating AI into teaching practices should treat “Prompt Engineering” as a foundational skill. Providing students with prompting templates and examples, which can potentially help bridge the gap between shallow interaction and meaningful conversation. Without this explicit training, students are likely to remain in a state of unstructured prompting habits even when the TA-AI is capable of providing deeper scaffolding.
5.4.5. Scalability, Transferability, and Constraints
The multi-agent architecture is designed to be scalable, as the web-based platform can accommodate larger class sizes without requiring additional instructor effort for real-time scaffolding. The TA-AI operates independently for each student, and the Analytics-AI aggregates data automatically. However, scalability is constrained by the reliance on external API calls to OpenAI, which introduces latency and cost considerations at scale. Regarding transferability, the system’s design principles are not specific to assembly language or computer science. The prompt-based scaffolding approach and the hint-quota mechanism can be adapted to other STEM disciplines where MCQ assessments are used and where students benefit from guided reasoning rather than direct answers. That said, the prompt engineering and difficulty calibration would need to be re-designed for each new domain context. Key constraints include the dependency on instructor-defined difficulty classifications, the fixed hint quota that may not suit all learners equally, and the current limitation to MCQ-based assessments. Extending the system to open-ended tasks would require substantial modifications to both the TA-AI prompt design and the Analytics-AI pattern recognition.
5.5. Limitations
While this study provides practical insights for computer science educators, several limitations to these findings must be acknowledged. First, the sample size for the prompt usage and quiz correctness regression model was relatively small, which restricted the statistical power to detect more subtle performance effects. This also limited the interpretation of the TA-AI’s influence on students’ performance to mostly descriptive trends and not generalizable, concrete effects in the population. The sample from graduate students may also possess more developed SRL skills than undergraduates, which should be considered in results interpretation. Second, the data were collected over a specific time frame (two weeks), and students likely require more time to develop prompting skills and move past “answer-seeking” to have observable prompt refinement outcomes. Additionally, since the TA-AI was specifically set to withhold direct answers, the effects of quiz success may shift if a more permissive AI system is used. Finally, the current data is limited to closed-ended MCQ assessments, future research should explore whether these help-seeking patterns hold true in open-ended programming design projects. Investigating whether explicit prompt training leads to a measurable change in students’ ZPD over a full semester remains an important next step for this project.
6. Conclusion
Overall, the TA-AI application is found to be contextually helpful such that students are more likely to prompt it when they’re in trouble, and the prompting strategies reflect the question difficulty they face. The data analysis identified a “sweet spot” in quiz difficulty that aligns with the ZPD, where TA-AI scaffolding could be most effective in activating students’ metacognitive thinking. Although the immediate impact on quiz correctness was modest, the move toward higher-level interactions indicates that TA-AI can foster strategic persistence, which is found to be positively associated with consistent engagement and beneficial in the long term within the computer science education domain. The value of TA-AI in CSE may lie less in its ability to provide answers to students, but more in its potential to support the SRL process and reflective thinking.
Acknowledgement
Results presented in this paper were obtained using the Chameleon testbed supported by the National Science Foundation (Keahey et al., 2020). This research was supported by the Social Sciences and Humanities Research Council of Canada (SSHRC), grant numbers 430-2024-00269 (PI: Michael Lin), 435-2026-1299 (PI: Michael Lin), 430-2026-00030 (PI: Marco Ho). Any opinions, findings, conclusions, or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the sponsors.


