Introduction

AI chatbots such as ChatGPT are increasingly integrated into educational settings (Akgun & Greenhow, 2022; Al Shloul et al., 2024; Güner & Er, 2025; Karataş et al., 2024; Oates & Johnson, 2025; Shoeibi, 2023), and discourse around their use has largely centered on concerns about academic integrity (Bin-Nashwan et al., 2023; Cotton et al., 2024; Currie, 2023). Despite low institutional registration rates for ChatGPT EDU, student use of AI tools is widespread, with many reporting experience with platforms like ChatGPT. Reluctance to disclose AI use may stem from inconsistent faculty attitudes and privacy concerns. A zero-tolerance stance toward AI use diminishes the relevance and appeal of traditional university education by failing to develop career-relevant skills. Educators must find a balance between maintaining instructional control and accommodating student preferences.

With current technology, AI chatbots can deliver interactive, socially responsive instruction. Since 2025, some platforms—such as ChatGPT-4o—feature persistent memory, enabling continuous dialogue across a course. This supports the activation of prior knowledge, allowing both students and AI to retrieve and apply long-term information in real time while engaging with new material—an essential process for meaningful learning and knowledge transfer (Bransford & Schwartz, 1999). Vygotsky’s Zone of Proximal Development (ZPD) defines learning as a social process where the student is pushed just beyond their current skill level by a more knowledgeable other (Vygotsky, 1978). Properly configured through structured prompts tailored by the instructor, AI can engage in complex pedagogically grounded interactions. It is not merely a tool for answering questions. This study examines both the benefits and challenges of integrating AI chatbots into postsecondary classrooms. While earlier tools used in MOOCs, such as Zybooks, provided some interactive elements, AI chatbots offer a more dynamic, scalable, and personalized alternative. They deliver real-time support that may enhance critical thinking, problem-solving, and collaborative learning.

Technological change in education is not without risk. Historically, disadvantaged populations have faced algorithmic bias, surveillance, and unequal access to digital resources (Hamilton, 2019; Seyyed-Kalantari et al., 2021). It is not clear if AI chatbots improve outcomes for individuals with disabilities (Wang et al., 2024) and may instead reinforce disparities related to gender and socioeconomic status (Humlum & Vestergaard, 2025). Nonetheless, employers increasingly expect AI fluency, making it vital for universities to provide structured, equitable learning opportunities. Without thoughtful implementation, AI tools risk widening existing gaps in access and literacy.

Our approach addresses both cognitive and equity-related challenges of AI chatbot integration by promoting guidance grounded in pedagogical principles. Scaffolding—where instructional support is gradually reduced as learners gain competence—has been shown to enhance student learning (Naznin et al., 2025; Shanto et al., 2025). Building on this, we propose that instructors design flexible learning environments aligned with the ADDIE framework (Analysis, Design, Development, Implementation, Evaluation), enabling lesson-by-lesson differentiation through ongoing evaluation of student feedback and performance (Branson et al., 1975). This structure aims to prevent the exclusion of students who may lack the resources or skills to fully engage with AI tools.

We do not advocate unregulated student use of AI chatbots, nor suggest that instructors treat them as substitutes for classroom instruction or assessment. Rather, AI can be leveraged to support and extend active learning strategies, improving student engagement and knowledge retention (Al Shloul et al., 2024; Deng et al., 2025; Liao et al., 2024). To this end, we promote system-level prompting, which democratizes access to AI by removing the need for programming expertise or in-depth knowledge of large language models (LLMs). This study describes a grant-supported implementation of AI tutors in two undergraduate computer architecture courses for Computer Science majors (Assembly Language and Microprocessor Organization and Design), at a medium-sized Hispanic-Serving Institution (HSI) and presents results from three class sections where the intervention was applied.

Literature Survey

Since late 2022, ChatGPT is the most widely used AI chatbot in education due to its accessibility and advanced language capabilities (Belkina et al., 2025; Deng et al., 2025; V. R. Lee et al., 2024; Tossell et al., 2024). Even commercial platforms marketed as academic tools such as Chegg, Khan Academy, Course Hero, and Grammarly are built to use ChatGPT on the back end. As such, any review on AI-based classroom strategies in K-12 education will necessarily center on ChatGPT’s integration into learning environments and students’ direct experiences with it.

Unstructured use of AI in classroom settings presents well-documented risks. Studies show that unguided AI use encourages shallow engagement, cognitive offloading—i.e., letting the AI think on your behalf, and weakened problem-solving abilities (Marian et al., 2023; Shum, 2024). A survey of AI-chatbot classroom interventions found early adopters encountered issues with over-reliance, and outsourcing of cognitive and metacognitive skills (Qian, 2025). One study found that frequent use of AI tools was associated with lower critical thinking, and this relationship was explained by cognitive offloading (Loble & Stephens, 2025). Others note the rise of lazy thinking and reduced effort when students rely too heavily on machine-generated output (Brender, Hamamsy, Uittenhove, F., & Bumbacher, 2025). These concerns suggest that over-reliance can reduce student effort, impair memory, weaken ownership of the learning process and result in lower academic outcomes (Fan et al., 2024; Woodruff, 2024).

Education must reinforce instructor-driven learning. As society adopts AI-in-the-Loop systems (Natarajan et al., 2025), postsecondary education requires an Instructor-in-the-Loop model in which teachers structure and regulate chatbot use to maintain rigor. Human-in-the-Loop approaches defer to humans in uncertainty (Memarian & Doleck, 2024); we extend this by positioning instructors as content experts and designers of AI-integrated lessons, shifting them from passive users to orchestrators of AI-mediated learning. Well-designed prompts can improve AI chatbot interactions and should be managed through theory-grounded frameworks aligned with educational goals (D. Lee & Palmer, 2025). Although some studies use chatbots for problem-based learning (Zhu et al., 2024), guided discussion (Tang et al., 2024), or writing tasks (Fan et al., 2024), to our knowledge, only one other study structured student-use of ChatGPT with the ADDIE model (Lim, 2024). A recent intervention using prompt engineering, scaffolding, and structured interaction with AI tutors showed promise, but did not examine how instructor-authored prompts shape student perceptions over time (Kestin et al., 2025). Our work distinguishes itself by using custom-designed chatbots, with attention to student perceptions over time, and regulated feedback. Our study focuses on the following research questions:

  • RQ1: How do students perceive the impact of the AI tutor on reasoning and understanding?

  • RQ2: How does instructor AI design (prompting) shape student experiences with the AI tutor?

  • RQ3: Does the AI tutor support equitable participation and reduce reliance on external resources?

Instructional Design

The ADDIE instructional design framework is a widely adopted model for creating and refining effective educational experiences (Abernathy, 2019; Ding & Toran, 2025; Spatioti et al., 2022). Each phase serves a specific function: Analysis identifies learner needs and challenges; Design sets instructional goals and strategies; Development produces learning materials; Implementation delivers instruction; and Evaluation measures effectiveness to guide revisions. ADDIE is not a strictly linear process. The Evaluation phase allows instructors to revisit and adjust earlier stages based on outcomes and feedback. During Implementation, instructors gather real-time feedback through surveys and classroom observations. This data informs Evaluation, where they assess whether AI tools support or impede learning.

Based on these findings, instructors may revise elements from earlier phases to adjust AI use, refine prompts, reconfigure GPT tools, or modify in-class activities. AI tutoring with LLMs remains experimental. While emerging frameworks provide limited guidance, a flexible and iterative approach is essential. The adaptive application of ADDIE is especially critical in classroom contexts where student engagement and comprehension vary. Rather than postponing evaluation until a course concludes, instructors use ADDIE as a continuous feedback loop to ensure AI integration remains aligned with instructional goals and responsive to changing student needs.

Example Application

In our approach, ChatGPT functions as a pair programming partner, procedurally generating assessment questions and delivering immediate feedback. The student and AI alternate turns: during the AI’s turn, it models an example question; during the student’s, it poses a question the student must answer and articulate their reasoning. Over time, the AI gradually withdraws support, also known as scaffolding, requiring the student to respond with increasing independence. The conversation log is submitted as part of the coursework. This method draws on the principles of pair programming, a collaborative coding practice in which two programmers share a single workstation: one as the driver, who writes code, and the other as the navigator, who reviews each line, suggests improvements, and considers the broader problem-solving strategy (Medel et al., 2021). This system promotes active collaboration and helps learners develop problem-solving skills through peer discussion and real-time feedback. The system is inherently designed to avoid the behavior where a student asks a question and receives a direct answer. To begin the lesson the student submits a system message to a LLM, such as ChatGPT. OpenAI’s Custom GPT feature allows instructors to provide a system message seamlessly, without requiring students to manually copy the system message for each conversation. An example prompt follows:

You are a module for a class in Computer Architecture based on the Hennessy and Patterson textbook. Start the game by reminding the player that you are not perfect, and you may both make mistakes. The goal is to learn from mistakes you both make. After 8 questions, stop the conversation and ask the user to upload the conversation to Canvas. Take turns with the player. On your turn, model a question for the player and explain your thought process. Invite the student to correct your approach if it is wrong. On their turn, present the prompt for a question. The user must provide the answer with sufficient explanation of their reasoning–which you may ask more questions to force them to articulate their thought process. Supports given to the user (scaffolding) decrease over time. Do not end the turn unless the user wants to move on and has no more questions. Topic: Branch prediction using a branch table buffer. Example questions: Why does moving the branch prediction logic from the MEM stage to the ID stage reduce the number of pipeline stalls, and what hardware change is required to enable this optimization? How does the branch prediction mechanism use a branch table buffer to improve processor efficiency, and what role does the 2-bit history play in this process? In what ways does speculative execution (i.e., predicting the next instruction) balance the trade-off between performance gains and the cost of misprediction in pipelined processors?

The prompt employs several well-established strategies to elicit tutor-like behavior from a LLM. Prompt conditioning constrains the scope and persona of the AI, placing it in a structured role with explicit turn-taking between student and system. This helps reduce hallucination drift by specifying both topic and function. Positive prompting instructs the model to act out pedagogical behaviors such as modeling expert reasoning, encouraging articulation of thought processes, and providing scaffolding. Including example questions at the end of the prompt functions as few-shot prompting, allowing the model to infer the expected depth, structure, and content of subsequent interactions and to generalize this pattern across the assessment. Providing the AI with only the course textbook and a few examples was sufficient to generate satisfactory assessment questions. Finally, the prompt introduces an uncommon but pedagogically meaningful strategy of epistemic humility by explicitly acknowledging that the AI may be wrong and inviting the student to identify and discuss errors, thereby reframing the interaction as a collaborative learning process (Celi et al., 2026).

Methods

A quantitative analysis was performed with the population receiving an intervention in the form of AI-driven assignments. These out-of-class assessments, similar to quizzes, were designed to engage students with course content before lectures. The AI chatbot activities, guided by structured prompts from the instructor, were incorporated directly into the course design. Attitudinal surveys were administered at multiple points across the semester to monitor student perceptions. Based on interim feedback, the instructor revised the prompts to improve clarity, relevance, and alignment with course objectives.

Participants

Participants were consenting students over the age of 18, enrolled in two courses: Assembly Language and Computer Architecture. The study was conducted over two semesters with Assembly Language offered in Spring 2025 and Computer Architecture offered in Spring and Fall 2025. Assembly Language is a lower-division course typically taken by sophomores. It covers number systems, representation of integer and fractional data, principles of reduced instruction set architectures, basic digital logic, and assembly language programming. Computer Architecture Organization and Design is an upper-division course generally taken by juniors. Building on Assembly Language concepts, it includes the design of a reduced instruction set microprocessor, including instruction set architecture, control unit, interrupt handling, and memory organization. The course also introduces optimization topics such as parallelism and security. Both courses are required for Computer Science majors. Assembly Language is not required for Computer Engineering majors, and both courses are electives for Information Systems majors. All courses were taught by the same instructor, a senior faculty member with twelve years of teaching experience, six years being classes taught in a flipped classroom format using active learning (Lage et al., 2000).

During the first week of class, a demographic survey was distributed, which also included exclusion criteria such as age and prior participation in the study. Participant demographics are summarized in Table 1. The typical student is a Junior non-White Hispanic/Latino, enrolled in Computer Architecture. The average age was 22.3 years (median: 21). Students were enrolled in an average and median of 14 semester units and reported working an average of 11.33 hours per week (median: 10). A total of 78.0% indicated prior experience with AI platforms such as ChatGPT, Claude, Deepseek, Gemini, Grok, and Microsoft Copilot.

Table 1.Demographics of participants in study. Consenting: percentage of individuals who gave informed consent.
Property Distribution
Major 4 Computer Engineers (10%), 32 Computer Scientists (80%), 4 Information Systems (10%)
Course 15 Assembly Language (19.5%), 62 Computer Architecture (80.5%)
Gender 33 Male (80.5%) and 7 Female (17.1%)
Hispanic/Latino(a) 26 Hispanics/Latino(a)s (65.0%) and 14 non- Hispanics/Latino(a)s (35.0%)
Race/Ethnicity 1 American Indian, Alaskan Native, Native Hawaiian or other Pacific Islander (2.3%), 7 Asian (16.3%), 1 Black or African American (2.3%), 17 White (39.5%), 3 preferred to not say (7.0%), 9 other (20.9%), and 3 two or more races (7.0%).
Standing 11 Seniors or later (26.8%), 25 Juniors (61.0%), 4 Sophomores (9.8%), and 1 preferred to not say (2.4%)
Class Sizes 23 in Fall 2025 Computer Architecture, 15 in Spring 2025 Assembly Language, 39 in Spring 2025 Computer Architecture
Consenting 13 in Fall 2025 Computer Architecture (56.5%), 15 in Spring 2025 Assembly Language (100%), and 29 in Spring 2025 Computer Architecture (74.4%)

Data Collection and Survey Instrument

Courses had attitudinal surveys conducted periodically throughout the semester as a part course requirements. Ethical consent was obtained (CSUB IRB 25-150). The surveys gathered feedback that the instructor used to evaluate the quality of the AI chatbot. The study was announced only after all academic obligations for the semester have been fulfilled and grades have been submitted to avoid coercion. The attitudinal survey contains the following questions:

  1. The AI tutor has enhanced my overall learning experience in this course. (RQ1)

  2. The AI tutor encourages me to think critically and solve problems independently. (RQ1)

  3. The AI tutor helps me recall and apply previous concepts relevant to the task at hand. (RQ1)

  4. The AI tutor is user-friendly and accessible during class activities. (RQ3)

  5. The feedback provided by the AI tutor is as helpful as feedback from the instructor. (RQ2)

  6. The prompts provided by the instructor help me effectively utilize the AI tutor. (RQ2)

  7. The AI tutor supports my learning regardless of my background or prior experience with similar tools. (RQ3)

  8. I feel prepared before attending lecture, especially when I use prompts or AI tools to guide my preparation. (RQ1/2)

  9. I can understand examples covered in lecture, particularly those that incorporate AI tools or resources. (RQ2)

  10. The course materials and AI tutor reduce my need to seek additional resources beyond what is provided by the instructor. (RQ3)

The same attitudinal survey was administered at three points in the semester (T1 start, T2 mid, T3 end). The analytic unit was a student’s survey response at a given timepoint. Since participation varied across waves, results are interpreted as pooled time trends and should not be viewed as a fully within-student longitudinal estimate. Response counts declined across waves, with T3 item-level responses ranging from 8 to 17 depending on the item.

The survey instrument measured students’ perceptions of the AI tutor and its role in supporting the instructional goals of the intervention. Items were aligned with the research questions and with the constructs introduced earlier in the manuscript: articulation, reflection, and independent reasoning. Because the instrument measured attitudes rather than direct performance outcomes, the alignment is used to clarify interpretation of the survey results rather than to claim direct measurement of all constructs.

Survey Instrument Reliability

Internal consistency was assessed for the 10-item attitudinal survey as a measure of students’ perceived instructional support from the AI tutor. Cronbach’s alpha was calculated separately at each timepoint using complete cases because the respondent pool and item completion varied across T1, T2, and T3. Although anonymized respondent IDs were available, participation varied across survey waves, limiting the interpretability of test–retest stability. Test-retest stability was therefore not treated as the primary reliability criterion because the survey was administered during an instructional intervention intended to change students’ perceptions over time; instability across T1, T2, and T3 may reflect expected development in student attitudes, attrition, and changing course experiences rather than measurement unreliability.

Statistical Analysis

To assess the sensitivity of student responses to change over time, two complementary ordinal logistic regression models were estimated. Ordinal logistic regression was specified as a proportional-odds model with a logit link, where higher response categories correspond to more favorable student perceptions. In the first model, time was treated as an ordered numeric predictor representing measurement wave to test for an overall monotonic trend in agreement. This specification is sensitive to gradual, cumulative shifts in perception without assuming that changes occur at specific intervals. In the second model, time was treated as a categorical factor with T1 as the reference category, enabling direct comparisons between later timepoints and the initial measurement. This specification is more sensitive to non-linear or delayed effects, such as changes that emerge later in the semester. Both models fit a proportional-odds ordinal logistic regression of the form:

log(P(Y≤k)P(Y>k))=θk−β for k=1,…,K−1 Where Y is the ordinal survey response, k indexes the response threshold, θk are threshold parameters, and β represents the effect of time on the log-odds of higher agreement. A positive value of β indicates increased odds of selecting a higher agreement category at later timepoints. All models were estimated under the proportional-odds assumption, which implies that the effect of time is constant across response thresholds.

Ordinal logistic regression was selected because survey responses were measured on a five-point Likert-type scale with ordered but non-interval categories. This approach is well suited for detecting changes over time when responses may shift gradually or unevenly across categories. Because not all participants responded at every timepoint, the data included missing responses, uneven participation across measurement waves, and limited repeated observations per participant. The use of proportional-odds ordinal logistic regression under such conditions is well established in the statistical literature (Agresti, 2010; Long & Freese, 2006). Separate ordinal models were estimated for each conceptually related survey item to characterize patterns of change in student perceptions; no formal adjustment for multiple comparisons was applied, and results should be interpreted descriptively. The proportional-odds assumption implies that the effect of time is constant across response thresholds. The study observed small and uneven wave-specific sample sizes reducing statistical power, particularly for T3 comparisons.

AI Disclosure

The work is a majority, original contribution by the human authors. Humans performed data cleaning using PowerShell and executed code in a Python 3 environment using the Python Data Analysis Library (PANDAS) on a Windows PC. Figure and tables were generated by humans using Microsoft Office. ChatGPT-4o and 5.2 were used to suggest changes to grammar, spelling and tone, using a prompt to avoid auto-generation—e.g., “Suggest editorial changes but do not automatically rewrite the text provided,” much in the spirit of our work. Thus, AI did not generate the content contained in this paper.

Figure 1
Figure 1.Responses to a Likert question measuring self-identified critical thinking and problem solving skill among students using AI tutors, over time.
Figure 2
Figure 2.Responses to a Likert question measuring effectiveness of feedback from instructor designed AI tutors compared to the instructor, over time.
Figure 3
Figure 3.Responses to a Likert question measuring self-identified course preparedness of students using AI tutors, over time.
Table 2.Descriptive statistics of survey responses by timepoint. Mode: Modal response. f: Proportion of respondents at that timepoint.
T1 T2 T3
Question Mode f N Mode f N Mode f N
1 N 0.441 33 A 0.500 24 A 0.647 8
2 A 0.471 33 A 0.417 24 SA 0.647 8
3 A 0.471 33 A 0.417 24 A 0.353 17
4 A 0.500 33 A 0.522 24 A 0.647 17
5 N 0.324 33 A 0.391 23 A 0.588 17
6 A 0.412 33 A 0.417 24 A 0.588 17
7 A 0.441 33 A 0.625 24 A 0.471 17
8 N 0.412 33 A 0.391 24 A 0.688 17
9 N 0.412 33 A 0.565 28 A 0.625 17
10 A 0.294 33 A 0.304 23 A 0.556 16
Table 3.Descriptive statistics, and results of ordinal logistic regression with ordered numerical trend. β: Estimated slope of predictor. SE: Standard error. OR: Odds ratio. 95% CI: 95% confidence interval.
Question β SE p OR 95% CI
1 0.325 0.334 0.330 1.384 [0.720, 2.661]
2 0.679 0.347 0.050 1.971 [1.000, 3.890]
3 -0.034 0.271 0.901 0.967 [0.569, 1.644]
4 -0.076 0.277 0.785 0.927 [0.539, 1.596]
5 0.772 0.283 0.006 2.165 [1.243, 3.772]
6 0.219 0.267 0.412 1.245 [0.738, 2.101]
7 -0.260 0.281 0.357 0.771 [0.444, 1.340]
8 0.480 0.281 0.088 1.166 [0.931, 2.801]
9 0.191 0.275 0.487 1.210 [0.707, 2.073]
10 0.334 0.264 0.207 1.396 [0.832, 2.343]
Table 4.Results of ordinal logistic regression with categorical time. β: Estimated slope of predictor. SE: Standard error. OR: Odds ratio. 95% CI: 95% confidence interval.
T1 to T2 T2 to T3
Question β SE p OR 95% CI β SE p OR 95% CI
1 -0.155 0.492 0.753 0.857 [0.327, 2.248] 1.086 0.747 0.146 2.963 [0.6859, 12.795]
2 -0.169 0.517 0.744 0.845 [0.307, 2.326] 2.112 0.792 0.008 8.261 [1.749, 39.016]
3 -0.463 0.513 0.367 0.630 [0.230, 1.720] 0.020 0.552 0.971 1.020 [0.346, 3.012]
4 -0.773 0.450 0.122 0.462 [0.173, 1.229] 0.081 0.576 0.888 1.084 [0.3510, 3.350]
5 0.071 0.493 0.886 1.073 [0.409, 2.820] 1.828 0.606 0.003 6.2183 [1.895, 20.410]
6 -0.685 0.516 0.184 0.504 [0.183, 1.386] 0.623 0.548 0.256 1.865 [0.637, 5.463]
7 -0.281 0.509 0.580 0.755 [0.279, 2.045] -0.514 0.575 0.371 0.598 [0.1938, 1.846]
8 0.434 0.506 0.391 1.5433 [0.573, 4.160] 0.970 0.571 0.089 2.637 [0.861, 8.075]
9 -0.307 0.476 0.519 0.736 [0.290, 1.870] 0.543 0.567 0.338 1.721 [0.567, 5.221]
10 -0.084 0.505 0.868 0.920 [0.342, 2.473] 0.743 0.536 0.166 2.102 [0.734, 6.014]

Results and Discussion

The results in Table 2 show generally high agreement across items, with several items increasing in agreement over time; some items shift from Neutral at T1 to Agree at later timepoints, while others begin at Agree and remain there. Descriptive patterns alone cannot establish change in attitudes over time as they may reflect agreement bias or self-selection rather than attitude change. Instead, conclusions regarding changes over time are drawn from ordinal time-series analyses. The survey demonstrated acceptable to high internal consistency across administrations: T1 α = .812, T2 α = .938, and T3 α = .930. These estimates are interpreted descriptively, particularly at T3, where the complete-case sample was small.

In Question 2, students were asked, “The AI tutor encourages me to think critically and solve problems independently.” Responses shifted upward over time (see Figure 1). At T1, responses clustered around neutral (N) and shifted to agreement (A). Ordinal logistic regression with time as an ordered trend indicates a positive linear effect on agreement (β=0.68, OR=1.97). This suggests that the trend is observed among respondents and is consistent with improved perceptions (p=0.050). Regression with time as categorical indicates where the change occurs, with strong evidence existing between T1 and T3 (β=2.11, OR=8.26, p<0.01). Students self-identified a positive impact of the AI tutor on their ability to think critically and solve problems independently, and it became more pronounced toward the end of the semester.

In Question 5, students responded to the question, “The feedback provided by the AI tutor is as helpful as feedback from the instructor.” Attitudes began at neutral (N) and shifted to agreement or strong agreement (A/SA) by timepoint T3 (see Figure 2). Ordinal logistic regression indicates a positive shift over time (β=0.772, OR=2.16, p<0.01), suggesting increasing alignment toward positive perception among remaining respondents. As with Question 2, there is no strong evidence of the difference between time T2 and T3, but T3 shows an increase relative to T1 (β=1.828, OR=6.21, p<0.005). The AI tutor was tuned via instructor-generated prompts. When posed Question 6, “the prompts provided by the instructor help me effectively utilize the AI tutor,” there was no statistically significant difference. However, the estimated effect was positive and supported a trend in the same direction as Question 5 (β=0.219, OR=1.245).

In Question 8, students reported if they, “[felt] prepared before attending lecture, especially when I use prompts or AI tools to guide [their] preparation.” In Figure 3, responses were initially centered around neutral (N) at T1, with a plurality already indicating agreement (A/SA). At T3, agreement remained high with no strong disagreement (SD). Across models, estimated effects were consistently positive. An ordinal logistic regression with ordered time indicated a positive trend over time, though this effect did not reach conventional levels of statistical significance (β=0.480, OR=1.615, p<0.1). When time was modeled categorically, there was no statistically detectable difference between T1 and T2 or T1 and T3.

For RQ1, students who completed the relevant survey items self-identified a positive impact of the AI tutor on their reasoning and conceptual understanding, particularly with respect to critical thinking and independent problem solving. Across timepoints, observed responses consistently shifted toward higher agreement, with the strongest alignment observed at the final measurement. However, this final measurement also had the smallest respondent pool. While some effects reached conventional levels of statistical significance and others reflected marginal or trend-level evidence, the direction of change was coherent across related items. Taken together, these patterns suggest that this implementation was experienced as distinct from typical AI use in educational settings. Rather than functioning primarily as an answer-generating tool, the AI tutor emphasized scaffolding, reflection, and independent reasoning. This distinction appears to have been recognized and favorably received by students, particularly among those who remained with the study at later data collection timepoints, who may represent a more engaged subset of the original participants. It is also possible that later measurements coincide with periods of increased academic engagement, such as preparation for practical examinations, which may have amplified students’ attention to course-related tools.

For RQ2 and RQ3, responses indicated favorable perceptions of instructor-designed prompts and the AI tutor’s accessibility and inclusivity. However, high levels of agreement across timepoints suggest the presence of an agreement bias, which limited the instrument’s ability to discriminate differences in student experience. As a result, conclusions for these research questions should be interpreted cautiously, and future work should employ more discriminating item designs to better isolate the effects of instructor mediation and equity outcomes.

Prompt Refinements

A key feature of the ADDIE model is continuous evaluation and refinement of curriculum. In prompt-engineering, this is supported by the concept of revision (Cruz et al., 2025). When generating prompts for AI chatbot assignments, a template of basic instructions provided guardrails and instructional framing for all conversations. This template was modified based on feedback and observations of the instructor. In the following we discuss modifications made according to the ADDIE model and revision.

The original prompt implemented a traditional quiz. Questions were either multiple-choice or free response, such as if the user needed to provide a numerical answer. Unlike a traditional quiz, the user was expected to articulate the reasoning behind their answer, and ChatGPT would not proceed to the next question until it felt that the user had supported their answer with sufficient discussion and reflection. However, it was found that the models used at the time (ChatGPT-4o, o1, o3 and 5) had difficulty with manipulating binary numbers. The prompt was altered in two ways. First, the chatbot was given a guardrail based on epistemic humility: Start the game by reminding the player that you are not perfect, and you may both make mistakes. The goal is to learn from mistakes you both make… Invite the student to correct your approach if it is wrong. Second, the chatbot’s behavior was altered to have a turn-taking behavior where the chatbot and the user alternate in responding to a generated question: Take turns with the player. On your turn, model a question for the player and explain your thought process… On their turn, present the prompt for a question. This resulted in a modelling behavior the students found beneficial. Framed in this way, mistakes made by the AI chatbot are not sources of misconception, but learning opportunities.

Limitations

Participation declined across data collection timepoints, with sample sizes decreasing from 33 at T1 to as few as 8 at T3. This degree of attrition reduces statistical power, limits the representativeness of the final respondent group, and introduces the possibility of self-selection effects and attrition bias. Observed shifts in response distributions should be interpreted cautiously, as they may reflect the perceptions of a more engaged subset of students. Though ordinal logistic regression models the ordered nature of Likert responses, the analyses treat observations as independent. Because some students contributed responses at multiple timepoints, unmodeled within-student correlation may understate uncertainty in standard error estimates. Mixed-effects or Generalized Estimating Equations (GEE)-based ordinal models were considered; however, sparse repeated observations and uneven participation across timepoints limited feasibility. These results are best interpreted as exploratory indicators of overall trends rather than estimates of within-student change. Reliability evidence was limited by attrition and small complete-case samples at later timepoints, especially T3; future studies should include larger matched samples to support stronger temporal reliability analysis.

The proportional-odds assumption inherent to the ordinal logit model was specified a priori. Formal diagnostic testing was constrained by small cell sizes at later timepoints, though this modeling approach is commonly applied in educational survey research under similar conditions. Multiple survey items were analyzed without formal correction for multiple comparisons; however, p-values are reported transparently and interpreted with effect sizes, directionality, and consistency of patterns across items, consistent with the exploratory nature of the study.

Conclusion

This study demonstrates that AI tutors can support student reasoning and engagement in computer architecture courses when their use is deliberately structured and instructor-mediated. When grounded in Vygotsky’s Zone of Proximal Development and implemented through the ADDIE framework, LLMs can function as scaffolded learning partners rather than answer-generating tools. Survey results indicate that students increasingly perceived the AI tutor as supporting critical thinking, independent problem solving, and effective feedback, particularly toward the end of the semester. While participation attrition and the exploratory design limit claims, the consistency of observed trends suggests that instructional design—especially instructor-authored prompts and controlled interaction patterns—plays a strong role in shaping productive AI use. These findings underscore the importance of an Instructor-in-the-Loop model of learning and suggest that the educational value of AI in engineering classrooms lies less in the technology itself than in the instructor-designed classroom structures that govern its use. Future work should include more direct measures of learning, reasoning, and engagement, including pre/post assessment items or performance-based measures aligned with the AI-tutoring activities.

Support for this work was provided by the California Learning Lab’s ELEVATE grant program. The statements, findings, and conclusions expressed are those of the authors and do not necessarily reflect the views of the California Learning Lab or the State of California.