1. Introduction

The integration of artificial intelligence (AI) into education has significantly transformed learning experiences (Bailey, 2023; Luckin et al., 2016; Molnar & Szüts, 2018). AI-driven learning technologies—such as intelligent tutoring systems and conversational agents—are now commonly used to provide scalable, on-demand academic support, personalized feedback, and adaptive learning experiences. In computer science and engineering education, these tools offer significant potential to assist students as they navigate abstract concepts, syntax-intensive programming languages, and complex problem-solving tasks. However, the rapid adoption of large language model (LLM)–based tools has also raised concerns regarding students’ over-reliance on direct-answer systems, which may undermine the development of essential problem-solving and critical-thinking skills (Aleven et al., 2018; Kersey et al., 2021).

Recent studies have shown that while AI tools such as conversational code assistants can quickly generate correct solutions, their unstructured use in programming courses may encourage surface-level learning and academic dependency rather than meaningful engagement with underlying concepts (Ai et al., 2024; Graesser et al., 2017; Shen & Tamkin, 2026; Zviel-Girshin, 2024). This challenge is particularly pronounced in introductory programming courses, where students are still developing foundational computational thinking skills. As a result, there is a growing need for pedagogically aligned AI systems that support learning without replacing the cognitive processes central to engineering and computer science education.

To address this gap, this study explores the design, deployment, and evaluation of AI-powered chatbots developed specifically as learning assistants for undergraduate computer science and engineering instruction. Rather than providing direct solutions, the designed chatbots emphasize guided reasoning through structured hints, step-by-step explanations, reflective prompts, and Socratic questioning. Grounded in constructivist learning theory (Bada & Olusegun, 2015) and instructional scaffolding principles (Holden & Sinatra, 2014; Taber, 2018), these chatbots are designed to encourage active engagement, conceptual understanding, and independent problem-solving within authentic classroom contexts.

This research began with an undergraduate C++ programming course and evaluated multiple AI platforms to identify effective approaches for implementing guided-learning chatbots in engineering education before extending to additional courses. Rasa was initially investigated as a rule-based conversational framework capable of structured dialogue management and intent recognition, while the GPT-based approach was explored for its ability to support flexible natural language interaction, contextual reasoning, and adaptive instructional responses. After preliminary development and internal testing, the GPT-based assistant was selected as the primary instructional system for deployment and evaluation in classroom due to its superior conversational flexibility and ability to generate more contextually relevant tutoring interactions. The study examines how students interact with these systems, how guided AI dialogue influences comprehension and problem-solving behaviors, and what limitations arise when deploying chatbots in real classroom environments. By integrating comparative testing for chatbot interactions, student surveys, and semi-structured interviews, this work provides empirical evidence on the role of guided AI assistance as a pedagogically sound alternative to direct-answer AI tools.

Accordingly, this paper seeks to answer the following research questions:

  1. How do undergraduate computer science and engineering students interact with an AI-powered chatbot designed for guided learning rather than direct answer provision?

  2. In what ways does the chatbot influence students’ conceptual understanding and problem-solving skills?

  3. What challenges and limitations emerge when implementing AI-driven chatbots as learning assistants in classroom-based computing education?

By addressing these questions, this study contributes to the growing body of research on AI-enhanced engineering education and offers practical insights for educators seeking to integrate AI tools that support learning while preserving academic rigor. The findings aim to inform the design of future AI learning assistants that align technological capabilities with established educational theory and classroom practice.

2.1. AI in Engineering and Computing Education

Artificial intelligence has become an increasingly influential component of modern educational environments, particularly in STEM and engineering education, where learners must develop both conceptual understanding and procedural problem-solving skills (Chng et al., 2023; Leon et al., 2025; Nawaz, 2024; Zhai & Krajcik, 2024). AI-based educational technologies support a range of instructional functions, including personalized tutoring, formative feedback, automated assessment, and adaptive learning pathways. Prior research suggests that when designed appropriately, AI systems can enhance student engagement, improve learning efficiency, and provide scalable instructional support beyond traditional classroom constraints (Almusaed et al., 2023; Saleem et al., 2025).

In computer science and programming education, generative AI tools have gained rapid adoption due to their ability to produce syntactically correct code, explain programming concepts, and respond to student queries in natural language (Dickey et al., 2024; Franklin et al., 2025). These capabilities offer clear benefits, including reduced barriers to entry for novice programmers and immediate feedback during independent practice. However, the literature (Becker et al., 2023; Franklin et al., 2025) also highlights significant risks associated with the unstructured use of generative AI in programming courses. When AI systems provide complete solutions or code snippets without pedagogical guidance, students may bypass essential cognitive processes, leading to shallow learning, reduced problem-solving skill development, and over-reliance on automated assistance.

Recent studies emphasize that the educational impact of AI in engineering education depends not only on technical performance but also on pedagogical alignment (Naixin et al., 2024; Wan et al., 2025). AI tools that fail to incorporate instructional intent may conflict with learning objectives by prioritizing efficiency over understanding. As a result, researchers increasingly call for AI-supported learning systems that balance accessibility and support with mechanisms that preserve productive struggle and conceptual reasoning.

2.2. Educational Chatbots and Intelligent Tutoring Systems

Educational chatbots and intelligent tutoring systems (ITS) represent a prominent class of AI-supported learning tools designed to provide interactive, conversational assistance to students (Graesser et al., 2017; Molnar & Szüts, 2018). In STEM education, these systems have been used to answer conceptual questions, guide problem-solving processes, and offer feedback during coding tasks. Prior research demonstrates that chatbots can improve learner engagement and persistence, particularly in introductory programming courses where students often experience high cognitive load and frustration. Unlike static instructional materials, chatbots can respond dynamically to student input, enabling personalized and just-in-time support by interpreting student queries and generating responses that align with instructional objectives (Popenici & Kerr, 2017; Winkler & Söllner, 2018).

However, the effectiveness of chatbot-based instruction varies significantly based on system design. Many existing chatbots function primarily as direct-answer tutors, delivering complete explanations or code solutions in response to student prompts. While efficient, such approaches risk encouraging passive learning behaviors and diminishing opportunities for students to reason through problems independently (Bailey, 2023; Chelghoum & Chelghoum, 2025; Ross, 2023). In contrast, guided-learning systems—drawing on principles from intelligent tutoring research—emphasize structured hinting, step-by-step reasoning, and reflective questioning. For example, AI-powered tutors can help students develop problem-solving skills by prompting them to think critically rather than offering immediate solutions (Fadhil & Villafiorita, 2017; Roll & Wylie, 2016). Studies have shown that chatbots offering structured guidance can enhance comprehension and retention by encouraging students to engage in step-by-step problem-solving (Graesser et al., 2017; Molnar & Szüts, 2018).

Despite these promising findings, gaps remain in the literature regarding classroom-deployed chatbot systems in authentic undergraduate engineering contexts. Furthermore, relatively few works examine how students perceive and adapt to non-solution-based chatbot assistance when generative AI tools capable of providing direct answers are readily available. Meanwhile, AI-driven education also raises ethical concerns, including issues of data privacy, algorithmic bias, and the potential for students to misuse chatbots for direct solutions rather than learning concepts (Chelghoum & Chelghoum, 2025). Previous research highlights the importance of structuring AI tools to foster engagement and comprehension rather than passive dependence (Bailey, 2023; Ross, 2023). Addressing these gaps is essential for informing responsible chatbot design in computing education.

2.3. Theoretical Framework

This study is grounded in constructivist learning theory, founded on seminal work in cognitive development and social constructivism (Bada & Olusegun, 2015; Vygotsky, 1978), which posits that learners construct knowledge through active, contextual engagement. In constructivist learning environments, students develop understanding by interacting with meaningful tasks, articulating reasoning, and integrating new information with prior knowledge. These principles are particularly relevant in programming education, where learning outcomes depend on iterative experimentation, debugging, and conceptual transfer.

Instructional scaffolding (Taber, 2018; Wood et al., 1976) further informs the design of effective AI-supported learning systems. Scaffolding involves providing temporary, adaptive support that enables learners to accomplish tasks they cannot yet complete independently. As learner competence increases, scaffolds are gradually removed to promote autonomy. In the context of AI chatbots, scaffolding can be operationalized through progressive hints, contextual explanations, and adaptive prompts that respond to student input without revealing full solutions.

Socratic questioning (Paul & Elder, 2007, 2019), complements scaffolding by encouraging metacognitive awareness and deeper reasoning. Through carefully designed questions, learners are prompted to explain their thought processes, identify errors, and consider alternative approaches. When embedded within conversational AI systems, Socratic dialogue supports reflection and conceptual clarity while maintaining learner agency.

Our chatbot design integrates constructivist learning theory, instructional scaffolding, and Socratic questioning, aligning these theoretical foundations with practical system features. By prioritizing guided reasoning, structured hinting, and reflective prompts over direct answer provision, the chatbots are intentionally designed to support constructivist learning processes in undergraduate programming education. This alignment between theory and design differentiates the present work from generic AI tutoring tools and provides a principled basis for evaluating the educational effectiveness of guided AI dialogue in engineering and computing classrooms. Ultimately, by evaluating a tool specifically tailored to mitigate the risk of over-reliance on AI-generated solutions, this study contributes to the growing body of research on responsible AI integration in engineering education.

3. Methodology

This study employed a mixed-methods research design to investigate the development, deployment, and educational impact of AI-powered chatbots implemented as guided learning assistants in an undergraduate computer science course. The methodology integrates a system design grounded in learning theory with empirical data collection from authentic classroom use. This section details the chatbot development process, instructional design principles, deployment context, and data collection procedures used to evaluate student interaction, perception, and learning experiences.

The proposed approach combines customized GPT-based chatbot design with engineering education-specific instructional prompting to support student learning and problem solving. Unlike many existing chatbot studies that rely on general conversational AI interactions, this study emphasizes structured reasoning support, guided problem decomposition, and scaffold-inspired prompting strategies tailored to engineering coursework. The chatbot prompts were iteratively refined through repeated educational testing scenarios to improve instructional clarity, contextual relevance, and response quality.

In addition, this study differs from prior prompt-engineering research by focusing not only on chatbot implementation, but also on classroom deployment and evaluation across multiple engineering courses using survey feedback, interaction examples, and qualitative student perceptions. The study thus contributes both a practical instructional chatbot framework and an evaluation of its educational use in authentic learning environments.

3.1. Chatbot Design and Development

The chatbots were designed to function as guided learning assistants rather than direct-answer tutors. Design decisions were informed by constructivist learning theory, instructional scaffolding, and Socratic questioning, emphasizing reasoning processes over solution delivery. Across all implementations, the chatbot responses were structured to encourage students to think through programming problems by providing conceptual prompts, progressive hints, and reflective questions instead of complete code solutions.

Multiple AI platforms were explored to evaluate their suitability for educational chatbot deployment in undergraduate computer science and engineering education. The chatbots employ natural language processing (NLP) to provide guided explanations rather than direct answers, thereby encouraging active learning (Sinha & Cassell, 2015). The chatbot’s core functionalities include:

  • Natural Language Understanding (NLU): Identifying user intent and extracting key entities from student queries.

  • Guided Problem-Solving: Providing hints and explanations rather than direct answers.

  • Interactive Learning Paths: Adapting responses based on student’s progress and responses.

The first implementation was developed using the Rasa framework, selected for its robust support of rule-based dialogue management, intent recognition, and contextual conversation tracking. Rasa enabled fine-grained control over conversational flow, allowing the chatbot to maintain instructional intent and adapt responses based on student input. Following the methodology outlined in the official documentation (GmbH, 2024), the Python environment was configured, and the Rasa platform was installed. The chatbot was subsequently built by defining intents, creating training examples, and structuring conversational stories before undergoing iterative testing. Specific procedural steps are detailed in the Results and Discussion section.

A second customized chatbot was implemented using OpenAI’s GPT-based language models. To tailor this platform to our pedagogical requirements, advanced prompt engineering techniques (Ekin, 2023) were employed to constrain model behavior, including explicit role definitions, output limitations, and instruction templates designed to prevent direct answer generation. Systematic adjustments focused on refining role assignments, establishing clear and specific instructions, enforcing explicit constraints, and controlling output verbosity. These prompts were iteratively refined through rigorous testing and feedback cycles to ensure consistent alignment with guided-learning objectives.

3.2. Instructional Context and Deployment

The study was initially conducted in an undergraduate introductory C++ programming course at a small liberal arts institution. The chatbot was introduced as a supplementary learning tool; participation was optional and not required for course completion. Students accessed the chatbot through a web-based interface, allowing for asynchronous use outside scheduled class hours.

The chatbot supported course-aligned topics, including control structures, loops, conditional logic, and basic algorithmic reasoning. Beyond technical programming assistance, the chatbot was configured to provide general course-related information, such as instructor contact details and assignment expectations, to enhance usability and encourage sustained engagement.

Importantly, the chatbot was intentionally positioned as a learning assistant rather than an assessment or grading tool. This design choice reduced performance anxiety and encouraged exploratory interaction, aligning with the study’s focus on learning processes rather than performance outcomes alone. The design piloted in this introductory course was subsequently extended to several other engineering and computer science courses.

3.3. Research Design and Data Collection

To capture both the quantitative and qualitative dimensions of student interaction with the chatbot, this study employed a mixed-methods research design. The chatbot was made accessible to students via web-based interfaces distributed through QR codes, shared URL links in class announcements, and direct emails, allowing for asynchronous assistance at any time. Data collection occurred throughout the duration of the courses via comparative chatbot testing, post-interaction surveys, and informal student interviews.

  • Comparative Testing: Specialized features and overall system performance were evaluated by comparing the custom-built solutions against baseline AI platforms.

  • Surveys: Post-interaction surveys utilizing structured, Likert-scale items were administered to assess perceived usefulness, ease of use, engagement, and the extent to which the chatbot encouraged reflective thinking and problem-solving. The survey instrument was reviewed and approved by the Institutional Review Board (IRB). Participants included undergraduate computer science and engineering students enrolled across three courses: introductory C++ programming, assembly language, and a freshman college experience course.

  • Interviews: To complement the quantitative survey findings, voluntary, informal interviews were conducted with students who interacted with the chatbot during the semester. These sessions were conversational rather than strictly protocol-driven, allowing participants to discuss their experiences, challenges, and perceptions in an open-ended format. This qualitative component provided deeper insight into student experiences with guided-learning AI assistants and helped contextualize patterns observed in the survey data. Students from the participating courses were invited during the instructional period, with the explicit understanding that involvement was entirely optional and would not affect their course standing or grades.

To analyze the qualitative data, interview notes and summaries were iteratively evaluated using a thematic analysis approach. Recurring concepts and patterns across participant comments were grouped into broader categories, and emerging themes were compared across interviews to ensure interpretative consistency. To strengthen trustworthiness, these qualitative observations were triangulated with quantitative survey responses and classroom interaction to determine whether consistent patterns emerged across multiple data sources. Finally, representative student comments were continually revisited during the analysis to confirm that interpretations remained grounded in participant perspectives rather than researcher assumptions. All student participation was voluntary, and no identifiable personal data were included in the analysis. Furthermore, all survey responses were anonymized, and the study protocol was approved by the Institutional Review Board (IRB) in compliance with institutional ethical standards for educational research.

3.4. AI Use Disclosure

During the preparation of this manuscript, ChatGPT (OpenAI, version GPT-5) and Gemini (Google, version Gemini 3) were used for linguistic refinement, including grammatical correction and stylistic rephrasing. The authors have carefully reviewed and edited the generated text to ensure factual accuracy, and the final content reflects their own interpretations.

4. Results and Discussion

Both the Rasa framework and OpenAI GPT-based platforms were initially explored during the development phase of the tutoring system. Rasa was evaluated for its structured intent-based dialogue management capabilities, while the GPT-based approach was investigated for its flexibility in natural language interaction and contextual instructional support. Following preliminary testing, the GPT-based tutoring assistant was selected for classroom deployment and formal evaluation due to its ability to provide more adaptive and contextually rich learning interactions. Consequently, all survey responses and interview feedback reported in this study were collected from student interactions with the GPT-based assistant. The discussion of Rasa in this manuscript is therefore intended to describe the exploratory development process rather than a separately evaluated experimental condition.

This section presents and interprets findings from the mixed-methods evaluation of the guided-learning chatbot across an undergraduate C++ programming course, an assembly language course, and a freshman college experience course. Results from comparative testing, surveys, and student interviews are synthesized and discussed in relation to the study’s research questions and theoretical framework.

4.1. Framework Comparison and Final System Selection

The first implementation utilized Rasa, an open-source framework designed for building sophisticated conversational AI. Rasa was selected because its architecture supports complex dialogue management, making it ideal for providing context-aware feedback and maintaining student engagement in an active learning process. A key advantage of the Rasa platform is its highly customizable dialogue flows, which enable the chatbot to guide students through scaffolding without providing direct answers. Furthermore, the framework’s structured stories and intents successfully tracked conversational context, allowing the system to adapt dynamically to a student’s level of understanding.

During the development process, several challenges emerged regarding the implementation of Rasa. The framework required significant initial setup, particularly for training NLU models and defining conversation flows. Furthermore, scaling and managing large user bases required additional infrastructure overhead.

The second implementation utilized OpenAI’s generative pre-trained transformers, initially deploying GPT-3.5-Turbo (configured with a temperature of \(0.7\) and a maximum token limit of \(1000\)) and subsequently upgrading to GPT-5.5. System prompts were customized using structured instruction templates. This approach leveraged advanced natural language generation and comprehension capabilities, allowing the system to process complex programming queries effectively. Additionally, the LLM generated detailed conceptual explanations, proving highly valuable for educational contexts. The platform also offered streamlined deployment through access to robust pre-trained models; prompts were iteratively engineered and refined to strictly prevent the chatbot from providing full code solutions to students.

A comparison of these two approaches is presented in Table 1.

Table 1.Comparison of the Rasa and GPT Chatbot Design Approaches
Criteria Rasa (Open Source) Customized OpenAI ChatGPT
Beginner Friendliness Moderate High
Coding Requirement Medium–High Low
Strength Full control & orchestration Best conversational UX
Setup Complexity High Low
Hosting Self-hosted / On-prem Managed via OpenAI
Best Learning Outcome Chatbot engineering Rapid AI prototyping
Best For Workflow-heavy enterprise bots Fast conversational assistants

During the initial development phase, both the Rasa framework and OpenAI GPT-based platforms were evaluated as potential architectures for the educational chatbot system. The decision-making process considered multiple technical and pedagogical factors, including coding complexity, setup requirements, conversational flexibility, deployment efficiency, and instructional effectiveness.

Rasa provided advantages in structured workflow orchestration, open-source customization, and on-premise hosting capabilities. However, it required higher setup complexity, greater coding effort, and more extensive intent and dialogue management design. In contrast, the OpenAI GPT-based approach offered lower implementation complexity, rapid prototyping capabilities, and stronger conversational interaction quality with minimal manual dialogue engineering. Preliminary testing indicated that the GPT-based assistant produced more natural, context-aware, and instructionally supportive responses for engineering education use cases.

Based on these considerations, the GPT-based architecture was selected as the primary tutoring platform for deployment and evaluation in this study. The final system emphasized guided conversational support, adaptive response generation, and rapid instructional content refinement, which aligned more effectively with the educational objectives of the project.

4.2. Chatbot Interaction Test Cases

Several test cases were conducted with students to evaluate the performance of our customized chatbot against the baseline OpenAI platform. Hyperlinks to the interaction logs are embedded directly within the figure captions. For example, in response to the user prompt “Loop through numbers 1 to 100 and return the results”, Figure 1 displays the direct code output generated by the standard OpenAI ChatGPT model. In contrast, Figure 2 highlights our customized chatbot’s structured, step-by-step guidance. This side-by-side comparison underscores the intentional design shift from immediate answer delivery to a scaffolded learning experience. Furthermore, while the standard ChatGPT model utilizes Python by default, the reconfigured version explicitly prioritizes C++ to align with the specific course curriculum.

An additional test case was evaluated using the prompt “Given the head of a linked list, remove the \(n\)-th node from the end of the list and return its head”. The corresponding outputs are illustrated below: the baseline OpenAI ChatGPT model provided a direct code solution (Figure 3), whereas the customized chatbot delivered scaffolded guidance (Figure 4). Again, while the baseline ChatGPT model utilizes Python by default, our reconfigured version explicitly prioritizes C++ to align with the course curriculum.

Additionally, the chatbot incorporates course syllabus data—such as instructor contact details and administrative policies—to help students navigate course instructions effectively, as illustrated in Figure 5.

Figure 5
Figure 5.Professor Information

4.3. Student Interaction and Engagement with the Chatbot

To assess the effectiveness of the learning assistant, a study was conducted with a group of computer science and engineering students who used the bot as a supplementary learning tool. At the end of each conversation, a survey was issued. All survey data, usage analytics, and interview feedback presented in this study were collected from student interactions with the GPT-based tutoring assistant rather than the earlier Rasa prototype.

Students were permitted to complete the optional survey multiple times, with each submission corresponding to a distinct interaction or usage experience with the AI tutor; consequently, individual participants may have contributed multiple responses. While more students were introduced to the chatbot throughout the investigation, survey completion remained optional, and not all users submitted responses. Data were explicitly analyzed at the interaction level rather than the individual student level to capture the granular variation across diverse tutoring scenarios and usage contexts. These repeated responses from individual users were intentionally retained in the dataset, as they represented independent interaction events and provided insights into how student perceptions of the tutor evolved over time. Survey results are summarized in Table 2 with 21 usage experiences from students. All survey categories received a mean score between 3.86 and 4.62 out of 5, indicating that participating students had a consistently positive experience.

Feedback collected through post-interaction surveys indicated that students found the chatbot valuable for simplifying complex concepts, providing clear and actionable examples, and troubleshooting common programming errors. The conversational approach enabled students to explore topics at their own pace, with scaffolded hints and tailored explanations available on demand. Overall, the learning assistant proved to be a valuable resource for supporting student learning. By delivering personalized, context-sensitive responses and fostering active engagement, the chatbot demonstrates considerable potential as an educational tool—enhancing students’ independent learning capabilities and cultivating their proficiency in the subject matter.

Table 2.Survey Results from 21 Usage Experiences from Students
Questions Score SD
How well did the chatbot help you understand the subject material? (1-5: Not at all helpful / Slightly helpful / Moderately helpful / Very helpful / Extremely helpful) 4.62 0.80
Did interacting with the chatbot improve your problem-solving skills or critical thinking in this subject? (1-5: Strongly disagree / Disagree / Neutral / Agree / Strongly agree) 4.38 0.67
How confident are you in your understanding of the concepts after using the chatbot? (1-5: Not confident at all / Slightly confident / Somewhat confident / Very confident / Completely confident) 4.33 0.73
How often did the chatbot encourage you to think through the steps of a problem instead of giving a direct answer? (1-5: Never / Rarely / Sometimes / Often / Always) 4.24 1.22
When you didn’t understand a concept, how effective was the chatbot in explaining or guiding you toward the solution? (1:5: Not at all effective / Slightly effective / Moderately effective / Very effective / Extremely effective 4.29 0.85
How engaging was the chatbot in helping you stay focused and motivated to learn? (1-5: Not engaging at all / Slightly engaging / Somewhat engaging / Very engaging / Extremely engaging) 4.33 0.91
Did the chatbot ask questions that prompted you to reflect on or apply what you learned? (1-5: Never / Rarely / Sometimes / Often / Always) 3.86 1.24
How easy was it to use the chatbot for learning? (1-5: Very difficult / Difficult / Neutral / Easy / Very easy) 4.62 0.67
Would you use the chatbot again for learning or problem-solving? (1-5: Definitely not / Probably not / Not sure / Probably yes / Definitely yes) 4.62 0.59

Analysis of survey responses and interview data indicates that students engage more actively with course material when guided through problem-solving steps rather than receiving direct answers. Data collection focused on evaluating student engagement, comprehension improvement, and potential challenges associated with chatbot usage in an educational setting (Aleven et al., 2018; Molnar & Szüts, 2018). Key observations include:

  • Engagement: Students reported that the chatbot’s structured hints encouraged critical thinking and active problem-solving.

  • Conceptual Understanding: Survey responses indicated that students gained a deeper grasp of the concepts when explanations were interactive, iterative, and contextualized.

  • Challenges: A subset of students experienced initial difficulty adjusting to the scaffolded learning style, expressing a preference for immediate code solutions over guided instruction.

Survey results further support these findings. Students reported that the chatbot encouraged them to work through problems step by step, yielding an average score of \(4.24 \pm 1.22\) out of \(5\) on items measuring guided reasoning. High scores for ease of use (\(4.62 \pm 0.67\)) and willingness to reuse the system (\(4.62 \pm 0.59\)) indicate strong overall usability and user acceptance. These results suggest that guided-learning chatbots can sustain student engagement even when they deliberately withhold direct answers.

Furthermore, survey and interview data also indicate that the chatbot positively influenced students’ conceptual understanding and confidence in programming. Students rated the chatbot’s effectiveness in helping them understand course material at an average of \(4.62 \pm 0.80\), while post-interaction confidence averaged \(4.33 \pm 0.73\). Interview responses reinforced these findings, with students noting that guided explanations helped them “understand why the code works” rather than simply copying solutions. Notably, students also reported improvements in problem-solving and critical thinking (\(4.38 \pm 0.67\)), suggesting that the chatbot fostered transferable skills beyond individual assignments. This outcome supports the study’s central premise that guided AI dialogue cultivates deeper cognitive engagement than direct-answer AI tutors in introductory programming contexts.

Although still positive, the lowest-rated survey dimension pertained to the chatbot actively asking reflective or application-oriented questions (\(3.86 \pm 1.24\)). This relatively lower mean and higher standard deviation indicate that while the chatbot successfully prompted reflection for some students, it did not do so with uniform consistency across all interaction sessions—highlighting an opportunity for future prompt engineering refinements.

These findings align with prior research on instructional scaffolding and intelligent tutoring systems, which demonstrates that step-by-step guidance and delayed solution disclosure support deeper learning outcomes. For instance, Roll and Wylie (2016) emphasized that interactive learning tools promoting engagement and problem-solving significantly improve student comprehension. Similarly, Aleven et al. (2018) found that intelligent tutoring systems providing structured hints rather than direct answers enhance student retention and problem-solving capabilities. In this study, the chatbot’s structured hints functioned as cognitive scaffolds, enabling students to progress independently while reducing the frustration associated with complex programming tasks.

Qualitative interview findings and open-ended survey responses addressing Research Question 2 reveal noticeable instructional gains in student conceptual mapping, analytical thinking, and programming mastery. Table 3 summarizes the qualitative coding scheme for student responses regarding chatbot usage and potential enhancements, categorized by primary themes, sub-themes/codes, and frequency of occurrence.

Table 3.Qualitative Code Summary Table
Theme Code / Category Description Count
Pedagogical & Learning Design Socratic / Active Learning Love the chatbot to avoid giving direct answers, ask follow-up questions, and guide students to figure out answers themselves. 3
Personalization & Progress Reflect the needs about adaptive learning paths, tracking of weak areas, and content tailored to individual knowledge levels. 1
Clearer Examples & Explanations Requests better, non-skeleton examples with real data, and simpler breakdowns for beginners. 2
Technical & System Performance System Stability & Limits Mentions crashes/freezing and requests an increase in message limits or usage quotas. 2
Domain Knowledge Accuracy Notes that while institutional info is good, course-specific accuracy needs improvement. 1
User Satisfaction No Improvements Needed / Highly Satisfied Expresses that the chatbot is currently working well, helpful for coursework/definitions, or has no suggestions. 7
Excluded /Invalid Vague / NonResponse Inputs that contain no feedback, system prompt residue, or general non-answers ("N/A"). 5

The findings suggest that the GPT-based tutoring assistant robustly supported conceptual understanding by guiding students through intermediate reasoning steps rather than simply providing final answers. Qualitative student feedback revealed specific areas for implementation improvement. First, requests for Clearer Examples & Explanations highlight the need to develop instructional use cases and dedicated onboarding materials to train students effectively on how to interact with the AI tutor. Second, feedback regarding System Stability & Limits indicated that several students attempted to access the customized AI tutor via free personal ChatGPT accounts; while external platform constraints are not configurable on the backend, this limitation can be proactively addressed through student orientation guidelines. Finally, regarding Domain Knowledge Accuracy, future iterations will refine system prompts and domain-specific knowledge integration tailored to individual course requirements.

Specifically, thematic analysis of the qualitative data highlighted three core operational advantages of this scaffolded approach:

  • 24/7 Accessibility and Asynchronous Support: The chatbot served as a vital asynchronous bridge for students catching up on missed lectures, clarifying underlying concepts, or troubleshooting complex engineering assignments outside of scheduled classroom and lab hours.

  • Scaffolded Learning and Code Restriction: The intentional “no-code/hint-only” restriction successfully prevented students from mindlessly copy-pasting answers, forcing them instead to interpret problem requirements and explore alternative solution paths.

  • Cognitive Reconstruction: By responding to structured Socratic prompts, students were compelled to break down complex programming tasks into smaller procedural steps, manually review their own logic loops, and re-evaluate their code structure, facilitating breakthrough “aha!” moments critical for long-term syntax retention.

Beyond structural knowledge gains, the qualitative feedback indicated that students perceived the chatbot as highly beneficial for improving their problem-solving confidence, debugging technical mistakes, and uncovering the logic behind engineering calculations. Several interaction cases demonstrated that students naturally engaged in iterative questioning and refinement processes while working through programming roadblocks. Rather than interacting with the platform solely as an answer-generation system, learners utilized the chatbot as a dynamic, guided-learning support tool that reinforced meaningful conceptual connections across the course curriculum.

4.4. Student Adaptation to Guided Learning

While overall responses were positive, the data also reveal challenges related to student adaptation. Some students initially expressed frustration with the chatbot’s refusal to provide full solutions, particularly when compared to widely available generative AI tools. Interview responses indicate that this resistance diminished over time as students became accustomed to the guided format and recognized its benefits for learning.

This adjustment period highlights an important tension in AI-supported education: students’ expectations for efficiency versus pedagogical intent. This findings align with a study by Graesser et al. (2017), which found that while students benefit from AI-driven tutoring, some struggle with interactive learning styles, preferring direct answers instead. The findings suggest that transparency in chatbot purpose and consistent instructional framing are critical for successful adoption. When students understood that the chatbot was designed to support learning rather than provide answers, they were more likely to engage productively with the system.

These observations are consistent with prior research showing that learners may initially resist interactive tutoring systems but ultimately benefit from structured guidance once expectations are aligned.

4.5. Technical and Design Limitations

Despite these promising findings, integrating chatbots into educational frameworks introduces several notable challenges. A primary concern involves ensuring that system responses are contextually accurate and strictly aligned with the course curriculum. Prior research by Winkler & Söllner (2018) demonstrated that conversational agents frequently struggle to interpret nuanced student queries, leading to contextual misunderstandings and incorrect guidance. This phenomenon was mirror-sampled in our study; students occasionally reported that the chatbot’s explanations lacked sufficient precision, indicating a clear need for further NLP refinement.

Furthermore, data privacy and ethical considerations remain paramount in AI-driven education. As emphasized by Bailey (2023), educational chatbots must guarantee that student interactions remain confidential, secure, and unbiased. This concern emerged as a recurring theme in our qualitative data, with students expressing minor apprehensions regarding data security and potential algorithmic biases in generative responses.

A distinct architectural challenge encountered during development was the “black-box” nature of the OpenAI ecosystem. Unlike Rasa, which offers open-source transparency for fine-grained architecture tracking, the proprietary OpenAI API provides less visibility into how queries are internally processed. This lack of transparency increased the difficulty of auditing internal logic and tailoring the model to strict pedagogical goals. These operational and technical limitations are summarized below:

  • NLP Refinement: The critical requirement to advance NLU to ensure accurate, context-aware interpretations of student queries.

  • Curriculum Alignment: The necessity of locking system outputs to specific course content to maintain pedagogical relevance and effectiveness.

Several inherent limitations of GPT-based conversational systems must also be considered when interpreting these findings. First, chatbot responses exhibit structural variability across identical queries due to the probabilistic nature of LLM token generation; consequently, students may receive divergent explanations for similar problems. Second, the instructional efficacy of the platform remains highly sensitive to prompt engineering and system configuration. Minor variations in prompt structure, token weight, or contextual framing significantly influenced the relevance, clarity, and pedagogical quality of the generated text.

Finally, although the chatbot incorporated guided prompting and structured reasoning, the system did not feature fully adaptive instructional scaffolding capable of continuously diagnosing individual misconceptions or dynamically adjusting support based on a learner’s evolving understanding. The chatbot primarily delivered generalized conversational guidance rather than personalized pedagogical intervention informed by formal learner modeling. These constraints highlight clear opportunities for future work involving adaptive learning analytics, personalized tutoring strategies, and robust instructional control mechanisms within AI-supported educational environments. Research by Sinha & Cassell (2015) suggests that integrating such adaptive, personalized AI models can significantly optimize long-term educational outcomes.

5. Conclusion and Future Work

This study demonstrates that AI-driven chatbots can effectively support independent learning in computer science and engineering education by fostering active engagement, improving conceptual understanding, and mitigating over-reliance on direct-answer tools. By enforcing a guided, scaffolded approach, conversational agents offer a scalable strategy for enhancing introductory programming instruction while maintaining strict pedagogical alignment.

In the next phase of this research, the project will transition to Google Vertex AI (Google, 2025) to explore its capabilities in supporting advanced data collection and analytics for educational chatbot interactions. A primary motivation for investigating Vertex AI is its support for capturing metadata and detailed interaction logs, which are not readily accessible through standard OpenAI GPT deployments. Access to these data will enable a systematic analysis of student–chatbot conversations, including usage patterns, response trajectories, and learning behaviors. Additionally, Google Cloud services will be leveraged to centralize logs and export interaction data to BigQuery, facilitating scalable querying, visualization, and downstream learning analytics.

Future research will also focus on:

  • Enhancing the chatbot’s capabilities by improving NLP accuracy and expanding its knowledge base.

  • Conducting controlled studies to systematically assess the chatbot’s long-term impact on learning outcomes. While current findings indicate positive student engagement and strong perceived problem-solving support, assessing measurable academic gains remains a critical next step. Future work will deploy empirical measures—such as concept inventories, pre/post-testing, and comparative experimental designs—to rigorously evaluate the system’s impact on student learning outcomes. Such methods will help determine the precise extent to which these scaffolded AI interactions drive engineering performance.

  • Expanding to additional programming concepts beyond C++ to evaluate the chatbot’s adaptability across different computer science domains.

Following the development and deployment of all three platforms, a series of user studies will be conducted with undergraduate computer science students. These studies will involve detailed evaluations of chatbot-student interactions, collection of user feedback, and measurement of learning outcomes. Key metrics will include the chatbot’s ability to guide students through structured problem-solving processes, the clarity and usefulness of explanations provided, and the extent to which the chatbot promotes independent thinking. In addition to pedagogical effectiveness, the evaluation will also consider practical implementation factors such as ease of deployment, customization flexibility, and scalability. These insights will inform recommendations for adopting AI-driven learning assistants in broader educational contexts.


Acknowledgment

We would like to express our gratitude to Benedict College for its institutional support throughout this project. This work was sponsored in part by the Minority Serving - Cyberinfrastructure Consortium (MS-CC) under NSF Award #2234326, as well as the UNCF Henry C. McBay Faculty Research Fellowship Program. Finally, we sincerely thank the students whose participation in the user studies and assessments made this research possible.