Introduction

As computer science enrollments continue to surge globally, the widening diversity gap among student populations underscores a critical need for instructional strategies that are both scalable and equitable (Bassner et al., 2024; Córdova-Esparza et al., 2024). Within large computer science (CS) courses, supporting student learning through equitable, scalable strategies is a topic of interest. In this context, equitable learning refers to pedagogical approaches that accommodate diverse student backgrounds, varying prior programming experiences, and individualized learning paces, ensuring all students have a fair opportunity to succeed. While strict content pacing and automated grading schemes are frequently implemented as administrative solutions to manage growing class sizes, these strategies can inadvertently hinder equity by imposing a rigid, one-size-fits-all structure that limits students’ abilities to connect with their own backgrounds, learning preferences, and perceived experiences. Moreover, these strategies tend to prove challenging as they scramble the already varying signs of student agency: time investment, self-evaluation, and preparedness.

To counter the rigid nature of automated instruction in large-scale computing courses, active learning has emerged as a vital alternative; by placing students at the center of the instructional process, it fosters the personalization necessary to achieve more equitable outcomes (Zangari et al., 2025). In areas like programming and software development, strategies such as inquiry-based learning, peer instruction, and project-based learning (PBL) have helped students build stronger conceptual understanding and retain knowledge over time (Amaya Chávez et al., 2020; Córdova-Esparza et al., 2024). These methods also contribute to more inclusive learning environments, especially for students from underrepresented backgrounds. Alongside these strategies, having students create their own learning materials has emerged as a promising way to increase motivation, metacognition, and ownership over learning (Ariza, 2023; Vasilchenko et al., 2020). The self-flipped classroom model has further shown that giving students a more active role can strengthen confidence and engagement with course content (Vasilchenko et al., 2018, 2020).

Building upon the principles of active learning, artificial intelligence (AI) tools—particularly conversational agents—offer a highly scalable medium for students to engage in project-based, self-directed exploration (Bassner et al., 2024; Deslauriers et al., 2019). However, these tools can lack personalization and sometimes fall short in addressing students’ social and emotional needs, creating barriers to adoption (Algerafi et al., 2023; Kuhail et al., 2023). To address these gaps, recent work highlights the value of involving students in designing AI learning agents to make them more relevant, empathetic, and effective (Denny et al., 2024; Feijóo-García et al., 2023). When students co-design agents, they deepen their understanding of course concepts while practicing communication and anticipating peer needs (Bassner et al., 2024; Córdova-Esparza et al., 2024; Zangari et al., 2025). This approach supports agency and empathy and aligns with broader goals to create active and meaningful learning experiences in CS education (Denny et al., 2024; Feijóo-García et al., 2025; Kuhail et al., 2023).

This paper highlights the potential of integrating student-created intelligent agents into computing education to promote engagement, creativity, and deeper learning. Our work extends a previously published idea and pilot study conducted in two small-scale computer science courses, one focused on software design and engineering and the other on user interface (UI) software principles (Zangari et al., 2025). While this concept has been presented, it has yet to be explored in the context of large computer science courses. In this exploratory study, we report on the implementation of this approach in a large software design and engineering (SWE) course (over 700 students enrolled), presenting insights into students’ perceptions of learning, performance, time investment (dedication), and overall satisfaction and engagement with the strategy. This study serves as an exploration into how the self-flipped approach fares when implemented in real classroom contexts. While active learning requires students to dynamically participate in classroom activities, student-centered learning more broadly shifts the focus of instruction from the teacher to the student; our approach of having students co-design AI agents leverages both of these distinct concepts to create meaningful learning experiences in CS education.

Motivated by the need to bridge scalable AI tools with equitable, student-centered active learning, this study addresses the following research questions:

RQ1: How do undergraduate students engage in the design of student-created tutoring agents in a large software design and engineering (SWE) course?

RQ2: How do undergraduate students perceive the role of student-created tutoring agents in supporting their software engineering skill development?

Background

Active Learning in Computing Education

Active learning and its role in computer science education has been widely studied in recent years, with evidence continuously pointing to its effectiveness in improving student engagement and performance. When it comes to more abstract areas of CS such as programming and software development, active learning strategies were seen to enhance student conceptual understanding and retention when paired with inquiry-based learning, peer instruction, and team problem solving (Amaya Chávez et al., 2020). Prior studies (Córdova-Esparza et al., 2024) also note that active learning strategies are especially helpful for students with underrepresented backgrounds, helping in the fight to close equity gaps through more inclusive classroom environments.

Another popular subset of active learning is PBL; one that has proved to be effective in CS and engineering courses. When examining PBL in the engineering context, past work has determined that the strategy has allowed students to meaningfully improve their academic performance and perceptions of relevance/satisfaction (Amaya Chávez et al., 2020). These findings corroborate broader ideas in the CS education community which agree that passive instruction must make way for active learning in large-scale classrooms.

The pedagogical value of these approaches is clear: active learning and PBL learning strategies provide much needed instructional benefits for students from all backgrounds. Adopting them and further studying their effects in large-scale classes continues to be a developing area in CS education.

Student-created Artifacts and Self-flipped Classrooms

In recent years, education communities in science, technology, engineering, and mathematics (STEM) have begun to note the benefits of having students create their own instructional resources. Student-generated content has emerged as a promising strategy which allows students to better apply abstract concepts. Specifically, student-created videos have been the subject of numerous studies, explaining that these videos help students increase meta-cognition, content retention, and critical thinking. When it comes to promoting deeper engagement, strategies like iterative design and reflection have been shown to be successful in engineering classrooms (Ariza, 2023). Moreover, past work has also bridged this success to CS, where a functional programming course employed video creation, resulting in improved student motivation, agency, and satisfaction (Feijóo-García & Gardner-McCune, 2020; Vasilchenko et al., 2020).

Complementing this approach is the “self-flipped” classroom model. This model allows students to be leaders of their learning as the emphasis is placed in a learner-first way through student-created learning materials. This new mode of teaching has proven to be beneficial for promoting student confidence and understanding in course material. It has been shown numerous times that when students co-construct materials, they are more likely to grasp course concepts and material when compared to their passive-learning counterparts (Vasilchenko et al., 2018, 2020). These studies remark that student-created materials incentivize social learning and ownership over one’s learning. As will be later discussed in this paper, these findings align with our AI design component, where students engage with artifact-based learning and co-construction.

Use of AI Tutors in Higher Education

Integrating AI into educational contexts has led to increased interest in chatbots and other scalable tools to support student learning. From this interest, the CS education community has seen the emergence of AI tutors, a blend of AI and the traditional human tutor, yielding automated guidance and feedback for a variety of situations. This automated, and in many cases instant, feedback has been shown to motivate students that were previously struggling with traditional learning techniques and delivery (Deslauriers et al., 2019). Similarly, other AI-driven chatbots have been employed successfully to aid student comprehension through alternative content delivery means, something especially powerful in programming contexts (Bassner et al., 2024).

Though these AI tutors have been seeing increased adoption and success among the CS education community, they are not without limitations. Due to many AI tutors relying on large language models (LLMs) to produce their flagship instant feedback, these tutors lack personalization. This inability to appeal directly to each student sometimes limits the AI tutor’s ability to effectively improve student engagement and generate performance feedback (Kuhail et al., 2023). Adding to this issue is the fact that students often are hesitant to fully adopt/rely on AI tutors when there are apparent gaps in social norms, usefulness, or even anxiety (Algerafi et al., 2023). These emotional and social barriers remain key considerations in the CS education community surrounding the future of AI tutors and their potential mainstream adoption. What’s clear is that these AI tools allow educators in CS contexts to take powerful steps towards delivering content more actively in large-enrollment settings.

Co-design of Intelligent Conversational Agents

As AI systems have become increasingly involved in education, particularly in the realm of CS, it has become clear that students should have a say in design conversations. Underpinning this notion is past work suggesting that students prefer AI tutors/agents possessing strong clarity, empathy, and responsiveness. Some student groups have even suggested that these preferences go beyond racial and gender groups (Denny et al., 2024). When student preferences are acknowledged in the design phase, it ensures that these AI educational tools are not just competent but conform to student needs as well (Feijóo-García et al., 2023).

Apart from ensuring AI tutors/agents are primed for student-engagement, it’s worth noting that students who build these agents themselves have reportedly benefited academically. When students engage with these agents, they are simultaneously sharpening their understanding of course concepts (Córdova-Esparza et al., 2024). This design phase typically requires students to bridge the abstract course concepts they’re familiar with into teachable pieces, allowing for self-reflection and robust learning scenarios (Zangari et al., 2025). This becomes even more apparent when students are instructed to design their agent for a peer, leading to students sub-consciously mastering course concepts as they anticipate user needs and hone their agent’s content delivery (Bassner et al., 2024). Furthermore, co-designing instructional agents motivates students to work on their communication skills, an approach that pedagogically promotes content knowledge and empathy (Córdova-Esparza et al., 2024; Feijóo-García et al., 2025; Kuhail et al., 2023).

Self-Efficacy and Identity in Computing

Self-efficacy in computing education, defined as a student’s belief in their capability to succeed in computing tasks (Bandura, 1977), plays a critical role in academic persistence, particularly for underrepresented groups. Prior research indicates that when students engage in the active creation of computational artifacts, they often experience a measurable increase in performance self-esteem and domain-specific confidence.

Educational Strategy

AI Use Affirmation

The authors acknowledge the use of Grammarly, ChatGPT 5.5 (tool by OpenAI) and Gemini 3.0 (tool by Google/DeepMind), for spell checking, grammar, style adherence, and editing assistance. Both tools were used to refine the language and enhance the manuscript’s clarity. Apart from Figure 1, no content was generated by AI for this work. Figure 1 was generated with Google’s Nano Banana image generation model.

Strategy Design

This paper presents an educational strategy that took place by the end of the Spring semester (term from January to May) of 2025 at a major Southeastern university in the United States (USA). The strategy was proposed as an individual take-home assessment for four sections of a large introductory undergraduate course in software design and engineering (SWE): Over 700 students were enrolled. The assessment was of individual doing and proposed to help students reinforce their understanding of software design principles and patterns. The strategy consisted of two parts. The first one (i.e., Part A) was completely scaffolded and had them follow and complete an online tutorial to build an intelligent conversational agent (i.e., AI agent) on Dialogflow CX (Google Codelabs, 2025). Dialogflow CX (i.e., customer experience) was used as it allowed students with no prior experience in conversational agent programming and design to still have a good chance of completing the assignment with a learning curve that could be overcame in the lapse of the activity. The tool offered different online resources and communities for students to debug issues and continue working on their assignment and was during the time of the study (Spring of 2025) an industry and academic standard in conversational agent development.

Part A gave them one week to work and was foundational to their technical skills, essential later for the second part. However, this first part did not include anything regarding the theoretical concepts of the course.

The second part of the assessment (i.e., Part B) took place the week after, and asked students to work on an open-ended non-scaffolded design format to build an AI agent to help someone similar to them learn or reinforce the course content. We also encouraged them to consider socio-phonetic cues such as the agent’s voice and accent but can also present themselves as larger speech indicators such as phrasing or tone. Part B was also take-home, and students had a week to work on it. Students were asked to follow this instruction:

  1. “Individually and using Dialogflow CX, create a conversational agent that helps explain the following design patterns: 1) Strategy, 2) Composite, and 3) Observer. Think of it as an assistant intended to answer questions about them (envision this assistant to support a student like you to learn these design patterns).”

To ensure consistency, the leading instructor (corresponding author) and all 57 undergraduate teaching assistants (UTAs) reviewed the student-created AI agents using a standardized grading rubric. Students were also asked to be part of a peer review process after all AI agents were designed and submitted. For this work, grades were not considered or analyzed to protect student anonymity for ethical purposes. Nonetheless, students who voluntarily agreed to respond to the post-assessment questionnaire on their experiences (n = 347) also responded to a set of questions that objectively evaluated their understanding on the topics considered for this assessment. The post-assessment questionnaire was distributed after students had completed and submitted their deliverables for the second part of the assessment and before the peer-review process started.

We report on a sample of 345 students (n = 2 were excluded due to self-reported nonsensical responses) and their experiences, in addition their self-reported perceptions of learning, perceived performance, estimated dedication to both parts of the assessment, and the correctness of their responses to questions about the three design patterns they had to study for the assessment.

Questionnaire

The post-assessment questionnaire had seven parts in total:

  1. The first part presented an informed consent document with three questions:

    1. Voluntary participation agreement.

    2. A close-ended question asking the student to select the class section (out of four) they were part of.

  2. The second part featured one question with seven 5-point Likert items concerning students’ performance self-esteem from the scale designed by Heatherton and Polivy (Heatherton & Polivy, 1991).

  3. The third part featured one question with ten 7-point Likert items from the Ten-Item Personality Inventory (TIPI) (Gosling et al., 2003).

  4. The fourth part referred to the second part (i.e., Part B) of the assessment, in which they had to design an AI agent to support a fellow student study and learn the course content. We evaluated the second part of the assessment first, as it was the most recent one they finished. Five questions were featured here:

    1. A close-ended question on a 7-point semantic scale about the student’s perceptions on their performance on the second part of the assessment.

    2. A close-ended question with six 7-point Likert scales regarding the student’s perceptions of the strategy behind the second part of the assessment and their learning process. These questions adapted the Likert scales from Diemer et al., measuring students’ perceptions of learning (Diemer et al., 2012).

    3. An open-ended question about positive aspects of the second part of the assessment as a strategy: “Tell us what you liked…”

    4. An open-ended question about negative aspects of the second part of the assessment as a strategy: “Tell us what you disliked…”

    5. An open-ended question asking students to estimate hours spent preparing for and completing the second part of the assessment.

  5. The fifth part referred to the first part of the assessment (i.e., Part A), in which they had to complete a scaffolded tutorial to build an AI agent with Dialogflow CX. Five questions were featured here, as it was done with the fourth part of this questionnaire.

  6. The sixth part featured four close-ended multiple-choice questions concerning the three design patterns students had to work on while designing and implementing their AI agents on Dialogflow CX. Responses to these questions did not impact students’ grades and had an option for them to indicate whether they knew or not the answer: We encouraged students to be honest on it. These questions, which are included in Appendix A alongside the assignment rubric and full questionnaire, were designed to evaluate students at the application and analysis levels of Bloom’s Taxonomy.

  7. The seventh part featured 11 questions about students’ age, ethnicity, gender, academic major, student origin (domestic to the United States or international), academic level (undergraduate or graduate), academic year, number of credits for the Spring of 2025, number of extracurricular activities, English proficiency, and frequently used languages.

Figure 1
Figure 1.Strategy Design layout, generated via Google’s Nano Banana image generation model

Data Analysis

Our analysis strategy considers quantitative and qualitative methods to evaluate students’ experiences and engagement toward the educational strategy we report on. We employed a mixed-methods design to evaluate this strategy. Specifically, we utilized a concurrent mixed-methods approach regarding student performance data and survey responses as they were collected concurrently, analyzed independently, and lastly integrated during the investigation’s discussion phase, providing a larger narrative regarding student engagement in the strategy.

Our quantitative analysis considers descriptive and inferential statistics that explore the interplay between students’ close-ended responses and their self-reported demographics. For hypotheses testing, we used non-parametric inferential methods (Fagerland, 2012) as our data came from ordinal scales and was not normally distributed (nonparametric). We also used Pearson’s correlation coefficient (Sedgwick, 2012) to analyze participants’ responses about the performance self-esteem scale and the personality inventory indicated in the Questionnaire section.

For our qualitative analysis, two researchers conducted a thematic analysis of the open-ended survey responses related to students’ likes and dislikes for both Part A and Part B of the assessment. Responses were coded into themes based on affinity, recurring patterns, and insights (Clarke & Braun, 2014). We employed a rigorous consensus-coding approach, where the research team iteratively discussed ambiguous excerpts until full agreement was reached, to leverage the diverse pedagogical expertise of the coders, prioritizing nuanced, collective interpretation over strict statistical agreement. To assess the reliability of the qualitative coding, an independent rater applied each question’s predefined codebook to a subset of 73 responses. Interrater agreement was evaluated using Cohen’s kappa (Landis & Koch, 1977). Agreement ranged from fair to moderate across the four open-ended questions. Specifically, the two questions on what students disliked yielded moderate agreement for both Part A (κ = 0.44) and Part B (κ = 0.44), whereas the two questions on what students liked yielded fair agreement for Part A (κ = 0.24) and Part B (κ = 0.34). These values indicate that the coding process involved a moderate degree of consistency for the “Dislike” responses and a lower, though still above-chance, level of consistency for the “Like” responses.

Student Population

We report on n = 345 students from our four class sections. Our students were between the 18-31 age range: 18-20 years old (n = 310), 21-25 years old (n = 33), and two students (n = 2) older than 26 years of age.

Among the students who voluntarily agreed to participate, almost all of them (n = 328–95.1%) were in computing majors (e.g., computer science, computer engineering, or computational media) – the course is mandatory for students in computing majors and those pursuing a minor. Participants’ reported gender distribution was as male (n = 246), female (n = 92), or other (n = 2). Four participants opted not to identify their gender (n = 4).

Participants were asked to self-identify their ethnicity: Asian (n = 225), Afro/Black American (n = 23), Hispanic/Latin American (n = 12), White American (n = 62) –non-Hispanic/non-Latin American, and Middle Eastern/North African (n = 6). A total of 17 students self-reported to two or more ethnic groups (n = 17).

In our pool of participants (n = 345) we had international (n = 40) and United States (USA) domestic (n = 305) students. We also had representation from predominantly English speakers, with both native (n = 212) and English as a second language (ESL)(n = 133) students.

Finally, most of the students (86.4%) reported being in their first (n = 148) or second (n = 150) academic years, generally taking between 12 to 19 academic credits: 77.1% of participants reported between 12 and 17 academic credits.

Outcomes (RQ2): Insights on Students’ Perceptions and Performance

We analyzed students’ responses to our close-ended questions looking to explore the interplay of these with their demographics. For most demographics, we observed no significant differences (p-value > 0.05) in regard to students’ perceptions of learning for none of the parts of the assessment, nor we did for the objective score we measured from the four close-ended questions proposed in the sixth part of the questionnaire. Significant differences in students’ perceptions of learning were only observed when comparing domestic and international, with international students reporting significantly higher perceived learning for Part A (Mann-Whitney U Test, U(343) = 7515, p<0.05) and Part B (Mann-Whitney U Test, U(343) = 7615, p<0.05).

Significant differences were also observed concerning students’ self-reported time in regard to their academic year (e.g., first-year, second-year) and their English fluency (English Learners vs. Native English Speakers).

English Learners spent significant more time on Part A compared to Native English Speakers (Mann-Whitney U Test, U(343) = 11838.5, p<0.05). However, no significant differences were observed regarding Part B and students’ self-reported dedication (Mann-Whitney U Test, U(343) = 11838.5, p>0.05). We also observed significant differences when analyzing students’ dedication on Part B regarding their academic year (Kruskal-Wallis Test, H(3) = 9.87, p<0.05). Pairwise comparisons using Dunn’s test found that first-year students significantly differed from third-year students (p<0.05): Third-year students reported significantly higher times than their peers. These findings suggest that students’ overall dedication could be related to the structure of the assignment (scaffolded vs. open-ended), as well as the academic maturity and commitment of the student: The course is a mandatory course proposed for second-year students.

We observed significant positive correlations between students’ perceived learning for both parts of the assessment, their perceived performance, and their objective learning scores from the close-ended multiple-choice questions they had to respond to. To analyze these correlations, we corrected our significant value using Bonferroni’s correction for multiple comparisons (i.e., six comparisons per item): α = 0.008 (Bland & Altman, 1995).

  • Objective Score vs. Part A: Perceived Learning (r = 0.161, p = 0.003).

  • Objective Score vs. Part B: Perceived Learning (r = 0.315, p <0.001).

  • Objective Score vs. Perceived Performance (r = 0.226, p <0.001)

It’s important to note that these correlation values, spanning r = 0.161 to r = 0.315, are modest in nature. Their interpretation moving forward is done carefully so as not to exaggerate these findings.

These observations suggest that students’ perceived learning positively align with factual understanding of course content. Future iterations will navigate further the relationship between engagement and learning when using student-created AI agents in CS classrooms.

We also observed significant correlations with students’ personality, performance self-esteem scores, and their perceptions of learning. First, we found a positive significant correlation when comparing performance self-esteem scores and their extroversion scores (r = 0.262, p < 0.001). We also found a significant positive correlation when comparing students’ performance self-esteem scores and their objective scores from the close-ended multiple-choice questions they responded to (r = 0.232, p < 0.001). Nonetheless, no significant correlations were observed between their responses and extroversion scores (r = -0.05, p = 0.356).

Figure 2
Figure 2.Performance Self-Esteem Scores Per Gender

Finally, we observed a significant difference when exploring students’ performance self-esteem about their genders (Kruskal-Wallis, H(2) = 9.87, p<.005) – we grouped non-binary participants and those who did not disclose their gender as “Other (n = 6).” Pairwise comparisons using Dunn’s test found that female students reported significant lower performance self-esteem scores than their male peers (p<.005) – see Figure 2. Future iterations will navigate the role of strategies centered on student-created artifacts and their effects on student self-efficacy, especially concerning female and other underrepresented groups.

Outcomes (RQ1) : Insights on Students’ Engagement and Preferences

For all reported thematic frequencies, the unit of analysis is the unique student; percentages represent the proportion of the total participating sample (n = 345) whose responses aligned with a given theme. We analyzed students’ preferences about our two-part assessment and the idea of working on student-created AI agents with Dialogflow CX. We found that students shared common interests across both parts of the assessment, but they also expressed different preferences for how they liked to engage with the material. We report on several insights on the benefits of our strategy, as well as on important challenges and concerns we identified.

Regarding the first part of the assessment (i.e., Part A), students generally expressed positive experiences about the Tutorial (22.0%, n = 76). That is that many students appreciated having clear, step-by-step instructions and scaffolds. This structured approach could have made the assignment easier to understand, especially for those students who were new to dialogue systems’ design and development. The second most frequent positive theme was Learning Dialogflow (16.5%, n = 57), showing that students enjoyed gaining skills with this specific platform. Learning Chatbot (13.6%, n = 47) was also frequent, confirming that the topic itself was interesting for many students:

  • “It was interesting to see a way to create chatbots. I liked it when everything was able to work the way I intended to after figuring out all the bugs and errors.” [R_155]

Nonetheless, the most common negative theme for Part A was also Tutorial (30.0%, n = 103). This suggests that although some students appreciated the guidance, others could have struggled with the tutorial or did not find it helpful. The second top negative theme was Time (20.1%, n = 69), indicating that many students felt that Part A took too long or that their time was not used efficiently. Finally, Irrelevant (13.1%, n = 45) appeared as a notable theme, showing that some students did not see the clear connection between this part of the assignment and the course objectives:

  • “It was really long, not relevant to the course, and repetitive.” [R_102]

For the second part of the assessment (i.e., Part B), students mainly valued the Open-Ended Nature of the assignment (22.3%, n = 77). This shows that students enjoyed having more freedom to design their own solutions and be creative. The second most common positive theme was Learning Chatbot (17.7%, n = 61), again confirming students’ strong interest in working with conversational AI. Even though some students shared Negative Feedback (9.3%, n = 32), most comments were generally positive and reflected that students liked applying what they learned in a practical way:

  • “I like that we were given freedom to design our own chatbot. The assignment was fairly straightforward, but the fact that we had relatively few limitations made the assignment kind of fun.” [R_231]

At the same time, Part B also had some challenges we observed. The most common negative theme was Unclear Instructions (23.7%, n = 82), suggesting that many students struggled to understand what was expected. The second theme was Irrelevant (17.6%, n = 61), meaning that some students questioned how the assignment related to the rest of the course. Tool Difficulty (12.4%, n = 43) also appeared, indicating that students had trouble using Dialogflow CX to complete their work:

  • “I thought making the chatbot overall did not help me much. I do not see myself using DialogFlowCX in the future much.” [R_244]

Despite these challenges, our findings suggest that students were highly engaged with the idea of designing their conversational AI agents and found real value in both learning the core concepts of the course and gaining hands-on experience as developers. Many students appreciated the combination of structured guidance (Part A) and opportunities for creative exploration (Part B). The consistent student interest in conversational AI agent development (i.e., Learning Chatbot) across regarding both parts of the assessment highlights that this topic can motivate CS students to learn and apply new skills in meaningful ways.

Discussion

This work serves as a valuable sample to examine the benefits and challenges of promoting the design and implementation of conversational intelligent agents (i.e., AI agents) as a strategy to promote engagement and learning in CS education. Our findings suggest that, although students identified and reported challenges and limitation in regard to the assignment format and its relevance to the course content (Software Design and Engineering–SWE), they also valued many aspects of the strategy and how it supported their learning processes.

On the positive end, students highlighted multiple benefits of the opportunity of designing a conversational AI agent to help a peer learn about SWE and its design patterns. Many appreciated the novelty of the assessment and how it enabled them to explore the design and engines of technologies they have used and interacted with (i.e., chatbots, dialogue agents) and did not know how they worked. Students predominantly valued the open-ended nature of the second assessment, a finding that directly corroborates Vasilchenko et al.‘s (2018) assertion that the self-flipped classroom model increases student ownership and confidence by empowering them to co-construct their own learning materials. Students also expressed that they enjoyed having agency on how their agents were going to be, appreciating the freedom the strategy gave them to design their own solutions and be creative. Overall, we observed that students were strongly interested in working with conversational AI, as the most common theme from the students’ experiences for both parts of the assessment, related to their benefits in working with AI and learning through the strategy.

Our analysis revealed a positive significant correlation between students’ performance self-esteem scores and their extroversion scores. This suggests that more extroverted students may naturally feel more confident in engaging with open-ended, conversational AI design. Consequently, instructors in large software engineering courses must be mindful of introverted students or those from underrepresented backgrounds, ensuring that open-ended design tasks include sufficient scaffolding to equitably build performance self-esteem across all personality types.

Nevertheless, some concerns were raised about the time students had to spend in the assessment and how relevant a tool such as Dialogflow CX could be for their professional development. Future iterations may consider navigating other course formats (e.g., large or small) and exploring the inclusion of strategy like ours in previous and more senior CS courses. We may also explore the relationship and effects of strategies like ours on student self-efficacy, perceived learning, and objective performance, looking to tailor a strategy that is not only engaging but also effective to help students learn CS concepts.

Overall, our findings highlight the potential of student-created AI agents to foster meaningful learning experiences in CS education. Still, careful consideration must be given to students’ situational factors. Future adaptations of this strategy should consider students’ time, commitment, and other factors. Active learning through student-created artifacts has the potential to be a powerful tool to support student agency and reflection, but as with any innovation within learning environments, must be thoughtfully designed to meet the diverse needs we may find in our CS classrooms.

Limitations

This paper has several limitations that should be considered when addressing its findings, results, and insights.

Regarding the use of the “objective learning score” in the student population subsection, the authors acknowledge that this measure is exploratory in nature regarding the measurement of student learning. Future studies should focus on using more structured instruments already trusted by the community to more effectively assess student learning and to communicate the findings in a more familiar way. Also, in relation to structured instruments, the authors also acknowledge that the student-created AI artifacts were not assessed for quality, accuracy, nor effectiveness in this study. Though this is an important consideration, the authors maintain that no anomalous submissions or conditions were observed while conducting this study. From a majority standpoint, the student submissions for their student-created AI artifacts were done thoughtfully and satisfactorily responding to the assignment description. Future investigations are planned to fully connect artifact outcomes with student experiences to help bridge this gap.

In the study’s communication of results, the authors also acknowledge that positive correlations between perceived learning and objective scores could potentially reflect students with increased abilities instead of highlighting the strategy’s effectiveness. This alternative explanation to the findings is one very worthy of further investigation. In future studies, this direction will be a primary focus to help determine the linkage between students with more advanced abilities compared to the effectiveness of the student-created AI artifact strategy.

Furthermore, to more closely allow for an organic shaping of the study’s offering in the lecture course, this study was designed to follow the pedagogical structure of the course, yielding methodology that did not include a control/comparison condition. Moreover, the volunteer sampling conducted for this study could have presented student self-selection bias, as the paper only reports findings and results from students who voluntarily decided to participate and successfully completed the study tasks (n = 345)—not the whole course population (N = 710). Although we report on a substantial sample (n = 345), future work must explicitly compare the participating sample’s demographics and prior academic performance against the full course population (N>710) to accurately quantify and mitigate potential self-selection bias. The authors acknowledge that this lack of control and intervention groups does sacrifice the study’s internal validity. Still, the overall experience reported in this paper adapts to the natural setup of the course assessed, with students reporting their perceptions in an organic learning environment. Please note that this study serves to promote an attractive area for further research in the use of student-created AI artifacts in STEM and computing education. Future iterations shall focus on employing proper experimental setups to more effectively address comparisons and extend this paper’s findings and results. This, with the goal of understanding the use of student-created AI artifacts in STEM and computing education, is important for future study focuses.

Finally, the interrater reliability ranged from fair to moderate across our four coding schemes, suggesting that some categories required subjective interpretation, particularly for the questions on what students liked of each part of the strategy. Accordingly, findings derived from these qualitative codes should be interpreted with appropriate caution.

Conclusions

This work explored the potential of integrating student-created intelligent agents into computing education as a way to promote engagement, creativity, and deeper learning. While some students shared concerns about the time required and questioned how relevant tools like Dialogflow CX might be for their professional development, many also valued the opportunity to better understand the technologies behind conversational agents and to apply design concepts in a hands-on way. The freedom to create their own solutions appeared to foster a sense of ownership and agency, showing that active learning strategies have the potential to help students connect course content to real-world contexts. Due to the exploratory nature of this study and the absence of a control group, these findings should be interpreted conservatively as preliminary and exploratory evidence of student engagement rather than causal proof of pedagogical superiority.

Moving forward, instructors considering similar approaches should aim to balance the novelty of these activities with clear connections to course objectives and professional skills, and provide enough support to address potential challenges. Future work will explore how strategies like this impact self-efficacy, perceived competence, and actual performance across different course formats and levels. Overall, our exploration points to potential value in integrating AI design experiences into CS education if they are thoughtfully planned to align with students’ needs, time constraints, and expectations. Ultimately, the implications of this work suggest that educators can leverage student-created AI as a dual-purpose tool, one that simultaneously provides scalable assessment mechanisms and fosters equitable, student-centered learning environments in large computer science cohorts.


Acknowledgements

This material is based upon work supported by the National Science Foundation (NSF) under Grant No. 2434428. Any opinions, findings, or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the NSF.

The authors would like to thank Mr. Jonathan E. Bruce, Mr. Jhonathan Sora-Cardenas, Mr. Christopher G. Leonard, and Mr. Gaurang Kamat for their assistance in assessing the interrater reliability of the four coding schemes used in this study.

The authors acknowledge the use of Grammarly, ChatGPT 5.5 (tool by OpenAI) and Gemini 3.0 (tool by Google/DeepMind), for spell checking, grammar, style adherence, and editing assistance. Both tools were used to refine the language and enhance the manuscript’s clarity. Apart from Figure 1, no content was generated by AI for this work. Figure 1 was generated with Google’s Nano Banana image generation model.