Introduction
Conducting a comprehensive literature review is a foundational step in the research process, serving several critical functions: synthesizing existing knowledge, identifying gaps and directions for future research, and supporting interdisciplinary inquiry (Ahern et al., 2019). In recent years, advances in artificial intelligence (AI) have added a new dimension to this process. Tools such as Elicit, Inciteful, SciSpace, Perplexity, Voyant Tools, DeepSeek, DALL·E, and ChatGPT have the potential to streamline literature reviews, enhance efficiency, and deepen the quality of analysis for both faculty and emerging scholars (Khalifa & Albadawy, 2024). Currently, the majority of students are incorporating generative AI into their academic work, Forbes (2025) reported that 90% of college students use AI tools for academic purposes, with nearly three-quarters noting an increase in their use over time. However, clear ethical, step-by-step guidelines for integrating AI into academic research remain largely absent, and many faculty members are uncertain about how to leverage these tools effectively to scaffold student learning (Saylam et al., 2023).
Our study explores the potential of AI-assisted tools to support the organization of literature reviews, particularly in academic settings. It proposes a structured step -by -step framework for integrating AI into the review process, emphasizing not only efficiency but also the need for critical engagement with AI-generated outputs and proposes to mitigate the previously mentioned gap. Simply organizing information using generative AI is not sufficient; students must also be able to critically evaluate the accuracy, credibility, and relevance of AI-generated content. Artificial intelligence (AI) is reshaping educational practice by enabling more efficient, data-informed, and individualized approaches to learning and research. AI tools provide unique opportunities to scaffold literature review, source evaluation, and synthesis by reducing manual workload and enhancing student engagement. These technologies can broaden access to scholarly inquiry—particularly for students at resource-constrained institutions—by simplifying information retrieval and structuring research pathways. However, to leverage this potential responsibly, it is essential for faculty members to mentor and teach students to critically assessing AI outputs, recognizing inherent biases, and adhering to ethical and disciplinary standards. When carefully embraced, AI can augment the intellectual rigor and reflective judgment cultivated through human mentorship.
For novice researchers especially undergraduate students’ literature reviews present both opportunities and challenges. Also, faculty at research undergraduate institutions (RUIs) often face difficulties due to limited research experience and varying interests. Literature reviews can serve as an entry point, identifying mutual interests, offering meaningful engagement with research while helping students build foundational analytical skills. AI can play a supportive role by simplifying labor-intensive tasks like retrieval, generating images, figures, tables, research connection maps and synthesis, thereby allowing students to focus on critical thinking and conceptual understanding. AI tools can support undergraduate engineering research training, assisting with tasks such as data analysis, design exploration, and literature review.
Recent research has increasingly examined the pedagogical implications of generative AI in higher education, particularly in relation to academic writing, epistemic risk, and verification practices (Cotton et al., 2024; Kasneci et al., 2023; Rudolph et al., 2023). Emerging research suggests that AI can serve as a scaffold for early-stage ideation and structural organization while simultaneously introducing risks such as hallucination, bias, and over-reliance. Building on this work, the present study offers a process-based demonstration of AI integration within undergraduate literature review development. Rather than evaluating learning outcomes, it contributes an applied instructional framework grounded in observed tool behavior. While early findings suggest a potential paradigm shift in how reviews are conducted, the authors underscore that human expertise, contextual insight, and ethical reasoning remain indispensable to scholarly research. As such, the integration of AI must be approached as a complement, not a replacement for critical human judgment (Zhang & Aslan, 2021). In this study, the authors investigate the usage of various AI tools in supporting literature reviews (with an example of perovskite solar cell {PSC)} and assess their strengths and limitations. Specifically, the authors propose data management, summarization and data visualization via AI. The research question asked for this study are - How can generative AI tools be integrated ethically and effectively to support novice researchers in the literature review process, and what guidelines should faculty follow? While the examples presented involve general PSC research contexts, the strategies are directly applicable to engineering education settings, including guiding students in research methods, supporting project-based learning, and mentoring undergraduates in engineering research design. The proposed framework aims to support literature review development through AI-assisted workflows. Also, this work will provide mentor professors/Project Investigators (PI) a practical and tested guide on mentoring students to use AI.
Background
The increasing volume of scholarly publications in this new era of research has posed a major challenge for researchers seeking to conduct comprehensive literature reviews. Also, teaching students poses another level of challenges for faculty members. The traditional literature review process involves systematically identifying, evaluating, and synthesizing existing scholarly work on a specific topic to establish a foundation for new research manually (Arslan et al., 2025). It begins with defining a clear research question or a domain, followed by searching academic databases and sources to gather relevant peer-reviewed articles, books, and reports. Researchers then critically analyze the selected works to understand trends, methodologies, theoretical frameworks, and gaps in the literature.
Artificial Intelligence (AI) has emerged as a powerful tool for navigating this complexity by automating several phases of the review process, especially with searching and organizing. Recent studies highlight how AI-based platforms such as Elicit, Research Rabbit, and SciSpace assist in identifying relevant literature through semantic search, clustering research themes, and generating concise article summaries (Aljohani & Cristea, 2023; Wang et al., 2022). These tools reduce the cognitive and manual workload traditionally associated with literature reviews, particularly benefiting novice researchers and interdisciplinary scholars who face barriers in accessing and organizing information across domains. Furthermore, ChatGPT and similar large language model (LLMs) based generative AI are increasingly used to scaffold early stages of review writing by outlining key themes or suggesting structure. While the potential of AI for enhancing productivity and scalability is significant, scholars caution against over-reliance on these tools without human validation (Choudhury et al., 2023). Concerns around biases in training data, lack of citation accuracy, and ethical implications remain at the forefront and there is gap focused on those conversations in the academic context. As such, researchers emphasize a hybrid approach where AI aids information processing while human judgment ensures academic rigor and contextual interpretation. Also, AI can generate figures, chart and infographics for the researchers, often AI tools can suggest further research references and map research connections through graphics as emerged from our work. A study by Zha et al. (2025) highlighted a close connection between students’ AI learning and their improved content knowledge which is worth exploring further. Some studies, such as Wagner et al. (2022) and Bolaños (2024), have examined the role of AI in the literature review process; however, they did not focus on providing step-by-step guidance or using prompts across diverse AI tools like us.
While a growing body of literature suggests that, when used responsibly, AI can open up and diversify access to research and foster deeper engagement in the scholarly process; discussions around critical thinking, ethics, and real-life hands-on example with different generative AI tools along with legal considerations remain limited within the academic context. This research aims to mitigate the existing gap.
Methodology
This study was informed by design-based research (DBR) principles rather than constituting a full design-based research study. Specifically, the authors adopted key DBR features, including problem-centered design, iterative refinement, and evaluation of instructional artifacts (Tinoca et al., 2022) to develop our step-by-step process for instructional practices. The authors used iterative design, evaluation, and refinement processes to examine AI-generated artifacts and generate practical instructional guidance for faculty and novice researchers. The purpose of this work was not to evaluate learner outcomes through repeated implementation cycles but rather to develop and document a practical framework for integrating generative AI into the literature review process for novice researchers.
Output (citation maps, summaries, mind maps, and references) were treated as qualitative artifacts.
These artifacts were systematically examined using qualitative content analysis to evaluate patterns of accuracy, coherence, completeness, and reliability across tools.
The PSC topic was selected due to the authors’ domain familiarity, which allowed for informed verification of AI outputs and identification of inaccuracies.
To strengthen methodological rigor the study incorporated iterative cycles across all the steps. For example, the figures or outputs included in the manuscripts were generated via multiple attempts, the authors have gone through several rounds of iterations to create the best suitable outcomes based on our own expertise. However, this paper only focused on the use of artificial intelligence to support and simplify engagement with academic literature
-
Exploratory Phase – Initial prompting of selected AI tools to generate literature searches, citation maps, summaries, paraphrases, and visualizations.
-
Evaluation Phase – Systematic review of outputs using predefined evaluative criteria (see Table), identifying hallucinations, inaccuracies, structural coherence, and alignment with verified sources done by a human observer (a student and sometimes by the faculty)
-
Refinement Phase – Revised prompting and cross-tool comparison to assess output variability, stability, and responsiveness to modified instructions.
Rather than relying solely on observational impressions, the analysis employed a structured evaluative framework across five dimensions which were verified by human observer:
-
Accuracy (factual correctness and verifiable citations)
-
Completeness (coverage of major themes in the PSC domain)
-
Transparency (traceability of sources and logic)
-
Coherence (structural clarity and logical flow)
-
Reliability (consistency across repeated prompts)
Prompts were repeated for selected tasks to evaluate output variability. While generative models are non-deterministic, limited replication enabled observation of consistency patterns and hallucination tendencies. In this work the authors are not presenting best ways to prompt generative AI for optimal results, rather the goal is to introduce methods to the new researchers to use gen AI in an ethical way.
The goal was to lower the entry barrier to engaging with academic literature.
Although this study adopts a qualitative, design-oriented approach, its purpose is illustrative for instructional purposes rather than empirical validation. The manuscript does not claim systematic outcome measurement; rather, it documents and analyzes AI-supported research workflows as they were observed in practice through the seven steps process. Each AI interaction was examined for accuracy and integrity. Patterns of strengths and weaknesses were identified across tools through repeated observation and comparative review.
The authors used perovskite solar cells (PSCs) as a sample topic (Gong et al., 2026) to construct the framework for the AI-assisted literature review process, however, the framework is generalized and can be used in any discipline. PSCs are a class of next-generation solar cells that use perovskite-structured materials as the light-absorbing layer. The interdisciplinary nature of PSC research demands an integrated approach, combining expertise from chemistry, materials science, device engineering, and data science. The second author has extensive experience in conducting high-impact literature review in the field on thin-film solar cells (Gong et al., 2012, 2017). Projects involving multiple students, each focusing on different aspects of PSC research—such as manufacturing techniques, space applications, simulation models, or data-driven methods—organizing these inputs into a unified review can be a daunting task. Tasks such as manually tracking the evolution of PSC efficiencies (Green et al., 2025) can be time-consuming but are ideal for undergraduates to explore using AI tools.
Seven Step Process
Conducting a literature review involves several key steps, as depicted in Fig. 1 (adapted from Dundalk Institute of Technology, 2025), to ensure a comprehensive and high-quality outcome. The first step involves defining the scope and purpose by clearly identifying the research question, focus area, and boundaries, such as time frames and topics. The next step is searching for relevant literature using academic databases like Google Scholar, Web of Science, or Scopus, employing keywords and Boolean operators to filter results effectively. Once sources are identified, evaluate and select them based on relevance, credibility, and quality, prioritizing peer-reviewed papers and reputable journals. Then, organize and categorize the literature by grouping studies with similar themes, methodologies, or theoretical approaches, using an outline to visualize connections. Next, analyze and synthesize the information by comparing findings, identifying trends, gaps, or inconsistencies, and determining their relevance to the research question. Finally, revise and edit the review for coherence, clarity, and logical flow. Ensure proper citation and formatting and seek feedback from peers or mentors to refine the final product. A well-conducted literature review not only organizes past research but also highlights contradictions, unanswered questions, and emerging themes, ultimately shaping the direction and significance of the new research endeavor. In traditional literature review all those steps are done manually. Also, all steps were verified by our human researcher teams and if student researcher had any concerns, the second author, the domain expert always oversaw those steps. So, the step-by-step process are documented based on analyzing the generated outputs
However, generative AI has revolutionized the literature review process by automating time-consuming tasks such as article discovery, thematic clustering, and citation organization, significantly enhancing research efficiency. Here the authors developed a step-by-step guideline on how AI technology can be embraced in each step, yet the critical thinking and higher order skills remains irreplaceable. Each step of the literature review process is presented by first describing the conventional manual workflow, followed by an AI-assisted alternative where applicable, including the prompt used and a critical examination of the resulting output after every step. The analytic component focuses on evaluating AI performance in terms of accuracy, completeness, citation reliability, and instructional usefulness based on experiences and observations of the authors rather than deriving emergent themes from qualitative data.
1: Identify the Topic (Define Research Scope)
The purpose of this step is to identify the domain areas via AI tools. The questions or a targeted domain explored in a literature review depend significantly on its purpose. For a review within a thesis or technical report, the primary focus is on relevance to the specific project. In contrast, a review intended for publication emphasizes making a meaningful contribution to the broader scientific field. The authors suggest researchers to explore different mind mapping tools to start with. For example, AI-powered mind mapping tools can help with brainstorming, organizing ideas, and structuring information. These tools use artificial intelligence to automate connections, generate topic suggestions, and enhance visualization. The authors suggest taking an article or multiple manuscripts on the selected domain and generate mind map or word cloud to first explore the important concepts related to the topic and those can guide researchers identify related domain which could be focused on the actual literature review writing. Mindmap can also potentially help to identify the gaps in the concept.
For concept exploration and literature review development, authors focused on the topic of perovskite solar cells (PSCs). Given these complexities, authors realized a visual to automate connections within a research area would be helpful for beginners. A mind map (Figure 2) was created using an AI tool (Map This) to provide a structured visual framework for organizing key topics in perovskite solar cell (PSC) research, including stability, machine learning, space applications, and manufacturing based on our manuscript.
The authors used the following prompt
“Create a mind map of review of perovskite solar cell including stability, machine learning, applications and manufacturing”. The authors created this mind map using our own work in progress manuscript. As shown in Figure 2, the mind map illustrates how these topics and their subtopics interconnect, such as how stability concerns influence space applications or how machine learning can optimize manufacturing processes to improve both efficiency and stability. The mind map can be created based on one manuscript or multiple manuscripts that are publicly available and already published or using a text that students are currently working on. The mind map also aids in identifying research gaps, such as the lack of specific machine learning algorithms tailored to PSC design or the unique challenges faced by PSCs in extraterrestrial environments, including radiation and temperature extremes. Additionally, the hierarchical structure of the mind map ensures that the articles can be logically divided into sections and subsections. For instance, it allows for a detailed discussion on experimental stability studies, followed by an exploration of environmental stressors such as oxygen, moisture, and UV light. This structured approach not only enhances the writing process but also facilitates collaborative work among multiple contributors, ensuring consistency and clarity across sections.
As an alternative approach, a word cloud (Fig. 3) can be generated using Pro Word Cloud (embedded for free within the word) to visualize the key topics of the draft paper. The authors just prompted to create a word cloud by choosing insert and then insert add ons and then finally choosing the Pro word cloud. For this case, the word cloud accurately highlighted several keywords, such as “PSC,” “stability,” “PCE,” “phase,” and “challenges” in large fonts. However, other key areas such as “space applications,” “manufacturing,” and “device modeling” were less prominent. The word cloud included non-technical words such as ‘can,’ ‘one,’ and ‘type.’ Since the word cloud tool allows filtering some words as well before the cloud is generated, researchers can utilize this feature or manually omit non-technical words, focusing only on relevant technical terms. For our study, the authors decided not to consider non-technical words after the cloud was generated. Therefore, researchers should exercise caution when analyzing the word cloud to ensure that meaningful terms are emphasized while filtering out irrelevant ones. The authors suggest that researchers be a little mindful of the words and omit the irrelevant words from consideration. While students create the mind map, the authors suggest they use the key constructs related to the topic to generate one. Since we were already researching the PSC, the authors had the manuscript ready and used it as the input to create this map.
Also, the authors used DALLE.AI and asked to “create an structured visual framework for organizing key topics in perovskite solar cell” and that generated a visual.
While doing this step, the authors noted those three AI tools created successful outputs, there is a risk that AI tools can generate inaccurate or non-existent information, often referred to as hallucination. AI tools rely on pattern recognition and predictive modeling rather than verified sources, which can lead to the inclusion of misleading concepts and incorrect relationships. To mitigate this risk, students should cross-check AI-generated content with credible academic sources and domain experts before incorporating it into their research. For this case, the mind map generated via Map this, and word cloud were pretty accurate. Please refer to Table 1 to cross-check and verify information generated by generative AI tools. After each step the authors provided guidance on how to critically evaluate AI generated information.
Step 2: Review Discipline Style
The traditional method of literature search and style identification starts with coming up with the keywords for searching research articles in databases like Web of Science to perform keyword-driven searches and identify relevant publications and finally identifying journals.
This approach was supplemented by domain expertise, as second author’s prior experience publishing Advanced Sustainable Systems provided valuable insights into the journal’s thematic focus and formatting requirements. By systematically reviewing references, abstracts, and full-text articles, a robust foundation for identifying key works and trends within the field is built. This method, while thorough, is time-intensive and heavily dependent on the researcher’s ability to manually synthesize and interpret large volumes of information.
There is no current academic AI search that can locate journals meeting our specific needs. However, the big academic publishers have their own journal finder tools. Links, along with our search results, are presented in Table 2. The authors asked the generative AI to create a list of journals that accept literature review articles focused on PSC along with impact factors and publishing companies. ChatGPT provided the initial list of suggested journals, after which our team of undergraduate researchers manually reviewed each journal to fact check. They verified the impact factor, assessed whether the journal falls within the scope of PSC, and checked if it had previously published any work related to PSC (which the authos did not ask ChatGPT) and finally created this table based on their manual search. The first column ChatGPT generated represented the publisher groups. The second column lists the journals suggested by ChatGPT. The third column provides the impact factors for each journal and the final column indicates the number of manuscripts has been published on each journal on PSC, and our undergraduate researchers created this particular column. The authors noticed that the ChatGPT was pretty accurate in terms of suggesting the list and the impact factor. However, the authors also noticed that ChatGPT suggested journals like Advances in Materials and Processing Technologies or Nature Communications which have never published review on PSC yet fall within the scope of PSC related work. Finally, this table was created by our students. The authors chose ChatGPT because the tool known for writing efficiency
Therefore, the authors suggest that while generative AI tools can assist researchers by generating a list of potential journals for submission, these lists may sometimes include journals that do not traditionally focus on the topic of interest. As a result, the authors recommend that researchers manually verify whether each suggested journal aligns with the scope of their topic and published that topic in the past, as our team has done. Please refer to Table 1 for verifying AI generated content.
Table 2. Search results of Web of Science using key word “perovskite solar cells” and document type: “review” -AI generated table after our manual search. (Author’s own work)
Elsevier: https://journalfinder.elsevier.com/
Springer: https://journalsuggester.springer.com/
Taylor & Francis: https://authorservices.taylorandfrancis.com/publishing-your-research/choosing-a-journal/journal-suggester/
Wiley: https://journalfinder.wiley.com/search?type=match
Step 3: Search the Literature Using Literature Map (Litmaps)
The traditional literature search process—relying on initial keyword searches and then manually identifying related articles—often lacks visual or conceptual connection maps, making it difficult to assess relationships between sources and track the evolution of ideas.
Therefore, the authors propose the relevant literature via AI tools. Students can leverage AI-based citation mapping tools that illustrate how different studies and ideas are interconnected. For example, the authors uploaded the citation of Kanaya et al (2019) to the Litmaps and the AI tool generated the citation map. The Litmaps database uses articles with available Open Access metadata to search on and generate suggestions based on the availability, to ensure the best coverage for the search. In Litmap the user can search the citation using DOI or keywords and a map is generated- if the full article is open access Litmap likely provides more thorough citation maps. It helped us to understand how one paper is related to the others. The authors did not have to use any prompts as it works auto. The size of the bubble indicated the impact and number of the citation. Also, the map indicates the relevance of all the research articles. Also, it provides students insights into which paper students should read next in order to explore the concept of PSC. This can be replicated with any journal paper/conference paper across disciplines for an efficient literature review. Also, the author want to make a point to students not to upload full paper due to copyright law, instead only referring to the publicly available data like abstract or references. The authors used Litmap because that was the only AI tool that could generate a citation map at that time.
The citation map in Figure 4(a) illustrates the interconnections among key research articles focusing on perovskite solar cell performance in space environments. Notably, Yoon et. Al (2024) lacks a direct citation connection to Angmo (2024), but both share references to foundational works by Kirmani (2022), Ho-baillie (2022), and Tumen-Ulzii (2020). This indirect linkage emphasizes shared thematic threads related to radiation-hardness experiments, deployment opportunities, and degradation mechanisms of perovskite solar cells in extreme conditions. The most influential work in the map is Kanaya (2019), as shown in Fig. 4(b), which serves as a foundational reference for five subsequent studies, including Verduci (2022), Ho-baillie (2022), Luo (2021), Tang (2023), and Angmo (2024). These studies build upon Kanaya’s investigation into proton irradiation tolerance and the protective measures necessary for perovskite absorbers in space applications. Collectively, these citations underline a consistent focus on enhancing device stability, recoverability, and performance under proton irradiation, demonstrating Kanaya’s significant impact on the field.
Also, elicit was used and prompted “Find papers in “perovskite solar cells in UV test” and filter the paper in the most recent 2 years. Now I have two papers at least in a single paragraph”.
Students can utilize various AI-powered mind mapping tools to create citation or concept connection maps that visually represent relationships among articles and ideas. By selecting different domains within the same overarching concept, students can use these tools to discover related research, uncover hidden connections, and identify emerging themes—enhancing both the depth and breadth of their literature review. Once again, the authors emphasize on consulting Table 1 for verifying any information generated by AI tools.
Step 4: Manage Your References
Efficient reference management is a critical component of academic research, and AI-powered tools have greatly accelerated this process. Platforms like Zotero, EndNote, and Mendeley, as well as newer AI-enhanced tools such as Research Rabbit and Connected Papers, allow users to automatically organize citations, generate bibliographies in various formats, and even discover related literature. However, ensuring the accuracy of these automatically generated references remains a key challenge.
Also, some practice guidelines were generated for managing references with AI tools-
Verify Sources: Always prioritize first-hand literature search to verify over choosing AI generated summaries to maintain the integrity and accuracy of research. Generative AI can often hallucinate information.
Review Citations: Double-check automatically generated citations and their content for accuracy before submitting any work.
Many AI tools can be leveraged to manage references. for example, Perplexity, Elicit and Zotero/endnote and all those AI powered search engine for academic discoveries. For Perplexity users should verify publication dates. Perplexity also carries the risk of relying on second-hand sources than original work. Zotero is a reference management software that facilitates literature organization, citation management, and PDF storage. It allows users to search for literature using metadata such as DOI, ISBN, or titles and supports automatic metadata extraction when saving articles from databases or journal websites. Even though Zotero does not have any built-in artificial intelligence (AI) functionality, that said, a growing number of third-party plugins within Zotero and integrations have emerged to extend Zotero’s capabilities—particularly for interacting with and summarizing PDFs using AI tools. These plugins connect Zotero with AI-powered applications.
Automatic citation generation may sometimes produce errors; manual verification is advised (Please refer to Table 1). Recently the authors noticed Perplexity and Elicit producing more academic and correct citation.
When we prompted ChatGPT can generate a reference list. For example, the prompt was used as
“Create a reference list in IEEE style with important articles on perovskite solar cell (PSC)”
ChatGPT responded -Here’s a reference list in IEEE style featuring some notable articles on perovskite solar cells (PSCs): However, upon fact checking the authors discovered that most of the articles are hallucinated and incorrectly cited. Therefore, the authors clearly observed inaccuracies and lack of consistencies while creating a reference list. However, the authors also agree with the new model and “Deep Thinking” function, this problem can be mitigated significantly.
Additionally, when AI tools are used to generate or process documents, questions of intellectual property arise—particularly concerning ownership and copyright of AI-generated content, that’s why it is critical to avoid uploading full document to any generative AI and the authors suggest uploading only publicly available information like abstract and references. To ensure responsible use of generative AI, all AI-assisted activities in this study were limited to publicly accessible materials, including open-access articles, author-generated drafts, bibliographic metadata, and abstract-level information. No confidential student data, unpublished datasets, or proprietary research documents were uploaded, and the manuscript emphasizes that instructors should similarly restrict AI interactions to publicly available content while avoiding copyrighted full-text materials.
Journal Article: (ChatGPT generated)
Green, M., Emery, A., Lee, D. H., & Warta, W. (2016). Solar cell efficiency tables (version 44). Progress in Photovoltaics: Research and Applications, 24(1), 3–12. https://doi.org/10.1002/pip.2728
The Correct citation should be- Green, M. A., Emery, K., Hishikawa, Y., Warta, W., & Dunlop, E. D. (2014). Solar cell efficiency tables (version 44). Progress in Photovoltaics: Research and Applications, 22(7), 701–710. https://doi.org/10.1002/pip.2525
Review Article: (ChatGPT generated)
Snaith, H. J. (2013). Perovskite solar cells: An emerging technology. Journal of Physical Chemistry Letters, 4(21), 3623–3630. https://doi.org/10.1021/jz4020162
The Correct Citation- Snaith, H. J. (2013). Perovskites: the emergence of a new era for low-cost, high-efficiency solar cells. The journal of physical chemistry letters, 4(21), 3623-3630**. https://doi.org/10.1021/jz4020162**
Research Article:
Choi, S. K. R. M. K., Ho-Baillie, A. W. Y., & Green, M. A. (2017). Stability of perovskite solar cells. Nature Energy, 2, Article 17032. https://doi.org/10.1038/nenergy.2017.32
This article could not be verified and a wrong citation, this research group never did this work.
Choi, S. K. R. M. K., Ho-Baillie, A. W. Y., & Green, M. A. (2017). Stability of perovskite solar cells. Nature Energy, 2, Article 17032. https://doi.org/10.1038/nenergy.2017.32
contains inaccuracies.
However, when the first author used ChatGPT with the prompt “Generate the reference list as per APA,” the tool successfully reformatted the citations. For this case the articles had already been correctly identified, they were originally in inconsistent formats. ChatGPT helped standardize them into accurate APA style, saving significant time and effort.
By automating citation formatting with generative AI, researchers can redirect their focus toward critical analysis and writing, rather than spending disproportionate time on those tasks. These tools may support reference management and reduce time spent on citation formatting tasks.
Step 5: Critically Analyze and Evaluate
Critically Analyze and Evaluate is a crucial step in the literature review process that means beyond merely summarizing existing studies and this is mostly manually done. AI cannot—and should not—play a dominant role at this stage, as it primarily relies on human judgment, critical thinking, and domain-specific expertise which researchers develop over the time.
At this stage, students are expected to engage deeply with the literature by examining the strengths, limitations, methodologies, theoretical frameworks, and findings of each source. Critical analysis involves identifying patterns, contradictions, and gaps in the research, while evaluation requires assessing the credibility, relevance, and contribution of each work to the field. This step encourages higher-order thinking by prompting students to question assumptions, compare perspectives, and situate individual studies within broader scholarly conversations based on the outputs that AI already generated. At this point, the authors suggest the researchers make sure to follow each evaluation checklist and cross-checking list as mentioned to analyze any information generated by AI once all the articles are gathered (refer to Table 1). Students should follow those process in each step (from step 1), however, at this point it becomes more critical to be mindful while analyzing and evaluating each article.
As generative AI tools become increasingly integrated into academic work, it is essential that students learn to critically evaluate AI-generated content. Doing so not only upholds academic integrity but also promotes responsible, ethical, and informed engagement with emerging technologies. To support this, the authors propose a structured approach that includes clearly defined steps for the responsible use of generative AI and the check list becomes more critical at this point (consult Table 1). From the outset of a literature review assignment, students should consult the protocol and evaluation checklist throughout all the steps. This ensures they are aware of best practices in verifying sources, assessing credibility, and integrating AI assistance with their own critical thinking and academic judgment.
Step 6: Synthesize with AI support
This step requires active engagement and critical thinking from students. While AI can assist, the thematic synthesis of literature remains a largely manual and reflective process. Students must evaluate which information is most relevant to their research focus and thoughtfully decide how to organize and frame it. The authors encourage students to group literature thematically—by methodology, theoretical lens, findings, or research context. AI tools such as Elicit, Connected Papers, and Research Rabbit can support this process by suggesting clusters of related articles, generating mind maps, or outlining content based on semantic similarities. For example, if researchers enter the manuscript title inciteful will suggest research related to the topic.
INPUT: The literature review title to get the mapping research connections were inserted
However, it is ultimately the student’s responsibility to determine how these themes are connected, how ideas evolve across studies, and how these insights shape the foundation of their own research, though students can utilize AI tools to summarize. Therefore, from this step the authors heavily rely on human analysis and judgement.
Also, here is an example of how generative AI helped us summarize our work. The authors have two examples from ChatGPT and Perplexity.
Original
Efficiency is usually regarded as the most important aspect of a solar cell, the more efficient the more power you get for the material you use. So then it only makes sense to make the highest efficiency solar cells you can. Table 2 Shows us the recent advances in the in the efficiency of the perovskite solar cells. In 2017 we see we have achieved 21.2% power conversion efficiency. While according to the Office of Energy Efficiency and Renewable Energy Silicon Solar cells have laboratory power conversion efficiencies of around 25%. This means that we are incredibly close to achieving the same power conversion efficiency to the standard commercialized solar cell and that was three years ago. Today perovskite solar cells have been measured to reach efficiencies of 25.2% Laalioui, Alaoui, Dads, Assali, Ikken and Outzourhit (2020), this has surpassed the commercialized silicon solar cell efficiency. So with the efficiency of perovskite solar cells rising at an incredible rate it is no surprise that so much time and effort has gone into creating the most efficient perovskite solar cell.
Edited version using ChatGPT
Efficiency is widely regarded as a crucial aspect of solar cell technology; higher efficiency translates to greater power output relative to the material used. Consequently, it is logical to strive for the highest efficiency in solar cell development. Table 2 illustrates recent advancements in the efficiency of perovskite solar cells. In 2017, a power conversion efficiency of 21.2% was achieved. According to the Office of Energy Efficiency and Renewable Energy, laboratory power conversion efficiencies for silicon solar cells are approximately 25%. This indicates that perovskite solar cells were already nearing the efficiency of commercial silicon solar cells just three years ago.
As of today, perovskite solar cells have reached efficiencies of 25.2% (Laalioui et al., 2020), surpassing the efficiency of commercial silicon solar cells. Given the rapid increase in perovskite solar cell efficiency, it is understandable that substantial research and resources have been dedicated to optimizing their performance.
Edited version using Perplexity
Efficiency is usually regarded as the most important aspect of a solar cell; the more efficient the cell, the more power you can generate from the material used. Therefore, it makes sense to strive for the highest efficiency in solar cells. Table 2 shows recent advances in the efficiency of perovskite solar cells. In 2017, we achieved a power conversion efficiency of 21.2%. According to the Office of Energy Efficiency and Renewable Energy, silicon solar cells have laboratory power conversion efficiencies of around 25%. This means we are incredibly close to achieving comparable power conversion efficiency to that of standard commercialized solar cells, and this was three years ago. Today, perovskite solar cells have been measured to reach efficiencies of 25.2% (Laalioui, Alaoui, Dads, Assali, Ikken, and Outzourhit, 2020), surpassing the efficiency of commercialized silicon solar cells. Given that the efficiency of perovskite solar cells is rising at an incredible rate, it is no surprise that significant time and effort have been invested in creating the most efficient perovskite solar cell possible.
One of the key challenges in evaluating LLM output is assessing the quality of the result. In this case, both versions preserved the essential information, and most LLMs would likely perform similarly in terms of content accuracy. However, when it comes to presentation style, it really comes down to personal preference. Therefore, it will be up to the student and professor to choose one.
Step 7: Write the Literature Review
Then, the authors ask students to summarize the research done previously, identify the gaps and suggest how the explored area will mitigate some of the gaps. Generative AI tools can be incredibly helpful in improving clarity, structure, and coherence of the texts but should be used ethically and thoughtfully. While AI can support this stage, students must maintain authorship. The review should reflect their critical thinking, synthesis, and scholarly voice. AI is a co-pilot—not the driver—in producing high-quality academic work. Tools like ChatGPT can assist in structuring paragraphs, paraphrasing ideas, and improving transitions between themes, helping students frame their synthesis more coherently. For enhancing grammar, style, and readability, Grammarly offer real-time feedback and academic tone adjustments. Zotero, particularly when used with its note plugins, can streamline the process of organizing notes, exporting grouped references, and generating bibliographies. However, it’s important to note that asking AI tools like ChatGPT or DeepSeek to generate content on specific, nuanced topics such as PSC (the researched topic) may result in inaccurate or fabricated information, at times might be difficult to verify. These tools are most effective when used to refine language and tone, not to generate original content on complex or technical subjects without human verification.
The first step to write the review is to make sure the reference is not hallucinated. The second step is to determine whether the reference provides the summary and if there is relationship between the reference and the AI summary. The third step is to determine whether the AI interpretation is accurate. The next step is to verify again, if other papers/sources can be used to confirm the AI-generated content, please utilize them.
At this critical juncture when we can find wealth of information online it’s very important to differentiate between facts and opinions (claims, theories, and explanations). Fact checking can be done using above mentioned tools (Table 1), however, checking on reasoning might not be that easy for the novice researchers and the authors recommend working closely with your mentors and professors as it’s a time taking process to master it and needs high level domain knowledge.
The authors want to emphasize that instead of requiring students to manually repeat the literature review, the process should position AI as a supplementary tool, but the human judgement remains the primary tool, especially feedback from the PI and other veteran researchers. AI can be used effectively to generate initial research questions or summaries, but in-depth analysis and over-viewing the final products should remain human-led to ensure accuracy, critical evaluation, and academic rigor. Therefore, students must verify each information after they are generated which should still save significant time and manual searches.
Here is an example of how we did it-
\SciSpace paraphraser:
Original paragraph
Device Simulation and Machine Learning Analyses
Thus far, only a few hundred permutations have been explored experimentally, making it nearly impossible to manually test all potential perovskite materials in solar cells. As a solution, the inverse design approach integrates device-level simulations and machine learning models to effectively harness the potential of perovskite materials in solar cells. The quest for high-performance PSCs involves screening materials for stability, processability, and optimal optoelectronic properties. Additionally, the role of transport layers and appropriate deposition methods is much needed. The complexity of these demands calls for an innovative approach that can quickly identify the ideal material combinations and processing techniques to optimize performance metrics, such as PCEs, device stability, and scalability. All-inorganic compounds like FASnI3 and Cs2TiBr6 are notable promising materials due to favorable optical and electrical properties and excellent stability in differing environments, and a simulation involving these materials is beneficial to discover the combination that yields the greatest efficiency. A study conducted by R. K. Shukla et al. (2024) presents a numerical simulation of lead-free perovskite solar cells using a particular device architecture (FTO/ZnO/FASnI3/Cs2TiBr6/PEDOT/PSS/Spiro-OMeTAD/Au) to optimize the thicknesses of the electron transport layer (ETL), light absorber layers, and hole transport layer (HTL) to achieve the highest possible efficiency.
We prompted “simplify the paragraph”.
New simplified paragraph
Device Simulation and Machine Learning Analyses
Only a limited number of perovskite material permutations have been examined, hindering comprehensive experimental testing. The inverse design methodology combines device simulations with machine learning to optimize perovskite usage in solar cells. The pursuit of efficient PSCs necessitates evaluation of materials for stability, processability, and optoelectronic characteristics. Moreover, understanding transport layers and deposition techniques is crucial. The complexity of these requirements demands innovative strategies to swiftly identify optimal material combinations and processing methods for enhanced performance metrics. All-inorganic compounds such as FASnI3 and Cs2TiBr6 exhibit favorable properties and stability, warranting simulations to identify combinations for maximum efficiency. R. K. Shukla et al. (2024) conducted a numerical simulation of lead-free perovskite solar cells with a specific architecture to optimize layer thicknesses for peak efficiency.
Now, we prompted “expand the previous paragraph”
New expanded paragraph
Device Simulation and Machine Learning Analyses
To date, the experimental exploration of perovskite materials has been limited to only a few hundred permutations, which renders it exceedingly challenging, if not nearly impossible, to manually conduct comprehensive testing of all conceivable perovskite materials that could potentially be utilized in solar cell applications. In light of this predicament, the innovative inverse design methodology has been developed, which adeptly integrates advanced device-level simulations alongside sophisticated machine learning models, thereby effectively unlocking and harnessing the vast potential of perovskite materials in the context of solar cells. The pursuit of achieving high-performance perovskite solar cells (PSCs) necessitates a meticulous screening process for various materials, focusing not only on their stability and processability but also on their optimal optoelectronic properties, which are critical for enhancing overall device performance. Furthermore, it is imperative to consider the significant impact of transport layers and the selection of appropriate deposition methods, which cannot be overlooked in this complex endeavor. The intricate nature of these requirements underscores the necessity for a groundbreaking and innovative approach that can swiftly and accurately identify the most suitable combinations of materials and processing techniques, all aimed at optimizing essential performance metrics, including power conversion efficiencies (PCEs), device stability, and scalability in real-world applications. Among the various materials under consideration, all-inorganic compounds such as FASnI3 and Cs2TiBr6 have emerged as particularly promising candidates, owing to their advantageous optical and electrical properties, as well as their remarkable stability when subjected to varying environmental conditions; thus, conducting simulations that involve these materials is highly advantageous for uncovering the combinations that yield the maximum efficiency in solar cell applications. A noteworthy investigation conducted by R. K. Shukla et al. (2024) elucidates a numerical simulation aimed at lead-free perovskite solar cells, employing a specific device architecture (FTO/ZnO/FASnI3/Cs2TiBr6/PEDOT/PSS/Spiro-OMeTAD/Au) to meticulously optimize the thicknesses of the electron transport layer (ETL), light absorber layers, and hole transport layer (HTL), ultimately striving to achieve the highest feasible efficiency attainable in the realm of solar energy conversion .
The authors observed that the AI tools consistently accelerated early-stage literature review development , particularly for new researchers. Performance became more consistent in interpretive and comparative tasks, such as synthesizing findings across multiple studies, extracting methodological details, or generating comparisons. While often helpful, outputs in these domains occasionally simplified nuance, overstated claims, or required careful cross-verification with original sources. By contrast, the most consistent failures emerged in high-precision bibliographic tasks, especially citation generation and reference accuracy, where hallucinated and incomplete citations were repeatedly observed. These patterns indicate that AI tools are not uniformly reliable or unreliable; rather, their utility is task dependent.
Also, all AI-assisted activities reported in this manuscript were conducted in 2024, and the findings should be interpreted within that temporal and technical context. The study utilized OpenAI’s ChatGPT (GPT-3.5 and GPT-4 models as available in 2024) for text generation, summarization, paraphrasing, and citation testing; Perplexity (Academic/Pro version, 2024) for retrieval-augmented responses; and Elicit (2024 platform configuration) for structured academic extraction and comparison. Visual generation was conducted using DALL·E3 as integrated within ChatGPT in 2024. Citation mapping and network visualization were performed using Litmaps (2024 web-based version), while word cloud visualizations relied on publicly available 2024 frequency-based text analysis tools. Reference organization and formatting demonstrations were conducted using the 2024 desktop versions of Zotero and EndNote. Because generative AI systems are continuously updated and often non-deterministic, hallucination patterns, citation accuracy, synthesis quality, and stylistic outputs may vary across versions and over time. Explicitly situating the study within the 2024 AI ecosystem strengthens reproducibility, clarifies interpretive boundaries, and enhances the longitudinal value of the manuscript’s conclusions.
Although the examples presented in this study focuses on perovskite solar cell research, the framework is designed to be transferable across disciplinary contexts. Researchers or instructors seeking to replicate the process may begin by selecting a topic relevant to their field and then systematically apply the seven-step framework. Please note the authors had the topic already defined. First, a research question is identified and refined using disciplinary knowledge and AI-supported brainstorming tools such as ChatGPT, Claude, word cloud or Gemini. Second, foundational disciplinary knowledge, terminology, and key concepts are explored using AI-assisted inquiry and traditional scholarly sources. Third, relevant literature is identified through a combination of AI-supported search tools (e.g., Elicit, Consensus, Perplexity, Litmaps) and traditional databases appropriate to the discipline. Fourth, references are organized and managed using reference management software such as Zotero, Mendeley, or EndNote. Fifth, sources and AI-generated outputs are critically analyzed to evaluate methodological rigor, relevance, theoretical alignment, and evidentiary support. Sixth, findings are synthesized across studies to identify patterns, areas of agreement and disagreement, research gaps, and emerging themes. Finally, the literature review is drafted, revised, and refined using AI-assisted writing support where appropriate while maintaining scholarly judgment and author responsibility for all content.
Throughout the process, researchers should apply the verification procedures outlined in Table 1, including source verification, triangulation of claims, fact-checking, consultation of authoritative sources, and critical evaluation of AI-generated outputs. Outputs should be evaluated using the criteria of accuracy, completeness, transparency, coherence, and reliability.
Discussion
This study demonstrated how generative AI tools can be systematically integrated into the literature review process through a structured seven step framework designed for undergraduate researchers. By examining AI-generated outputs including citation maps, summaries, and visualizations—this work illustrated how AI can support early stages of literature exploration, organization, and synthesis while highlighting the importance of human verification and mentorship. The findings suggested that generative AI can significantly reduce the manual burden associated with literature discovery and organization, particularly for novice researchers navigating complex or interdisciplinary topics.
The integration of generative AI into the literature review process and overall research present both promising opportunities and notable challenges. AI tools can definitely streamline the initial stages of a literature review by accelerating the search, summarization, generating chart, images, mind maps and infographics and organization of scholarly content. These tools reduce the manual burden, making it easier for researchers especially for the novice scholars and those in interdisciplinary fields—to access, categorize, and engage with vast amounts of academic information which can be overwhelming. Additionally, AI can also assist in detecting emerging trends, creating visual representations like infographics, knowledge graph, tables, mind maps, and suggesting thematic groupings, which enhance the efficiency and depth of literature exploration. AI can often provide mentorship to the students about STEM persistence, and the mentorship can be tailored towards teaching use of AI in research process. (Okado et al., 2025)
However, there are limitations and risks. AI-generated content can sometimes lack nuance, context, or accuracy, AI can fake data and information that can be mistaken as rare, for example, for our case ChatGPT generated hallucinated referees while the authors asked ChatGPT to generate a list of articles focused on PSC in IEEE format. While fact-checking can be accomplished using Table 1, evaluating the reasoning behind arguments is often more challenging, especially for novice researchers. This process requires deeper critical thinking and domain-specific knowledge. Therefore, the authors strongly recommend that students work closely with their mentors and professors, as mastering this skill takes time and guided practice and no way authors are proposing that students do not need to maintain a regular mentor-mentee relationship. AI cannot replace human interaction and social and emotional aspects of learning.
The potential for bias in training data, superficial synthesis, or omission of critical works raises concerns about reliability and academic rigor. Furthermore, over-reliance on AI may undermine the development of critical thinking and analytical skills necessary for scholarly inquiry.
Our goal was to also contribute to the ongoing discussion on how to implement the use of AI in an ethical and responsible way. As the different AI tools come with subscriptions versions, students are more likely to use those, we as educators stand at a critical juncture and need to become more proactive of introducing the nuances and benefits of using AI technology to them.
Limitation of this Study and Future Scope
Also, there are several limitations of this work. This study is illustrative rather than experimental. The AI outputs analyzed represent specific prompts conducted in 2024 and do not reflect controlled repeated trials or parameter manipulation. Because generative AI systems are non-deterministic and continuously updated, behaviors documented here—including hallucination frequency and synthesis quality—may evolve over time. The authors used one prompt multiple times to standardize our documentation process. However, this work is not focused on the description of iterative refinement cycles or coding the outputs.
Additionally, the manuscript does not measure student learning outcomes or compare AI-assisted instruction to traditional pedagogical models. Future research should incorporate controlled prompting experiments, repeated trials, version-controlled comparisons, and empirical assessment of student learning outcomes. Longitudinal studies examining how AI-assisted scaffolding influences research literacy, critical evaluation skills, and epistemic awareness would further strengthen the instructional implications of this work.
The authors acknowledge that different AI platforms may rely on retrieval-augmented generation (RAG), curated metadata indexing, cached embeddings, or direct LLM inference. However, as end users of publicly accessible free versions of these platforms in 2024, the authors did not have access to internal architectural specifications, version-locked environments, or configurable system parameters. Our study therefore treats these platforms as black-box instructional tools, reflecting authentic usage conditions. Our focus was on how students interacted with outputs generated under typical real-world constraints and how human verification mitigated architectural opacity
This study did not aim to evaluate the technical performance or statistical behavior of generative AI models. The tools were used in their default configurations, reflecting typical undergraduate access, and the examples are presented as instructional demonstrations rather than experimentally controlled trials. Recognizing that generative AI systems are probabilistic and subject to architectural opacity and hallucination risk, the framework intentionally emphasized structured verification and critical evaluation, positioning domain expertise and human judgment as the final authority in assessing AI-generated outputs.
Although informed by design-based research principles, this study does not represent a complete DBR cycle because the framework was not implemented and evaluated across multiple instructional settings with learners. Future research should examine the framework through classroom-based implementation and iterative redesign.
The framework was demonstrated using a single case study domain (perovskite solar cells). Although the framework is designed to be transferable across disciplines, additional validation across fields would strengthen its generalizability. Also, the primary objective of this work was to familiarize students with generative AI for use in conducting literature reviews.
In the future the team plan to continue to evaluate the performance of multiple state-of-the-art LLMs (Schneider et al., 2026) focusing on the extraction of experimental metadata—including input features, target variables, dataset size, model types, and performance metrics—from peer-reviewed scientific papers and the evaluation framework grounded in the ordered 3C principles, i.e., compliance, completeness, and correctness of the AI tools.
Looking ahead, the future of AI in literature reviews lies in a balanced, hybrid approach—one where AI serves as a supportive tool rather than a replacement for human judgment and analytical thinking. Training students and researchers to critically evaluate AI outputs, verify sources, and apply domain expertise is essential. As AI technologies evolve, establishing ethical guidelines, promoting transparency in AI algorithms, and integrating AI literacy into research training will be vital for fostering responsible and impactful use of these tools in the scholarly landscape as AI is here to stay and transform the research.
Conclusion
This manuscript illustrates how generative AI tools can be incorporated into literature review processes through an example-driven framework. The study offers a transparent demonstration of how various AI tools function across stages of literature development. The examples show that AI systems can assist with tasks such as topic familiarization, thematic clustering, summarization, visualization, and workflow organization with lesser human supervision. At the same time, the documented instances of hallucinated citations, incomplete synthesis, and bibliographic inaccuracies underscore the necessity of verification and sustained human oversight.
Taken together, the findings suggest that generative AI tools hold significant instructional potential when positioned as scaffolding mechanisms rather than autonomous agents. Their value appears strongest in supporting idea generation, structural organization during early-stage literature engagement. However, high-precision scholarly tasks remain firmly dependent on human judgment and their disciplinary expertise.
The seven-step AI-assisted framework contributes in offering educators a transparent, adaptable structure for teaching responsible AI integration, grounded in both opportunity and limitation. As AI systems continue to evolve, future research should examine learning outcomes, reproducibility across model updates, and longitudinal impacts on student research development. This study provides a practical and critically informed foundation for integrating AI into research instruction while maintaining academic rigor and ethical responsibility.
Acknowledgments
The authors would like to express their gratitude to Luke Schneider for his valuable assistance in generating the literature map. Additionally, we acknowledge the support of the 2024 Penn State Institute of Energy and the Environment (IEE) Seed Grant Program, and the 2024 Penn State Behrend Research/Creative Activities Seed Grants, both of which provided essential funding for this research. Also, this project was partially supported by summer manuscript competition program funding from the Center for Excellence in Teaching and Learning at Kennesaw State University.
Generative AI tools such as ChatGPT and Co-Pilot were used to assist with framing and editing purposes only; however, all ideas and interpretations presented are original to the authors.

_research__fo.png)
.png)
_broad_citation_network_showing_connections_among_key_studies_on_perovskite_solar_cells.png)
