Processing math: 100%
Nkrumah, S. K., Tucker, S. M., Boyle, F., & Walsh, J. (2026). Comparing AI Approaches for Automated Skill Extraction: From Embeddings to Agentic Systems for Engineering Education. ASEE Computers in Education, 15(4), 26–48. https://doi.org/10.18260/B4B8-9A-80426
Download all (4)
  • Figure 1. ReAct agent workflow. The LLM produces a reasoning trace followed by either a search_esco tool call or a finish action; tool observations are appended to the state and control returns to the LLM. The loop terminates at finish or a 20-step budget (median 2 reasoning–action cycles).
  • Figure 2. Recall improvement from enhanced multi-layer matching across all AI system configurations
  • Figure 3. Average latency per sample across AI system configurations (log scale). Llama latencies reflect local GPU execution; Gemini latencies reflect cloud API access.
  • Figure 4. Enhanced precision versus enhanced recall across all 11 configurations. Points are coloured by paradigm. Dashed lines show iso-F1 contours. ReAct occupies the top-left (highest recall, lowest precision), tool-grounded Gemini Flash reaches the highest F1, and embeddings cluster at lower recall.

Abstract

The rapid pace of technological change has created a persistent skills gap between what graduates learn and what industry demands. Automated skill extraction from text offers a scalable approach to aligning education with workforce needs, and the arrival of modern artificial intelligence (AI) systems has accelerated this direction as new skills emerge and older ones fall out of use. Comparisons of modern approaches that span embedding, generative, and agentic paradigms remain sparse in the literature we surveyed. We report a comparative evaluation of 11 model configurations across these three paradigms, embedding-based models, generative large language models (LLMs), and an agentic AI system, using the European Skills, Competences, Qualifications and Occupations (ESCO) benchmark (926 samples, 1,696 annotated skills). The ReAct agentic configuration achieved the highest enhanced recall (89.7%), exceeding the best generative configuration (83.7%) and the best embedding configuration (49.1%). When precision is considered alongside recall, however, tool-grounded Gemini Flash achieved the highest enhanced F1 (51.6%) because ReAct’s coverage advantage is partly offset by lower precision (22.7%). Tool-grounded configurations also led on exact matching (62.8%), supporting precise alignment with the ESCO taxonomy. Our enhanced multi-layer matching pipeline improved measured recall by 11–72 percentage points across all configurations, indicating that exact-matching evaluations understate generative system capabilities. We offer scenario-specific deployment guidance for engineering educators applying these tools to curriculum design, career counselling, and workforce tracking, with the appropriate choice depending on whether outputs are surfaced to learners directly or triaged by human advisors.

1. Introduction

The engineering workforce is undergoing a profound transformation driven by technological advancement and evolving industry demands. According to the World Economic Forum (World Economic Forum, 2025), 59% of the global workforce will require training by 2030, with engineering fields particularly affected by the emergence of artificial intelligence and automation (Acemoglu & Restrepo, 2022; Green, 2024). This creates a fundamental challenge for engineering education: how can academic programs ensure their curricula align with the rapidly changing skills that employers require? Traditional approaches to curriculum alignment, employer advisory boards, alumni surveys, and periodic program reviews, provide valuable insights but cannot keep pace with the velocity of change in technical fields. Natural language processing (NLP) techniques offer a promising solution by automatically extracting skill mentions from job postings, resumes, and syllabi (Zhang et al., 2022). When mapped to standardised taxonomies like the European Skills, Competences, Qualifications and Occupations (ESCO) classification (see Section 2.2), which contains 13,939 distinct skills, these extractions can provide actionable intelligence for curriculum design and career guidance. The potential applications are substantial: curriculum committees could analyse job postings to identify emerging requirements, career services could provide personalised skill-gap analyses, and accreditation processes could systematically map student work to required competencies. However, comprehensive comparisons of modern AI approaches are lacking. Prior research has evaluated individual methods in isolation: Bidirectional Encoder Representations from Transformers (BERT), a language model that understands words in context by reading text in both directions simultaneously, much like how a human reads surrounding sentences to determine whether “Java” refers to a programming language or an island. BERT-based embeddings convert text into numerical representations that capture meaning, so phrases like “Python programming” and “Python development” are recognised as similar (Reimers & Gurevych, 2019), generative LLMs (Clavié & Soulié, 2023), or agentic systems (Yao et al., 2023). Engineering educators facing technology adoption decisions lack guidance on which approach offers the best trade-offs, that is why we undertook this research. This study addresses this gap through systematic evaluation of 11 model configurations across three paradigms: embedding-based approaches (SBERT, ModernBERT), generative LLMs (Gemini 2.5 Flash/Pro, Llama 3.2), and agentic AI systems (ReAct) using Gemini 2.5 Pro as the base model. We introduce an enhanced multi-layer matching pipeline employing eight progressive strategies to account for skill paraphrasing, a critical methodological contribution given that models may identify semantically correct skills using different surface forms than the gold-standard labels. For example, if the gold-standard label is “develop software” but the model predicts “software development,” both refer to the same skill, yet strict exact-matching evaluation would mark this as incorrect. Our pipeline recognises such equivalent expressions, providing a fairer assessment of model capabilities. This research addresses the following research questions:

  1. How do modern AI approaches compare for automated skill extraction from text?

  2. How effective is an enhanced multi-layer matching pipeline for improving skill recognition accuracy?

  3. What are the practical implications for engineering curriculum alignment and career guidance?

The remainder of this paper is organised as follows: Section 2 reviews related work in skill extraction, taxonomies, and AI approaches. Section 3 describes our methodology, including the 11 AI configurations and the enhanced matching pipeline. Section 4 presents results, and Section 5 discusses implications including practical deployment recommendations for engineering programs. We address limitations in Section 6 and conclude in Section 7.

2. Literature Review

2.1. Evolution of Skill Extraction

Skill extraction from text has evolved through several distinct phases. Early approaches relied on rule-based systems and keyword matching using hand-crafted lexicons (Kivimäki et al., 2013), but suffered from limited coverage and required continuous manual updates, with new emerging skills. Word embeddings (Mikolov et al., 2013) enabled semantic matching beyond exact strings, but struggled with multi-word expressions and context-dependent meanings.

Transformer-based models (Vaswani et al., 2017) like BERT (Devlin et al., 2019) marked a significant advancement by capturing contextual relationships. Researchers treated skill extraction as named entity recognition or semantic similarity tasks (Decorte et al., 2022; Zhang et al., 2022). To evaluate these approaches, the SkillSpan benchmark (Zhang et al., 2022) provided important data distinguishing between hard skills (technical competencies like “Python programming”) and soft skills (interpersonal abilities like “communication” or “teamwork”). However, evaluation revealed that soft skills remain particularly challenging to extract, even with transformer-based models, due to their diverse expression patterns, for instance, “works well with others,” “team player,” and “collaborative” all convey similar competencies but use completely different phrasing.

LLMs introduced zero-shot extraction capabilities (Clavié & Soulié, 2023; Decorte et al., 2023), enabling skill identification without task-specific training, meaning models can extract skills from any text without needing labelled examples as a prerequisite. This flexibility is valuable for engineering education, where skill requirements span diverse domains from software development to biomedical engineering. However, these approaches face challenges including hallucination (the generation of plausible but incorrect skills, such as inventing “advanced quantum networking” when the text only mentions basic network configuration), and difficulty aligning outputs with specific taxonomies. For instance, an LLM might output “coding” when the ESCO taxonomy specifies the precise label “develop software applications,” making it difficult to aggregate results or integrate with existing databases.

Most recently, agentic AI systems have emerged as a promising paradigm, combining LLM reasoning with external tool use rather than generating outputs in a single pass (Yao et al., 2023). The ReAct (Reasoning + Acting) framework exemplifies this approach by interleaving reasoning traces, where the model explicitly thinks through the problem, with actions like database lookups, then observing results before continuing. For skill extraction, this means an agent can reason that “managing cross-functional teams” likely implies leadership skills, query the ESCO database to verify this inference, and refine its output based on what it finds. Rather than inventing skill names, the agent validates its inferences against actual taxonomy entries, ensuring outputs align with standardised classifications. Crucially, this grounding in authoritative databases reduces hallucination. The approach also enables detection of implicit skills that simpler methods miss, as the agent can reason about competencies underlying described activities. Reflexion (Shinn et al., 2023) extends this paradigm with self-correction capabilities, enabling agents to recognise errors and adjust their reasoning within a single task. Despite these advances, comparisons of agentic approaches against embedding and generative baselines for skill extraction remain sparse in the published literature we surveyed. This study contributes one such comparison, framed around an engineering education deployment context.

2.2. Skill Taxonomies for Engineering Education

Standardised skill taxonomies provide essential infrastructure for automated skill extraction by defining the target label space and enabling cross-document comparison. Without a common vocabulary, extracted skills from different sources, such as job postings, resumes, and syllabi, cannot be meaningfully aggregated or compared. For instance, one document might mention “programming,” another “software development,” and a third “coding”; a standardised taxonomy maps these to a single canonical skill, enabling quantitative analysis across documents. The European Skills, Competences, Qualifications and Occupations (ESCO) taxonomy (European Commission 2023b, 2023a) represents the most comprehensive multilingual skill classification currently available, developed by the European Commission to support labour market transparency. ESCO contains 2,942 occupations and 13,939 skills organised hierarchically, spanning knowledge areas, technical competencies, and transversal skills like communication and teamwork. Crucially, ESCO’s explicit skill-occupation links enable direct mapping between educational outcomes and career pathways, for example, identifying which specific skills a student needs to develop for their target occupation. The O*NET database (U.S. Department of Labor, 2024) provides a US-focused alternative maintained by the Department of Labor, but lacks ESCO’s international applicability and multilingual support. Other taxonomies were considered: the Lightcast Open Skills library offers broader coverage ( 33,000 skills) but is primarily a commercial product with limited occupation-skill linkages, while the European e-Competence Framework (e-CF) provides well-defined competences but is restricted to the ICT sector. We selected ESCO due to its comprehensive coverage, hierarchical organisation enabling analysis at multiple granularity levels, and growing adoption in workforce development applications worldwide.

2.3. AI Approaches for Skill Extraction

Embedding-based approaches represent text as dense numerical vectors where semantically similar content appears nearby in the vector space, meaning “Python programming” and “Python development” would be positioned close together, enabling matching without exact string overlap. Sentence-BERT (SBERT) (Reimers & Gurevych, 2019) extended BERT for efficient sentence-level embeddings, while ModernBERT (Warner et al., 2024) offers architectural improvements for faster inference. For skill extraction, these models compute similarity between input text and pre-computed embeddings of all taxonomy skills, returning the closest matches. These approaches excel at processing speed (milliseconds per sample) but achieve lower accuracy than generative methods, particularly for implicit skills, since they can only match what is explicitly stated, not reason about competencies underlying described activities."

Generative LLMs extract skills through natural language generation, producing skill lists as text output rather than matching against pre-computed vectors. Models like Gemini (Gemini Team et al., 2024) and Llama (Touvron et al., 2023) can generate extractions without task-specific fine-tuning by leveraging knowledge acquired during pre-training, the initial phase where models learn language patterns and world knowledge from massive text corpora before being deployed for specific tasks. Different prompting strategies affect performance: zero-shot provides only task instructions, few-shot includes example extractions to guide the model, and chain-of-thought (Wei et al., 2022) encourages step-by-step reasoning before generating outputs. Tool-grounded approaches augment LLMs with database access, allowing the model to look up and validate skills against the taxonomy during generation. While flexible and capable of identifying implicit skills through reasoning, generative approaches have higher latency (seconds rather than milliseconds) and may produce outputs misaligned with target taxonomies, such as generating “coding” when the taxonomy requires “develop software applications.”.

Agentic AI systems represent the most recent paradigm, combining LLM reasoning with autonomous tool use and multi-step planning employed rather than the generation of all outputs in a single pass. The ReAct framework (Yao et al., 2023) interleaves reasoning traces (where the model explicitly thinks through the problem), with actions like database lookups, it then observes the results before continuing. For example, when processing text describing “led cross-functional team meetings,” an agent might reason this implies leadership and coordination skills, query the ESCO database to find matching entries, and refine its output based on actual taxonomy labels. This grounded approach reduces hallucination by validating inferences against authoritative databases rather than inventing skill names, while also enabling identification of implicit skills that embedding and basic generative approaches miss.

3. Methodology

3.1. Research Design

We conducted a comparative study evaluating 11 AI system configurations (see Table 1) across three paradigms, namely embedding-based, generative LLM, and agentic AI on a standardised benchmark derived from the ESCO taxonomy. Our experimental design prioritises recall over precision, though both metrics are valuable depending on the application context. Recall measures the proportion of actual skills correctly identified, essentially asking “of all the skills that exist in the text, how many did we find?” Precision measures what proportion of predicted skills were actually correct, asking “of everything we suggested, how much was right?” We prioritised recall for this study because educational contexts typically benefit from comprehensive skill identification. Missing relevant skills can be more consequential than suggesting additional skills for consideration, a curriculum analysis tool that fails to identify important skill gaps could leave students underprepared for the workforce, whereas suggesting extra skills gives educators more options to review and filter. Applications with different requirements, such as automated credential verification where false positives carry serious consequences, would appropriately prioritise precision instead. All experiments used identical input data and evaluation procedures to ensure fair comparison. Gemini systems were accessed via Google’s cloud API, while Llama was run locally on an NVIDIA RTX 4080 GPU; this infrastructure difference affects latency comparisons but not accuracy metrics. We measured both exact recall (requiring precise string matches) and enhanced recall (allowing semantically equivalent matches through our multi-layer pipeline), along with processing latency to characterise speed-accuracy trade-offs.

3.2. Dataset

Our evaluation uses the Skill Extraction benchmark (Decorte et al., 2022), an extension of the SkillSpan dataset (Zhang et al., 2022) with ESCO skill annotations. The benchmark comprises 926 job posting text samples containing 1,696 gold-standard skill annotations drawn from the 13,939-skill ESCO classification v1.1.0. Unlike typical Named Entity Recognition (NER) datasets, which identify and label specific items in text, for example recognising “Python” as a skill and “Google” as an organisation in a job posting, but without linking these to taxonomies, this benchmark provides direct mappings between text spans and standardised ESCO skill identifiers-, enabling rigorous evaluation of taxonomy-grounded extraction rather than mere mention detection.

The dataset combines three complementary subsets: TECH (technology sector job postings), HOUSE (household services), and TECHWOLF (diverse industries), together representing the variety of skill-bearing texts encountered in educational and workforce contexts. Each sample contains one to five gold-standard skills spanning domains such as software development, data science, mechanical engineering, and project management. Importantly, the benchmark includes both explicit skills (directly stated, e.g., “Python programming”) and implicit skills (competencies underlying described activities, e.g., “collaborate with stakeholders” implying communication skills), providing a rigorous test of extraction capabilities beyond simple keyword matching.

3.3. AI Systems Evaluated

We evaluated 11 configurations (Table 1) across three distinct AI paradigms:

Embedding-Based Systems: We tested SBERT using the sentence-transformers/all-MiniLM-L6-v2 checkpoint, a 22M-parameter general-purpose sentence encoder, and ModernBERT using the answerdotai/ModernBERT-base checkpoint in zero-shot configuration. Both systems computed cosine similarity between input text embeddings and pre-computed embeddings for all 13,939 ESCO skills, returning the highest-similarity matches.

Generative LLM Systems: We tested Google’s Gemini 2.5 Flash (API identifier gemini-2.5-flash) in three configurations: zero-shot, few-shot with ten examples, and tool-grounded with ESCO database access. Gemini 2.5 Pro (gemini-2.5-pro) was evaluated in two configurations: chain-of-thought, and few-shot chain-of-thought combining examples with reasoning. All Gemini calls used temperature 0 for deterministic output and were made via Google’s cloud API between 9 December 2025 and 1 January 2026. Meta’s Llama 3.2 3B-Instruct, served via the llama3.2 tag in Ollama (3.21B parameters, Q4_K_M quantisation, 128k-token context window), was evaluated in three configurations: zero-shot, few-shot, and tool-grounded, and run locally on an NVIDIA RTX 4080 GPU. Each configuration used carefully designed prompts specifying the extraction task and desired output format.

Agentic System: We implemented a ReAct (Yao et al., 2023) agent using Gemini 2.5 Pro as the reasoning backbone, with access to an ESCO skill lookup tool. The agent’s architecture has three components. First, the state accumulates the input text, the evolving reasoning trace, the history of tool calls, and the current list of candidate skills and knowledge terms; the state is passed to the LLM on every step. Second, the action space exposes two operations: search_esco(query), which performs a semantic lookup against the 13,939-skill ESCO index and returns the top-k matches with their canonical labels, and finish(skills, knowledge), which terminates the loop and returns the final extraction. Third, the control loop runs the LLM on the current state, parses the next action, executes it, appends the observation to the state, and repeats. Termination occurs when the agent emits finish or when the step budget of 20 reasoning–action cycles is reached; empirically the agent terminates after 1–3 reasoning–action cycles on most samples (mean 1.7, median 2 across a 100-sample subset with full trace metadata; see Figure 1). The agent is prompted to identify both explicit and implicit skills and to validate inferences through ESCO lookups before committing to them. We make no claim that this specific configuration is optimal; our goal is to situate one representative agentic system alongside embedding and generative baselines so that the paradigm’s trade-offs can be measured rather than asserted.

Figure 1
Figure 1.ReAct agent workflow. The LLM produces a reasoning trace followed by either a search_esco tool call or a finish action; tool observations are appended to the state and control returns to the LLM. The loop terminates at finish or a 20-step budget (median 2 reasoning–action cycles).

3.4. Enhanced Multi-Layer Matching Pipeline

A fundamental challenge in evaluating skill extraction is that AI systems may predict semantically correct skills using different surface forms than the gold-standard labels. For instance, a system predicting “software development” when the gold label is “develop software” has identified the correct skill but would receive no credit under exact string matching.

We address this limitation through an enhanced multi-layer matching pipeline employing eight progressive strategies, applied in order from most to least strict:

  1. Exact Match: Case-insensitive string equality between prediction and gold label

  2. Substring Match: Either string contains the other as a substring

  3. Fuzzy Partial: TheFuzz library (SeatGeek, 2023) partial ratio score ≥ 85

  4. Fuzzy Token Sort: TheFuzz token-sorted comparison score ≥ 85

  5. Fuzzy Token Set: TheFuzz token-set comparison score ≥ 85

  6. Sequence Matcher: Python difflib SequenceMatcher ratio ≥ 0.85

  7. Jaccard Similarity: Word-level overlap coefficient ≥ 0.7

  8. Cosine Similarity: SBERT embedding similarity ≥ 0.85

Each predicted skill is evaluated against all gold labels using these strategies in sequence, with the first successful match recorded. This approach provides nuanced assessment of system capabilities while maintaining rigorous matching standards.

3.5. Evaluation Metrics

We report precision, recall, and F1 under both exact and enhanced matching. Exact metrics use normalized string equality against ESCO labels and represent performance for applications requiring precise taxonomy alignment. Enhanced metrics use the eight-layer matching pipeline described above and represent performance for applications where semantic equivalence suffices. Average latency per sample is also reported to support deployment decisions.

We treat recall as the primary metric for engineering education use cases, with precision and F1 reported alongside it for deployment choice and for transparency regarding the precision–recall trade-off analysed in §4.4. This emphasis reflects the downstream cost profile in our target applications: in curriculum mapping, career guidance, and skills gap analysis, a missed skill can lead a student away from a relevant pathway or omit it from advising entirely, whereas a surplus candidate is typically filtered at low cost by a human advisor or a taxonomy-aware post-processing step before the result reaches the student. Recall therefore governs the worst-case outcome (a pathway the student never sees), while precision governs a reviewable one (a candidate that can be removed). We nonetheless report precision and F1 throughout because the appropriate balance depends on whether the output is surfaced directly to a learner or triaged by an advisor; the scenario-specific deployment recommendations in §5.2 use both.

4. Results

4.1. Overall Performance Comparison

Table 1 presents precision, recall, and F1 for all 11 AI system configurations, under both exact and enhanced matching. Reporting precision alongside recall surfaces a trade-off that a recall-only view obscures: the ReAct agentic system achieved the highest enhanced recall (89.7%), but its enhanced precision of 22.7% is the second-lowest among generative and agentic configurations. Tool-grounded Gemini Flash achieved the highest enhanced F1 (51.6%) by combining moderate recall (79.9%) with substantially higher precision (38.1%). We return to this pattern in §4.4.

Table 1.Performance of 11 AI system configurations on the ESCO benchmark (926 samples, 1,696 gold skills). P/R/F are reported under exact matching and under the enhanced multi-layer matching pipeline. Model identifiers are defined in §3.3.
Exact Enhanced
# System Config P R F1 P R F1 Latency
1 Gemini 2.5 Flash tool-⁠grounded 16.4% 62.8% 26.0% 38.1% 79.9% 51.6% 15,822ms
2 Gemini 2.5 Flash few-shot 1.7% 15.7% 3.1% 36.8% 82.0% 50.8% 2,662ms
3 Gemini 2.5 Pro few-shot-cot 2.5% 13.7% 4.2% 34.0% 79.1% 47.6% 6,742ms
4 Llama 3.2 3B few-shot 2.6% 12.4% 4.4% 31.7% 75.8% 44.8% 4,065ms
5 SBERT zero-shot 13.0% 24.5% 17.0% 38.4% 35.5% 36.9% 11ms
6 ReAct (Gemini 2.5 Pro) react 2.0% 39.6% 3.9% 22.7% 89.7% 36.3% 16,361ms
7 Llama 3.2 3B tool-grounded 7.0% 39.4% 11.9% 26.2% 58.3% 36.2% 58,248ms
8 Gemini 2.5 Flash zero-shot 0.6% 11.6% 1.2% 22.4% 83.5% 35.3% 3,677ms
9 Gemini 2.5 Pro cot 0.8% 12.8% 1.5% 19.7% 83.7% 32.0% 8,549ms
10 ModernBERT zero-shot 0.5% 0.6% 0.6% 23.3% 49.1% 31.6% 16ms
11 Llama 3.2 3B zero-shot 0.4% 8.6% 0.7% 9.1% 75.3% 16.2% 11,438ms

Model identifiers are listed in §3.3. Versions accessed December 2025 – January 2026.

Several patterns emerge. On enhanced F1, the top four configurations are all generative systems rather than the agentic ReAct, because ReAct’s coverage advantage is partly offset by lower precision. On enhanced recall, ReAct remains the clear leader, with a 6–14 percentage-point margin over Gemini and Llama configurations. Tool-grounded methods dominate exact precision and exact recall because they constrain outputs to valid ESCO entries rather than generating free-form skill names. Embedding models occupy a distinct regime: SBERT achieves competitive exact precision (13.0%) because its nearest-neighbour retrieval returns canonical ESCO strings, but its enhanced recall (35.5%) is far below generative systems.

Latency spans four orders of magnitude: embedding models process samples in 11–16 milliseconds, Gemini 2.5 configurations require 2.6–7.6 seconds via cloud API, and the ReAct agentic approach requires approximately 16 seconds per sample. Llama 3.2 3B latencies (4–58 seconds) reflect local execution on consumer hardware (RTX 4080) rather than optimised cloud infrastructure, so direct latency comparisons with Gemini are not appropriate; accuracy comparisons remain valid.

Few-shot prompting did not consistently improve over zero-shot on enhanced recall, though it substantially improved precision (Gemini 2.5 Flash: 22.4%  36.8%; Llama 3.2 3B: 9.1%  31.7%), suggesting examples primarily help constrain output vocabulary. Among generative systems, Gemini 2.5 models consistently outperformed Llama 3.2 3B across configurations.

4.2. Impact of Enhanced Matching Pipeline

Figure 2 illustrates the substantial impact of our multi-layer matching pipeline. Enhanced matching improved measured recall by 11–72 percentage points across all configurations.

Figure 2
Figure 2.Recall improvement from enhanced multi-layer matching across all AI system configurations

The most dramatic improvement occurred for Gemini Pro chain-of-thought (70.9 percentage points), suggesting chain-of-thought prompting produces varied skill descriptions matching gold labels semantically but not lexically. This finding has important methodological implications: evaluations using exact matching only substantially underestimate generative system capabilities.

4.3. Speed-Accuracy Trade-offs

Figure 3 presents the latency comparison. The four-order-of-magnitude difference between embedding models (11–16ms) and tool-grounded approaches (15,000–58,000ms) creates fundamentally distinct deployment scenarios. Note that Llama’s higher latencies (up to 58 seconds for tool-grounded) reflect local GPU execution rather than optimised cloud infrastructure. Among cloud-based systems, Gemini Flash zero-shot offers an attractive middle ground: 83.5% enhanced recall at 3.7 seconds per sample.

Figure 3
Figure 3.Average latency per sample across AI system configurations (log scale). Llama latencies reflect local GPU execution; Gemini latencies reflect cloud API access.

4.4. Precision–Recall Trade-off

Because recall alone cannot distinguish a system that finds the right skills from one that simply returns a longer list, we examined how each configuration balances precision against recall. Figure 4 plots all 11 configurations in this space, with iso-F1 contours for reference. Three distinct regimes are visible.

Figure 4
Figure 4.Enhanced precision versus enhanced recall across all 11 configurations. Points are coloured by paradigm. Dashed lines show iso-F1 contours. ReAct occupies the top-left (highest recall, lowest precision), tool-grounded Gemini Flash reaches the highest F1, and embeddings cluster at lower recall.

ReAct sits in the top-left: its 89.7% enhanced recall is paired with 22.7% enhanced precision, indicating that a non-trivial share of its coverage advantage comes from producing more candidates per sample. For example, on a single-sentence input naming Solidity as the gold skill, the ReAct agent returned 12 candidate skills and 13 candidate knowledge terms (25 outputs total). It matched the gold correctly, but also introduced many unrelated items. ReAct’s gains should therefore be interpreted as higher coverage with lower specificity, not as uniformly higher quality.

Tool-grounded Gemini Flash reached the highest enhanced F1 in our evaluation (51.6%): by constraining outputs to valid ESCO entries it raised precision to 38.1% while retaining 79.9% recall. Few-shot and chain-of-thought generative configurations cluster just below this point. Embedding models occupy a lower-recall regime but remain competitive on precision: SBERT’s 38.4% enhanced precision is comparable to the top generative systems, reflecting that nearest-neighbour retrieval returns canonical taxonomy strings rather than paraphrased outputs.

These results modify the interpretation of recall-only rankings. If the goal is exhaustive coverage of candidate skills for downstream human review (for example, broad curriculum mapping), ReAct’s higher recall is the relevant metric. If the goal is to produce a shorter, higher-confidence list of matches (for example, automated career suggestions shown directly to students), tool-grounded generation offers a more favourable balance. We return to these scenario-specific recommendations in §5.2.

4.5. Analysis of Agentic System Behaviour

The ReAct agent’s high recall stems from three mechanisms: (1) implicit skill detection: reasoning about skills implied by activities even when not explicitly mentioned; (2) validation through database lookup: querying ESCO to find actual entries matching inferences; and (3) self-correction: adjusting reasoning when lookups return unexpected results. These benefits come at latency cost: 1–3 reasoning–action cycles per sample, contributing to the 16-second average processing time (Table 1). As §4.4 notes, the same permissive generation that enables implicit-skill coverage also produces the lower precision observed in Table 1.

To ground this pattern, Table 2 shows four representative sample-level outputs drawn directly from our saved predictions. Case 1 illustrates how ReAct identifies an implicit skill that an embedding baseline misses entirely, returning noisy but usable output for fuzzy-matched ESCO terms. Case 2 shows the over-extraction pattern: a single-gold input yields 25 ReAct candidates while tool-grounded Gemini Flash returns the correct ESCO entry plus one near-miss. Case 3 shows the same pattern on a longer input. Case 4 shows a generative zero-shot configuration producing a 39-term candidate list for a one-gold input, where most candidates are near-duplicates generated from the same surface expressions in the text. Across all four cases, the correct gold skill is recovered by the higher-recall systems, but the ratio of signal to noise varies sharply by configuration.

Table 2.Representative sample-level outputs illustrating precision–recall behaviour across configurations. “|preds|” is the number of candidate skills plus knowledge terms returned. Gold skills matched under enhanced matching are shown in bold.
Input (excerpt) Gold System Output (|preds|)
Case 1. Implicit skill: gold term not in text
“Advanced knowledge of architecture and design across all systems” information architecture ModernBERT-base “*”, “Advanced”, “knowledge”, “all” (4)
ReAct (Gemini 2.5 Pro) architectural design, system design, solution design, system architecture, enterprise architecture, ... (8)
Case 2. Over-extraction on short input
“Cypress for frontend is a must have, pytest for backend, and truffle for Solidity is a good addition.” Solidity Gemini 2.5 Flash (tool-grounded) implement front-end website design, Solidity (2)
ReAct (Gemini 2.5 Pro) software testing, automated testing, test planning, ..., Solidity, frontend development, ... (25)
Case 3. Moderately specific input
“Hands-on experience in Core Java Spring Boot and Microservices” Java (computer programming) Gemini 2.5 Flash (tool-grounded) apply basic programming skills, Java (computer programming), web services (3)
ReAct (Gemini 2.5 Pro) developing in Java, Spring Boot, microservices, Core Java, Java, Java (computer programming), Microservice architecture, ... (25)
Case 4. Generative expansion of surface expressions
“Cypress for frontend is a must have, pytest for backend, and truffle for Solidity is a good addition.” Solidity Gemini 2.5 Flash (zero-shot) Frontend testing, Developing frontend tests, Implementing frontend test automation, ..., Solidity, Solidity programming language, ... (39)

5. Discussion

5.1. Addressing Research Questions

RQ1: How do modern AI approaches compare for automated skill extraction from text? The answer depends on which metric is prioritised. On enhanced recall, the ReAct agentic configuration leads, with 89.7% recall exceeding the best generative configuration (Gemini Pro chain-of-thought, 83.7%) by 6 percentage points and the best embedding configuration (ModernBERT, 49.1%) by more than 40. On enhanced F1, however, the leader is tool-grounded Gemini Flash (51.6%), because ReAct’s coverage advantage is paired with lower precision (22.7%) (see §4.4 and Table 1). These results indicate paradigm-specific strengths rather than a single dominant approach, and the interpretation of “best” therefore depends on whether coverage, precision, or a balance of the two is most valuable for the application.

For applications requiring alignment with ESCO taxonomy labels, such as generating standardised reports or populating databases, tool-grounded approaches provide the highest exact recall (62.8% for Gemini Flash) we observed. For applications that tolerate broader candidate lists intended for downstream human review, the ReAct agent offers the highest enhanced recall at a cost of 16 seconds per sample and lower precision. For real-time applications where latency dominates, embedding models provide millisecond responses with lower recall but competitive enhanced precision.

RQ2: How effective is an enhanced multi-layer matching pipeline for improving skill recognition accuracy? The multi-layer matching pipeline proved essential for accurate evaluation, improving measured recall by 11–72 percentage points across all models. This finding suggests that prior skill extraction evaluations using exact matching may have substantially underestimated generative model capabilities. The pipeline’s value is particularly pronounced for generative approaches that produce natural language descriptions; for tool-grounded approaches that output exact taxonomy labels, the improvement is smaller but still meaningful.

RQ3: What are the practical implications for engineering curriculum alignment and career guidance? Table 3 summarises our scenario-specific guidance. Each row is framed as a decision for a practitioner: what is being optimised, which configuration is the most natural starting point, and what trade-off to be aware of before deploying it.

5.2. Deployment Decision Guide

Table 3.Scenario-specific starting points for engineering education deployments. The “Trade-off to watch” column flags the failure mode of each choice so practitioners can anticipate where human review will be needed. Recommendations are starting points rather than definitive rankings.
If your goal is… Start with Why Trade-off to watch
Broad coverage of candidate skills for human advisor review (e.g. 1:1 career counselling) ReAct agent 89.7% enhanced recall; finds implicit skills 22.7% enhanced precision; advisor must prune ∼3–5× more candidates than needed
A shorter, higher-confidence list shown directly to students (e.g. automated suggestions) Tool-grounded Gemini 2.5 Flash 51.6% enhanced F1; outputs are valid ESCO entries 15.8s latency; cloud API cost per extraction
Curriculum–industry alignment against ESCO identifiers (database population) Tool-grounded Gemini 2.5 Flash 62.8% exact recall and 38.1% enhanced precision Same latency/cost caveats; novel skills outside ESCO are missed
High-volume batch analysis where downstream review is planned Gemini 2.5 Flash zero-shot 83.5% recall at 3.7s per sample 22.4% enhanced precision; not suitable for student-facing output
Low-latency or interactive tools (e.g. live form feedback) SBERT or ModernBERT 11–16ms per sample; no API cost Enhanced recall 35–49%; will miss implicit skills

5.3. Practical Trade-offs for Engineering Programmes

Engineering educators evaluating these tools can frame the decision around four trade-offs that recur across the scenarios above:

  • Coverage versus specificity. The higher the recall, the more candidates the system returns per input; ReAct’s 89.7% recall is paired with outputs 3–5 times longer than tool-grounded generation. Suitable when a human will review results; less suitable when output is student-facing.

  • Accuracy versus speed. ReAct requires 16 seconds per sample, SBERT 11 milliseconds; interactive applications may need the speed, batch applications can absorb the latency.

  • Accuracy versus cost. API-based systems incur per-request costs that scale with volume; SBERT runs locally without incremental cost. A hybrid pipeline, embedding-based filter followed by generative re-ranking on top candidates, can combine these.

  • Taxonomy alignment versus flexibility. Tool-grounded systems guarantee valid ESCO codes for database integration but cannot capture skills outside the taxonomy; generative systems surface emerging skills but require post-processing to align with ESCO.

In practice, most programmes will combine configurations rather than choose one, using a low-cost embedding filter to triage at scale and reserving a higher-recall or tool-grounded configuration for cases that reach an advisor or a student-facing surface.

5.4. Applications in Engineering Education

To make the deployment choices concrete, consider a programme coordinator preparing for an annual curriculum review. She wants to know which skills her upcoming software engineering capstone projects already develop and which employer-requested skills are under-represented. A practical pipeline would run a low-cost embedding model (SBERT, 11ms per sample) across several hundred scraped job postings to produce a shortlist of high-frequency skills, then pass that shortlist through tool-grounded Gemini Flash to align each candidate with a canonical ESCO identifier suitable for a committee report. For individual student advising later in the term, the advisor might instead opt for ReAct to generate a broader list of candidate skills from a student’s resume, knowing she will prune them in session. The same tools support different decisions; the choice is driven by who will see the output and how much review is available downstream.

Our findings support several such applications:

  • Data-driven curriculum design: Programs can systematically analyse job postings to quantify industry skill demand, identifying which competencies employers prioritise and at what frequency. By comparing these demand patterns against current course offerings, programs can pinpoint gaps, skills frequently requested by employers but underrepresented in curricula. The tool-grounded approach is particularly suitable here, as its high exact recall (62.8%) ensures extracted skills map directly to standardised ESCO identifiers, enabling precise alignment between industry requirements and academic content.

  • Scalable career services: Career centers can offer students personalised skill-gap analyses by extracting skills from their resumes and comparing them against target job descriptions. This enables advisors to provide specific, actionable guidance, such as “your resume demonstrates data visualization but lacks the statistical modelling skills this role requires”, rather than generic career advice. The ReAct agent’s high accuracy (89.7% recall) makes it well-suited for these higher-stakes individual consultations where missing a critical skill gap could affect a student’s job prospects.

  • Accreditation support: Programs must demonstrate that students achieve specific learning outcomes for accreditation bodies such as ABET (ABET, 2024) and Engineers Ireland (Engineers Ireland, 2024). Automated extraction can analyse student work products, capstone reports, project documentation, lab assignments; to map demonstrated competencies against required outcomes. This reduces the manual assessment burden on faculty while enabling programs to systematically track competency development across cohorts and identify where curriculum changes successfully address previously identified gaps.

  • Industry partnership development: Programs can strengthen relationships with industry partners by extracting skills from their job postings and demonstrating concrete curriculum alignment. For example, a program could show a prospective partner that 85% of their required technical competencies are already covered in existing courses. This data-driven approach also identifies collaboration opportunities, skills that partners need but the program doesn’t yet address could inform new course development, guest lectures, or internship projects that benefit both parties.

5.5. Equity and Implementation Considerations

Automated skill extraction supports competency-based engineering education by making skills explicit and measurable (Holmes et al., 2019). AI-powered extraction can process hundreds of syllabi and project descriptions to create comprehensive skill inventories that would be impractical to develop manually, freeing faculty time for curriculum design and student mentoring. For project-based learning, skill extraction can map project outcomes to ESCO competencies, helping students recognise marketable skills developed through academic work. Capstone courses particularly benefit, instructors wherehy, instructors can analyse industry job postings, extract required skills, and ensure projects develop market-relevant competencies. Equity considerations are paramount. First-generation students often lack access to professional networks that provide career intelligence informally. AI-powered tools can democratise access by making skill-career connections explicit and available to all students. However, programs should evaluate extraction accuracy across student populations and supplement AI recommendations with human judgement. We recommend that skill extraction tools supplement rather than replace human expertise. AI excels at processing volume and identifying patterns; human counsellors provide contextual understanding and nuanced judgement that AI cannot replicate.

6. Limitations

Several limitations constrain our results. Our evaluation uses a single English-only benchmark of 926 samples and 1,696 skills that may not fully represent the diversity of engineering domains and document types encountered in practice. Future work should extend evaluation to domain-specific corpora and multilingual contexts. AI capabilities evolve rapidly, and our evaluation reflects the model versions described in §3.3, accessed between December 2025 and January 2026. Subsequent model releases may exhibit different characteristics, though our comparative methodology remains applicable to future evaluations. A larger Llama variant was not tested due to local-GPU memory constraints; comparisons between Llama 3.2 3B and the Gemini 2.5 family should therefore be interpreted with this parameter-count difference in mind. A fundamental challenge lies in the accelerating pace of skill emergence versus taxonomy maintenance. The World Economic Forum estimates that 44% of worker skills will be disrupted by 2028, with skill half-lives declining from 10–15 years to under 5 years, and below 2.5 years for many technical skills (Boston Consulting Group, 2023; World Economic Forum, 2025). IBM estimates that 40% of the global workforce will need reskilling within three years (IBM Institute for Business Value, 2024). Yet standardised taxonomies like ESCO, which rely on our evaluation, update infrequently. ESCO has released only incremental updates since version 1.0 in 2017 (Chiarello et al., 2021). This creates a growing gap: emerging skills in AI, cybersecurity, and sustainability may not yet appear in the taxonomy, meaning extraction systems cannot identify skills that the reference standard does not define. Future work should explore dynamic taxonomy extension or hybrid approaches combining standardised classifications with emerging skill detection. Our recall-focused evaluation reflects educational contexts where comprehensive skill identification matters more than avoiding false positives. Applications requiring high precision would benefit from different evaluation emphasis. We did not compare AI extraction against human expert performance, limiting our ability to contextualise absolute accuracy levels. The enhanced matching pipeline thresholds were selected based on preliminary experiments but may not be optimal for all skill types or domains.

7. Conclusion

This study compares embedding, generative, and agentic AI approaches for automated skill extraction in an engineering education context. Evaluating 11 AI system configurations on the ESCO benchmark (926 samples, 1,696 gold-standard skills), we observe trade-offs that depend on the metric used. On enhanced recall, the ReAct agentic system leads (89.7%), with a 6–14 percentage-point margin over generative LLMs; on enhanced F1, however, tool-grounded Gemini Flash leads (51.6%), because ReAct’s coverage is paired with lower precision (22.7%). For exact taxonomy matching, tool-grounded approaches dominate (Gemini Flash: 62.8%). Our enhanced multi-layer matching pipeline improved measured recall by 11–72 percentage points across all configurations, indicating that exact-matching evaluations understate generative system capabilities for paraphrased ESCO labels.

These results suggest paradigm-specific strengths rather than a single winner. Embedding models are fast but shallow; generative LLMs leverage world knowledge but may misalign with taxonomies; tool-grounded generation constrains outputs to valid ESCO entries at the cost of latency; agentic systems maximise coverage but produce more candidates per input. Each paradigm occupies a distinct position in the precision–recall–latency–cost trade-off space, and the appropriate choice depends on whether the output is surfaced directly to a learner or triaged by a human advisor (§5.2).

For engineering educators, we offer the deployment recommendations in §5.2 as a starting point, subject to the precision requirements and cost constraints of the use case: Gemini Flash zero-shot for high-volume analysis where downstream review is available (83.5% recall at 3.7 seconds), ReAct for broad coverage in human-reviewed workflows (89.7% recall), tool-grounded generation where outputs must align with ESCO identifiers (62.8% exact recall, 51.6% enhanced F1), and embedding models for low-latency or low-cost deployments. Future work should extend the evaluation to multilingual corpora, establish human-expert performance baselines, and capture full reasoning traces to support a deeper error analysis of agentic behaviour. Used with appropriate human oversight, automated skill extraction offers engineering programmes a practical tool for curriculum alignment, career services, and workforce tracking, but the trade-offs documented here caution against treating any single configuration as universally preferred.


AI Tool Disclosure

This research employed several AI tools in different capacities. Google’s Gemini 2.5 Flash, Gemini 2.5 Pro, and Meta’s Llama 3.2 3B-Instruct were used as research subjects in the comparative evaluation; exact API identifiers are listed in §3.3. Their outputs form the primary data analysed in this study. Claude Code (Anthropic) was used in developing the experimental infrastructure, including data processing pipelines, evaluation scripts, and visualization code. All AI-generated code was reviewed, tested, and validated by the research team before use to ensure correctness and reproducibility. Grammarly was used to assist with writing clarity and sentence restructuring throughout the manuscript preparation process. We maintained full responsibility for the study design, interpretation of results, and all conclusions presented in this paper.

Accepted: June 23, 2026 EDT

References

Acemoglu, D., & Restrepo, P. (2022). Tasks, automation, and the rise in U.S. Wage inequality. Econometrica, 90(5), 1973–2016. https:/​/​doi.org/​10.3982/​ECTA19815
Google Scholar
Boston Consulting Group. (2023). Reskilling the workforce for the future. BCG Publications. https:/​/​www.bcg.com/​publications/​2023/​reskilling-workforce-for-future
Google Scholar
Chiarello, F., Fantoni, G., Hogarth, T., Giordano, V., Baltina, L., & Spada, I. (2021). Towards ESCO 4.0 – is the European classification of skills in line with Industry 4.0? A text mining approach. Technological Forecasting and Social Change, 173, 121177. https:/​/​doi.org/​10.1016/​j.techfore.2021.121177
Google Scholar
Clavié, B., & Soulié, G. (2023). Large language models as batteries-included zero-shot ESCO skills matchers. arXiv Preprint. https:/​/​doi.org/​10.48550/​arXiv.2307.03539
Google Scholar
Decorte, J.-J., Van Hautte, J., Deleu, J., Develder, C., & Demeester, T. (2022). Design of negative sampling strategies for distantly supervised skill extraction. Proceedings of the 2nd Workshop on Recommender Systems for Human Resources (RecSys-in-HR 2022), 3218, 26–32. https:/​/​ceur-ws.org/​Vol-3218/​RecSysHR2022-paper_4.pdf
Google Scholar
Decorte, J.-J., Verlinden, S., Van Hautte, J., Deleu, J., Develder, C., & Demeester, T. (2023). Extreme multi-label skill extraction training using large language models. arXiv Preprint. https:/​/​doi.org/​10.48550/​arXiv.2307.10778
Google Scholar
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186. https:/​/​doi.org/​10.18653/​v1/​N19-1423
Google Scholar
Gemini Team, Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., & others. (2024). Gemini: A family of highly capable multimodal models. arXiv Preprint. https:/​/​doi.org/​10.48550/​arXiv.2312.11805
Google Scholar
Green, A. (2024). Artificial intelligence and the changing demand for skills in the labour market. OECD Artificial Intelligence Papers, 14. https:/​/​doi.org/​10.1787/​88684e36-en
Google Scholar
Holmes, W., Bialik, M., & Fadel, C. (2019). Artificial intelligence in education: Promises and implications for teaching and learning. Center for Curriculum Redesign.
Google Scholar
IBM Institute for Business Value. (2024). Augmented work for an automated, AI-driven world. IBM Corporation. https:/​/​www.ibm.com/​thought-leadership/​institute-business-value/​
Kivimäki, I., Panchenko, A., Desber, A., Dhillon, R., Sahlgren, M., & Enne, H. (2013). A graph-based approach to skill extraction from text. Proceedings of the Workshop on Graph-Based Methods for Natural Language Processing, 79–87. https:/​/​doi.org/​10.18653/​v1/​w13-5011
Google Scholar
Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv Preprint. https:/​/​doi.org/​10.48550/​arXiv.1301.3781
Google Scholar
Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3982–3992. https:/​/​doi.org/​10.18653/​v1/​D19-1410
Google Scholar
SeatGeek. (2023). TheFuzz: Fuzzy string matching in Python. https:/​/​github.com/​seatgeek/​thefuzz
Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36, 8634–8652. https:/​/​proceedings.neurips.cc/​paper_files/​paper/​2023/​hash/​1b44b878bb782e6954cd888628510e90-Abstract-Conference.html
Google Scholar
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., & others. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv Preprint. https:/​/​doi.org/​10.48550/​arXiv.2307.09288
Google Scholar
U.S. Department of Labor. (2024). O*NET online. https:/​/​www.onetonline.org/​
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008. https:/​/​doi.org/​10.5555/​3295222.3295349
Google Scholar
Warner, B., Chaffin, A., Clavié, B., Weller, O., & others. (2024). Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference. arXiv Preprint. https:/​/​doi.org/​10.48550/​arXiv.2412.13663
Google Scholar
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824–24837. https:/​/​proceedings.neurips.cc/​paper_files/​paper/​2022/​hash/​9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html
Google Scholar
World Economic Forum. (2025). The future of jobs report 2025. World Economic Forum. https:/​/​www.weforum.org/​publications/​the-future-of-jobs-report-2025/​
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. International Conference on Learning Representations (ICLR). https:/​/​doi.org/​10.48550/​arXiv.2210.03629
Google Scholar
Zhang, M., Jensen, K., Sonniks, S., & Plank, B. (2022). SkillSpan: Hard and soft skill extraction from English job postings. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4962–4984. https:/​/​doi.org/​10.18653/​v1/​2022.naacl-main.366
Google Scholar

A. Illustrative ReAct Interaction Trace

The original experiment pipeline persisted each agent’s final output and message count but did not retain verbatim intermediate reasoning. The trace below is a reconstruction drawn from the agent’s action space, its Gemini 2.5 Pro system prompt, and the final prediction recorded for Sample 0 of the ESCO benchmark. We present it to make the agent’s control flow concrete rather than as a verbatim record. A trace-capturing re-run of a small sample is listed under future work.

Input: “Cypress for frontend is a must have, pytest for backend, and truffle for Solidity is a good addition.”
Gold: Solidity

Cycle 1. Thought. The sentence lists testing tools (Cypress, pytest) and a blockchain language (Solidity). I should check ESCO for entries matching each candidate so I do not invent labels.
Action. search_esco(“Solidity”)
Observation. Top match: Solidity (skill, knowledge type).

Cycle 2. Thought. The target skill is confirmed. The surrounding tools (Cypress, pytest, truffle) are also referenced; ESCO has related knowledge entries for software testing and blockchain development that I can report as candidate knowledge terms even though they are not gold.
Action. search_esco(“software testing”)
Observation. Matches: software testing, automated testing, test automation frameworks.

Terminate. Action. finish(skills=[software testing, automated testing, test planning, ..., using truffle], knowledge=[Cypress, pytest, truffle, Solidity, ..., blockchain development])

This reconstruction illustrates two behaviours visible in the aggregate results: (i) the agent’s verification step correctly pins Solidity to an ESCO entry (a precision benefit), and (ii) the agent continues searching adjacent concepts even after the gold skill is confirmed, which is the source of the 25-candidate output reported in Case 2 of Table 2.