1. Introduction
The engineering workforce is undergoing a profound transformation driven by technological advancement and evolving industry demands. According to the World Economic Forum (World Economic Forum, 2025), 59% of the global workforce will require training by 2030, with engineering fields particularly affected by the emergence of artificial intelligence and automation (Acemoglu & Restrepo, 2022; Green, 2024). This creates a fundamental challenge for engineering education: how can academic programs ensure their curricula align with the rapidly changing skills that employers require? Traditional approaches to curriculum alignment, employer advisory boards, alumni surveys, and periodic program reviews, provide valuable insights but cannot keep pace with the velocity of change in technical fields. Natural language processing (NLP) techniques offer a promising solution by automatically extracting skill mentions from job postings, resumes, and syllabi (Zhang et al., 2022). When mapped to standardised taxonomies like the European Skills, Competences, Qualifications and Occupations (ESCO) classification (see Section 2.2), which contains 13,939 distinct skills, these extractions can provide actionable intelligence for curriculum design and career guidance. The potential applications are substantial: curriculum committees could analyse job postings to identify emerging requirements, career services could provide personalised skill-gap analyses, and accreditation processes could systematically map student work to required competencies. However, comprehensive comparisons of modern AI approaches are lacking. Prior research has evaluated individual methods in isolation: Bidirectional Encoder Representations from Transformers (BERT), a language model that understands words in context by reading text in both directions simultaneously, much like how a human reads surrounding sentences to determine whether “Java” refers to a programming language or an island. BERT-based embeddings convert text into numerical representations that capture meaning, so phrases like “Python programming” and “Python development” are recognised as similar (Reimers & Gurevych, 2019), generative LLMs (Clavié & Soulié, 2023), or agentic systems (Yao et al., 2023). Engineering educators facing technology adoption decisions lack guidance on which approach offers the best trade-offs, that is why we undertook this research. This study addresses this gap through systematic evaluation of 11 model configurations across three paradigms: embedding-based approaches (SBERT, ModernBERT), generative LLMs (Gemini 2.5 Flash/Pro, Llama 3.2), and agentic AI systems (ReAct) using Gemini 2.5 Pro as the base model. We introduce an enhanced multi-layer matching pipeline employing eight progressive strategies to account for skill paraphrasing, a critical methodological contribution given that models may identify semantically correct skills using different surface forms than the gold-standard labels. For example, if the gold-standard label is “develop software” but the model predicts “software development,” both refer to the same skill, yet strict exact-matching evaluation would mark this as incorrect. Our pipeline recognises such equivalent expressions, providing a fairer assessment of model capabilities. This research addresses the following research questions:
-
How do modern AI approaches compare for automated skill extraction from text?
-
How effective is an enhanced multi-layer matching pipeline for improving skill recognition accuracy?
-
What are the practical implications for engineering curriculum alignment and career guidance?
The remainder of this paper is organised as follows: Section 2 reviews related work in skill extraction, taxonomies, and AI approaches. Section 3 describes our methodology, including the 11 AI configurations and the enhanced matching pipeline. Section 4 presents results, and Section 5 discusses implications including practical deployment recommendations for engineering programs. We address limitations in Section 6 and conclude in Section 7.
2. Literature Review
2.1. Evolution of Skill Extraction
Skill extraction from text has evolved through several distinct phases. Early approaches relied on rule-based systems and keyword matching using hand-crafted lexicons (Kivimäki et al., 2013), but suffered from limited coverage and required continuous manual updates, with new emerging skills. Word embeddings (Mikolov et al., 2013) enabled semantic matching beyond exact strings, but struggled with multi-word expressions and context-dependent meanings.
Transformer-based models (Vaswani et al., 2017) like BERT (Devlin et al., 2019) marked a significant advancement by capturing contextual relationships. Researchers treated skill extraction as named entity recognition or semantic similarity tasks (Decorte et al., 2022; Zhang et al., 2022). To evaluate these approaches, the SkillSpan benchmark (Zhang et al., 2022) provided important data distinguishing between hard skills (technical competencies like “Python programming”) and soft skills (interpersonal abilities like “communication” or “teamwork”). However, evaluation revealed that soft skills remain particularly challenging to extract, even with transformer-based models, due to their diverse expression patterns, for instance, “works well with others,” “team player,” and “collaborative” all convey similar competencies but use completely different phrasing.
LLMs introduced zero-shot extraction capabilities (Clavié & Soulié, 2023; Decorte et al., 2023), enabling skill identification without task-specific training, meaning models can extract skills from any text without needing labelled examples as a prerequisite. This flexibility is valuable for engineering education, where skill requirements span diverse domains from software development to biomedical engineering. However, these approaches face challenges including hallucination (the generation of plausible but incorrect skills, such as inventing “advanced quantum networking” when the text only mentions basic network configuration), and difficulty aligning outputs with specific taxonomies. For instance, an LLM might output “coding” when the ESCO taxonomy specifies the precise label “develop software applications,” making it difficult to aggregate results or integrate with existing databases.
Most recently, agentic AI systems have emerged as a promising paradigm, combining LLM reasoning with external tool use rather than generating outputs in a single pass (Yao et al., 2023). The ReAct (Reasoning + Acting) framework exemplifies this approach by interleaving reasoning traces, where the model explicitly thinks through the problem, with actions like database lookups, then observing results before continuing. For skill extraction, this means an agent can reason that “managing cross-functional teams” likely implies leadership skills, query the ESCO database to verify this inference, and refine its output based on what it finds. Rather than inventing skill names, the agent validates its inferences against actual taxonomy entries, ensuring outputs align with standardised classifications. Crucially, this grounding in authoritative databases reduces hallucination. The approach also enables detection of implicit skills that simpler methods miss, as the agent can reason about competencies underlying described activities. Reflexion (Shinn et al., 2023) extends this paradigm with self-correction capabilities, enabling agents to recognise errors and adjust their reasoning within a single task. Despite these advances, comparisons of agentic approaches against embedding and generative baselines for skill extraction remain sparse in the published literature we surveyed. This study contributes one such comparison, framed around an engineering education deployment context.
2.2. Skill Taxonomies for Engineering Education
Standardised skill taxonomies provide essential infrastructure for automated skill extraction by defining the target label space and enabling cross-document comparison. Without a common vocabulary, extracted skills from different sources, such as job postings, resumes, and syllabi, cannot be meaningfully aggregated or compared. For instance, one document might mention “programming,” another “software development,” and a third “coding”; a standardised taxonomy maps these to a single canonical skill, enabling quantitative analysis across documents. The European Skills, Competences, Qualifications and Occupations (ESCO) taxonomy (European Commission 2023b, 2023a) represents the most comprehensive multilingual skill classification currently available, developed by the European Commission to support labour market transparency. ESCO contains 2,942 occupations and 13,939 skills organised hierarchically, spanning knowledge areas, technical competencies, and transversal skills like communication and teamwork. Crucially, ESCO’s explicit skill-occupation links enable direct mapping between educational outcomes and career pathways, for example, identifying which specific skills a student needs to develop for their target occupation. The O*NET database (U.S. Department of Labor, 2024) provides a US-focused alternative maintained by the Department of Labor, but lacks ESCO’s international applicability and multilingual support. Other taxonomies were considered: the Lightcast Open Skills library offers broader coverage ( 33,000 skills) but is primarily a commercial product with limited occupation-skill linkages, while the European e-Competence Framework (e-CF) provides well-defined competences but is restricted to the ICT sector. We selected ESCO due to its comprehensive coverage, hierarchical organisation enabling analysis at multiple granularity levels, and growing adoption in workforce development applications worldwide.
2.3. AI Approaches for Skill Extraction
Embedding-based approaches represent text as dense numerical vectors where semantically similar content appears nearby in the vector space, meaning “Python programming” and “Python development” would be positioned close together, enabling matching without exact string overlap. Sentence-BERT (SBERT) (Reimers & Gurevych, 2019) extended BERT for efficient sentence-level embeddings, while ModernBERT (Warner et al., 2024) offers architectural improvements for faster inference. For skill extraction, these models compute similarity between input text and pre-computed embeddings of all taxonomy skills, returning the closest matches. These approaches excel at processing speed (milliseconds per sample) but achieve lower accuracy than generative methods, particularly for implicit skills, since they can only match what is explicitly stated, not reason about competencies underlying described activities."
Generative LLMs extract skills through natural language generation, producing skill lists as text output rather than matching against pre-computed vectors. Models like Gemini (Gemini Team et al., 2024) and Llama (Touvron et al., 2023) can generate extractions without task-specific fine-tuning by leveraging knowledge acquired during pre-training, the initial phase where models learn language patterns and world knowledge from massive text corpora before being deployed for specific tasks. Different prompting strategies affect performance: zero-shot provides only task instructions, few-shot includes example extractions to guide the model, and chain-of-thought (Wei et al., 2022) encourages step-by-step reasoning before generating outputs. Tool-grounded approaches augment LLMs with database access, allowing the model to look up and validate skills against the taxonomy during generation. While flexible and capable of identifying implicit skills through reasoning, generative approaches have higher latency (seconds rather than milliseconds) and may produce outputs misaligned with target taxonomies, such as generating “coding” when the taxonomy requires “develop software applications.”.
Agentic AI systems represent the most recent paradigm, combining LLM reasoning with autonomous tool use and multi-step planning employed rather than the generation of all outputs in a single pass. The ReAct framework (Yao et al., 2023) interleaves reasoning traces (where the model explicitly thinks through the problem), with actions like database lookups, it then observes the results before continuing. For example, when processing text describing “led cross-functional team meetings,” an agent might reason this implies leadership and coordination skills, query the ESCO database to find matching entries, and refine its output based on actual taxonomy labels. This grounded approach reduces hallucination by validating inferences against authoritative databases rather than inventing skill names, while also enabling identification of implicit skills that embedding and basic generative approaches miss.
3. Methodology
3.1. Research Design
We conducted a comparative study evaluating 11 AI system configurations (see Table 1) across three paradigms, namely embedding-based, generative LLM, and agentic AI on a standardised benchmark derived from the ESCO taxonomy. Our experimental design prioritises recall over precision, though both metrics are valuable depending on the application context. Recall measures the proportion of actual skills correctly identified, essentially asking “of all the skills that exist in the text, how many did we find?” Precision measures what proportion of predicted skills were actually correct, asking “of everything we suggested, how much was right?” We prioritised recall for this study because educational contexts typically benefit from comprehensive skill identification. Missing relevant skills can be more consequential than suggesting additional skills for consideration, a curriculum analysis tool that fails to identify important skill gaps could leave students underprepared for the workforce, whereas suggesting extra skills gives educators more options to review and filter. Applications with different requirements, such as automated credential verification where false positives carry serious consequences, would appropriately prioritise precision instead. All experiments used identical input data and evaluation procedures to ensure fair comparison. Gemini systems were accessed via Google’s cloud API, while Llama was run locally on an NVIDIA RTX 4080 GPU; this infrastructure difference affects latency comparisons but not accuracy metrics. We measured both exact recall (requiring precise string matches) and enhanced recall (allowing semantically equivalent matches through our multi-layer pipeline), along with processing latency to characterise speed-accuracy trade-offs.
3.2. Dataset
Our evaluation uses the Skill Extraction benchmark (Decorte et al., 2022), an extension of the SkillSpan dataset (Zhang et al., 2022) with ESCO skill annotations. The benchmark comprises 926 job posting text samples containing 1,696 gold-standard skill annotations drawn from the 13,939-skill ESCO classification v1.1.0. Unlike typical Named Entity Recognition (NER) datasets, which identify and label specific items in text, for example recognising “Python” as a skill and “Google” as an organisation in a job posting, but without linking these to taxonomies, this benchmark provides direct mappings between text spans and standardised ESCO skill identifiers-, enabling rigorous evaluation of taxonomy-grounded extraction rather than mere mention detection.
The dataset combines three complementary subsets: TECH (technology sector job postings), HOUSE (household services), and TECHWOLF (diverse industries), together representing the variety of skill-bearing texts encountered in educational and workforce contexts. Each sample contains one to five gold-standard skills spanning domains such as software development, data science, mechanical engineering, and project management. Importantly, the benchmark includes both explicit skills (directly stated, e.g., “Python programming”) and implicit skills (competencies underlying described activities, e.g., “collaborate with stakeholders” implying communication skills), providing a rigorous test of extraction capabilities beyond simple keyword matching.
3.3. AI Systems Evaluated
We evaluated 11 configurations (Table 1) across three distinct AI paradigms:
Embedding-Based Systems: We tested SBERT using the sentence-transformers/all-MiniLM-L6-v2 checkpoint, a 22M-parameter general-purpose sentence encoder, and ModernBERT using the answerdotai/ModernBERT-base checkpoint in zero-shot configuration. Both systems computed cosine similarity between input text embeddings and pre-computed embeddings for all 13,939 ESCO skills, returning the highest-similarity matches.
Generative LLM Systems: We tested Google’s Gemini 2.5 Flash (API identifier gemini-2.5-flash) in three configurations: zero-shot, few-shot with ten examples, and tool-grounded with ESCO database access. Gemini 2.5 Pro (gemini-2.5-pro) was evaluated in two configurations: chain-of-thought, and few-shot chain-of-thought combining examples with reasoning. All Gemini calls used temperature 0 for deterministic output and were made via Google’s cloud API between 9 December 2025 and 1 January 2026. Meta’s Llama 3.2 3B-Instruct, served via the llama3.2 tag in Ollama (3.21B parameters, Q4_K_M quantisation, 128k-token context window), was evaluated in three configurations: zero-shot, few-shot, and tool-grounded, and run locally on an NVIDIA RTX 4080 GPU. Each configuration used carefully designed prompts specifying the extraction task and desired output format.
Agentic System: We implemented a ReAct (Yao et al., 2023) agent using Gemini 2.5 Pro as the reasoning backbone, with access to an ESCO skill lookup tool. The agent’s architecture has three components. First, the state accumulates the input text, the evolving reasoning trace, the history of tool calls, and the current list of candidate skills and knowledge terms; the state is passed to the LLM on every step. Second, the action space exposes two operations: search_esco(query), which performs a semantic lookup against the 13,939-skill ESCO index and returns the top- matches with their canonical labels, and finish(skills, knowledge), which terminates the loop and returns the final extraction. Third, the control loop runs the LLM on the current state, parses the next action, executes it, appends the observation to the state, and repeats. Termination occurs when the agent emits finish or when the step budget of 20 reasoning–action cycles is reached; empirically the agent terminates after 1–3 reasoning–action cycles on most samples (mean 1.7, median 2 across a 100-sample subset with full trace metadata; see Figure 1). The agent is prompted to identify both explicit and implicit skills and to validate inferences through ESCO lookups before committing to them. We make no claim that this specific configuration is optimal; our goal is to situate one representative agentic system alongside embedding and generative baselines so that the paradigm’s trade-offs can be measured rather than asserted.
3.4. Enhanced Multi-Layer Matching Pipeline
A fundamental challenge in evaluating skill extraction is that AI systems may predict semantically correct skills using different surface forms than the gold-standard labels. For instance, a system predicting “software development” when the gold label is “develop software” has identified the correct skill but would receive no credit under exact string matching.
We address this limitation through an enhanced multi-layer matching pipeline employing eight progressive strategies, applied in order from most to least strict:
-
Exact Match: Case-insensitive string equality between prediction and gold label
-
Substring Match: Either string contains the other as a substring
-
Fuzzy Partial: TheFuzz library (SeatGeek, 2023) partial ratio score 85
-
Fuzzy Token Sort: TheFuzz token-sorted comparison score 85
-
Fuzzy Token Set: TheFuzz token-set comparison score 85
-
Sequence Matcher: Python difflib SequenceMatcher ratio 0.85
-
Jaccard Similarity: Word-level overlap coefficient 0.7
-
Cosine Similarity: SBERT embedding similarity 0.85
Each predicted skill is evaluated against all gold labels using these strategies in sequence, with the first successful match recorded. This approach provides nuanced assessment of system capabilities while maintaining rigorous matching standards.
3.5. Evaluation Metrics
We report precision, recall, and F1 under both exact and enhanced matching. Exact metrics use normalized string equality against ESCO labels and represent performance for applications requiring precise taxonomy alignment. Enhanced metrics use the eight-layer matching pipeline described above and represent performance for applications where semantic equivalence suffices. Average latency per sample is also reported to support deployment decisions.
We treat recall as the primary metric for engineering education use cases, with precision and F1 reported alongside it for deployment choice and for transparency regarding the precision–recall trade-off analysed in §4.4. This emphasis reflects the downstream cost profile in our target applications: in curriculum mapping, career guidance, and skills gap analysis, a missed skill can lead a student away from a relevant pathway or omit it from advising entirely, whereas a surplus candidate is typically filtered at low cost by a human advisor or a taxonomy-aware post-processing step before the result reaches the student. Recall therefore governs the worst-case outcome (a pathway the student never sees), while precision governs a reviewable one (a candidate that can be removed). We nonetheless report precision and F1 throughout because the appropriate balance depends on whether the output is surfaced directly to a learner or triaged by an advisor; the scenario-specific deployment recommendations in §5.2 use both.
4. Results
4.1. Overall Performance Comparison
Table 1 presents precision, recall, and F1 for all 11 AI system configurations, under both exact and enhanced matching. Reporting precision alongside recall surfaces a trade-off that a recall-only view obscures: the ReAct agentic system achieved the highest enhanced recall (89.7%), but its enhanced precision of 22.7% is the second-lowest among generative and agentic configurations. Tool-grounded Gemini Flash achieved the highest enhanced F1 (51.6%) by combining moderate recall (79.9%) with substantially higher precision (38.1%). We return to this pattern in §4.4.
Several patterns emerge. On enhanced F1, the top four configurations are all generative systems rather than the agentic ReAct, because ReAct’s coverage advantage is partly offset by lower precision. On enhanced recall, ReAct remains the clear leader, with a 6–14 percentage-point margin over Gemini and Llama configurations. Tool-grounded methods dominate exact precision and exact recall because they constrain outputs to valid ESCO entries rather than generating free-form skill names. Embedding models occupy a distinct regime: SBERT achieves competitive exact precision (13.0%) because its nearest-neighbour retrieval returns canonical ESCO strings, but its enhanced recall (35.5%) is far below generative systems.
Latency spans four orders of magnitude: embedding models process samples in 11–16 milliseconds, Gemini 2.5 configurations require 2.6–7.6 seconds via cloud API, and the ReAct agentic approach requires approximately 16 seconds per sample. Llama 3.2 3B latencies (4–58 seconds) reflect local execution on consumer hardware (RTX 4080) rather than optimised cloud infrastructure, so direct latency comparisons with Gemini are not appropriate; accuracy comparisons remain valid.
Few-shot prompting did not consistently improve over zero-shot on enhanced recall, though it substantially improved precision (Gemini 2.5 Flash: 22.4% 36.8%; Llama 3.2 3B: 9.1% 31.7%), suggesting examples primarily help constrain output vocabulary. Among generative systems, Gemini 2.5 models consistently outperformed Llama 3.2 3B across configurations.
4.2. Impact of Enhanced Matching Pipeline
Figure 2 illustrates the substantial impact of our multi-layer matching pipeline. Enhanced matching improved measured recall by 11–72 percentage points across all configurations.
The most dramatic improvement occurred for Gemini Pro chain-of-thought (70.9 percentage points), suggesting chain-of-thought prompting produces varied skill descriptions matching gold labels semantically but not lexically. This finding has important methodological implications: evaluations using exact matching only substantially underestimate generative system capabilities.
4.3. Speed-Accuracy Trade-offs
Figure 3 presents the latency comparison. The four-order-of-magnitude difference between embedding models (11–16ms) and tool-grounded approaches (15,000–58,000ms) creates fundamentally distinct deployment scenarios. Note that Llama’s higher latencies (up to 58 seconds for tool-grounded) reflect local GPU execution rather than optimised cloud infrastructure. Among cloud-based systems, Gemini Flash zero-shot offers an attractive middle ground: 83.5% enhanced recall at 3.7 seconds per sample.
4.4. Precision–Recall Trade-off
Because recall alone cannot distinguish a system that finds the right skills from one that simply returns a longer list, we examined how each configuration balances precision against recall. Figure 4 plots all 11 configurations in this space, with iso-F1 contours for reference. Three distinct regimes are visible.
ReAct sits in the top-left: its 89.7% enhanced recall is paired with 22.7% enhanced precision, indicating that a non-trivial share of its coverage advantage comes from producing more candidates per sample. For example, on a single-sentence input naming Solidity as the gold skill, the ReAct agent returned 12 candidate skills and 13 candidate knowledge terms (25 outputs total). It matched the gold correctly, but also introduced many unrelated items. ReAct’s gains should therefore be interpreted as higher coverage with lower specificity, not as uniformly higher quality.
Tool-grounded Gemini Flash reached the highest enhanced F1 in our evaluation (51.6%): by constraining outputs to valid ESCO entries it raised precision to 38.1% while retaining 79.9% recall. Few-shot and chain-of-thought generative configurations cluster just below this point. Embedding models occupy a lower-recall regime but remain competitive on precision: SBERT’s 38.4% enhanced precision is comparable to the top generative systems, reflecting that nearest-neighbour retrieval returns canonical taxonomy strings rather than paraphrased outputs.
These results modify the interpretation of recall-only rankings. If the goal is exhaustive coverage of candidate skills for downstream human review (for example, broad curriculum mapping), ReAct’s higher recall is the relevant metric. If the goal is to produce a shorter, higher-confidence list of matches (for example, automated career suggestions shown directly to students), tool-grounded generation offers a more favourable balance. We return to these scenario-specific recommendations in §5.2.
4.5. Analysis of Agentic System Behaviour
The ReAct agent’s high recall stems from three mechanisms: (1) implicit skill detection: reasoning about skills implied by activities even when not explicitly mentioned; (2) validation through database lookup: querying ESCO to find actual entries matching inferences; and (3) self-correction: adjusting reasoning when lookups return unexpected results. These benefits come at latency cost: 1–3 reasoning–action cycles per sample, contributing to the 16-second average processing time (Table 1). As §4.4 notes, the same permissive generation that enables implicit-skill coverage also produces the lower precision observed in Table 1.
To ground this pattern, Table 2 shows four representative sample-level outputs drawn directly from our saved predictions. Case 1 illustrates how ReAct identifies an implicit skill that an embedding baseline misses entirely, returning noisy but usable output for fuzzy-matched ESCO terms. Case 2 shows the over-extraction pattern: a single-gold input yields 25 ReAct candidates while tool-grounded Gemini Flash returns the correct ESCO entry plus one near-miss. Case 3 shows the same pattern on a longer input. Case 4 shows a generative zero-shot configuration producing a 39-term candidate list for a one-gold input, where most candidates are near-duplicates generated from the same surface expressions in the text. Across all four cases, the correct gold skill is recovered by the higher-recall systems, but the ratio of signal to noise varies sharply by configuration.
5. Discussion
5.1. Addressing Research Questions
RQ1: How do modern AI approaches compare for automated skill extraction from text? The answer depends on which metric is prioritised. On enhanced recall, the ReAct agentic configuration leads, with 89.7% recall exceeding the best generative configuration (Gemini Pro chain-of-thought, 83.7%) by 6 percentage points and the best embedding configuration (ModernBERT, 49.1%) by more than 40. On enhanced F1, however, the leader is tool-grounded Gemini Flash (51.6%), because ReAct’s coverage advantage is paired with lower precision (22.7%) (see §4.4 and Table 1). These results indicate paradigm-specific strengths rather than a single dominant approach, and the interpretation of “best” therefore depends on whether coverage, precision, or a balance of the two is most valuable for the application.
For applications requiring alignment with ESCO taxonomy labels, such as generating standardised reports or populating databases, tool-grounded approaches provide the highest exact recall (62.8% for Gemini Flash) we observed. For applications that tolerate broader candidate lists intended for downstream human review, the ReAct agent offers the highest enhanced recall at a cost of 16 seconds per sample and lower precision. For real-time applications where latency dominates, embedding models provide millisecond responses with lower recall but competitive enhanced precision.
RQ2: How effective is an enhanced multi-layer matching pipeline for improving skill recognition accuracy? The multi-layer matching pipeline proved essential for accurate evaluation, improving measured recall by 11–72 percentage points across all models. This finding suggests that prior skill extraction evaluations using exact matching may have substantially underestimated generative model capabilities. The pipeline’s value is particularly pronounced for generative approaches that produce natural language descriptions; for tool-grounded approaches that output exact taxonomy labels, the improvement is smaller but still meaningful.
RQ3: What are the practical implications for engineering curriculum alignment and career guidance? Table 3 summarises our scenario-specific guidance. Each row is framed as a decision for a practitioner: what is being optimised, which configuration is the most natural starting point, and what trade-off to be aware of before deploying it.
5.2. Deployment Decision Guide
5.3. Practical Trade-offs for Engineering Programmes
Engineering educators evaluating these tools can frame the decision around four trade-offs that recur across the scenarios above:
-
Coverage versus specificity. The higher the recall, the more candidates the system returns per input; ReAct’s 89.7% recall is paired with outputs 3–5 times longer than tool-grounded generation. Suitable when a human will review results; less suitable when output is student-facing.
-
Accuracy versus speed. ReAct requires 16 seconds per sample, SBERT 11 milliseconds; interactive applications may need the speed, batch applications can absorb the latency.
-
Accuracy versus cost. API-based systems incur per-request costs that scale with volume; SBERT runs locally without incremental cost. A hybrid pipeline, embedding-based filter followed by generative re-ranking on top candidates, can combine these.
-
Taxonomy alignment versus flexibility. Tool-grounded systems guarantee valid ESCO codes for database integration but cannot capture skills outside the taxonomy; generative systems surface emerging skills but require post-processing to align with ESCO.
In practice, most programmes will combine configurations rather than choose one, using a low-cost embedding filter to triage at scale and reserving a higher-recall or tool-grounded configuration for cases that reach an advisor or a student-facing surface.
5.4. Applications in Engineering Education
To make the deployment choices concrete, consider a programme coordinator preparing for an annual curriculum review. She wants to know which skills her upcoming software engineering capstone projects already develop and which employer-requested skills are under-represented. A practical pipeline would run a low-cost embedding model (SBERT, 11ms per sample) across several hundred scraped job postings to produce a shortlist of high-frequency skills, then pass that shortlist through tool-grounded Gemini Flash to align each candidate with a canonical ESCO identifier suitable for a committee report. For individual student advising later in the term, the advisor might instead opt for ReAct to generate a broader list of candidate skills from a student’s resume, knowing she will prune them in session. The same tools support different decisions; the choice is driven by who will see the output and how much review is available downstream.
Our findings support several such applications:
-
Data-driven curriculum design: Programs can systematically analyse job postings to quantify industry skill demand, identifying which competencies employers prioritise and at what frequency. By comparing these demand patterns against current course offerings, programs can pinpoint gaps, skills frequently requested by employers but underrepresented in curricula. The tool-grounded approach is particularly suitable here, as its high exact recall (62.8%) ensures extracted skills map directly to standardised ESCO identifiers, enabling precise alignment between industry requirements and academic content.
-
Scalable career services: Career centers can offer students personalised skill-gap analyses by extracting skills from their resumes and comparing them against target job descriptions. This enables advisors to provide specific, actionable guidance, such as “your resume demonstrates data visualization but lacks the statistical modelling skills this role requires”, rather than generic career advice. The ReAct agent’s high accuracy (89.7% recall) makes it well-suited for these higher-stakes individual consultations where missing a critical skill gap could affect a student’s job prospects.
-
Accreditation support: Programs must demonstrate that students achieve specific learning outcomes for accreditation bodies such as ABET (ABET, 2024) and Engineers Ireland (Engineers Ireland, 2024). Automated extraction can analyse student work products, capstone reports, project documentation, lab assignments; to map demonstrated competencies against required outcomes. This reduces the manual assessment burden on faculty while enabling programs to systematically track competency development across cohorts and identify where curriculum changes successfully address previously identified gaps.
-
Industry partnership development: Programs can strengthen relationships with industry partners by extracting skills from their job postings and demonstrating concrete curriculum alignment. For example, a program could show a prospective partner that 85% of their required technical competencies are already covered in existing courses. This data-driven approach also identifies collaboration opportunities, skills that partners need but the program doesn’t yet address could inform new course development, guest lectures, or internship projects that benefit both parties.
5.5. Equity and Implementation Considerations
Automated skill extraction supports competency-based engineering education by making skills explicit and measurable (Holmes et al., 2019). AI-powered extraction can process hundreds of syllabi and project descriptions to create comprehensive skill inventories that would be impractical to develop manually, freeing faculty time for curriculum design and student mentoring. For project-based learning, skill extraction can map project outcomes to ESCO competencies, helping students recognise marketable skills developed through academic work. Capstone courses particularly benefit, instructors wherehy, instructors can analyse industry job postings, extract required skills, and ensure projects develop market-relevant competencies. Equity considerations are paramount. First-generation students often lack access to professional networks that provide career intelligence informally. AI-powered tools can democratise access by making skill-career connections explicit and available to all students. However, programs should evaluate extraction accuracy across student populations and supplement AI recommendations with human judgement. We recommend that skill extraction tools supplement rather than replace human expertise. AI excels at processing volume and identifying patterns; human counsellors provide contextual understanding and nuanced judgement that AI cannot replicate.
6. Limitations
Several limitations constrain our results. Our evaluation uses a single English-only benchmark of 926 samples and 1,696 skills that may not fully represent the diversity of engineering domains and document types encountered in practice. Future work should extend evaluation to domain-specific corpora and multilingual contexts. AI capabilities evolve rapidly, and our evaluation reflects the model versions described in §3.3, accessed between December 2025 and January 2026. Subsequent model releases may exhibit different characteristics, though our comparative methodology remains applicable to future evaluations. A larger Llama variant was not tested due to local-GPU memory constraints; comparisons between Llama 3.2 3B and the Gemini 2.5 family should therefore be interpreted with this parameter-count difference in mind. A fundamental challenge lies in the accelerating pace of skill emergence versus taxonomy maintenance. The World Economic Forum estimates that 44% of worker skills will be disrupted by 2028, with skill half-lives declining from 10–15 years to under 5 years, and below 2.5 years for many technical skills (Boston Consulting Group, 2023; World Economic Forum, 2025). IBM estimates that 40% of the global workforce will need reskilling within three years (IBM Institute for Business Value, 2024). Yet standardised taxonomies like ESCO, which rely on our evaluation, update infrequently. ESCO has released only incremental updates since version 1.0 in 2017 (Chiarello et al., 2021). This creates a growing gap: emerging skills in AI, cybersecurity, and sustainability may not yet appear in the taxonomy, meaning extraction systems cannot identify skills that the reference standard does not define. Future work should explore dynamic taxonomy extension or hybrid approaches combining standardised classifications with emerging skill detection. Our recall-focused evaluation reflects educational contexts where comprehensive skill identification matters more than avoiding false positives. Applications requiring high precision would benefit from different evaluation emphasis. We did not compare AI extraction against human expert performance, limiting our ability to contextualise absolute accuracy levels. The enhanced matching pipeline thresholds were selected based on preliminary experiments but may not be optimal for all skill types or domains.
7. Conclusion
This study compares embedding, generative, and agentic AI approaches for automated skill extraction in an engineering education context. Evaluating 11 AI system configurations on the ESCO benchmark (926 samples, 1,696 gold-standard skills), we observe trade-offs that depend on the metric used. On enhanced recall, the ReAct agentic system leads (89.7%), with a 6–14 percentage-point margin over generative LLMs; on enhanced F1, however, tool-grounded Gemini Flash leads (51.6%), because ReAct’s coverage is paired with lower precision (22.7%). For exact taxonomy matching, tool-grounded approaches dominate (Gemini Flash: 62.8%). Our enhanced multi-layer matching pipeline improved measured recall by 11–72 percentage points across all configurations, indicating that exact-matching evaluations understate generative system capabilities for paraphrased ESCO labels.
These results suggest paradigm-specific strengths rather than a single winner. Embedding models are fast but shallow; generative LLMs leverage world knowledge but may misalign with taxonomies; tool-grounded generation constrains outputs to valid ESCO entries at the cost of latency; agentic systems maximise coverage but produce more candidates per input. Each paradigm occupies a distinct position in the precision–recall–latency–cost trade-off space, and the appropriate choice depends on whether the output is surfaced directly to a learner or triaged by a human advisor (§5.2).
For engineering educators, we offer the deployment recommendations in §5.2 as a starting point, subject to the precision requirements and cost constraints of the use case: Gemini Flash zero-shot for high-volume analysis where downstream review is available (83.5% recall at 3.7 seconds), ReAct for broad coverage in human-reviewed workflows (89.7% recall), tool-grounded generation where outputs must align with ESCO identifiers (62.8% exact recall, 51.6% enhanced F1), and embedding models for low-latency or low-cost deployments. Future work should extend the evaluation to multilingual corpora, establish human-expert performance baselines, and capture full reasoning traces to support a deeper error analysis of agentic behaviour. Used with appropriate human oversight, automated skill extraction offers engineering programmes a practical tool for curriculum alignment, career services, and workforce tracking, but the trade-offs documented here caution against treating any single configuration as universally preferred.
AI Tool Disclosure
This research employed several AI tools in different capacities. Google’s Gemini 2.5 Flash, Gemini 2.5 Pro, and Meta’s Llama 3.2 3B-Instruct were used as research subjects in the comparative evaluation; exact API identifiers are listed in §3.3. Their outputs form the primary data analysed in this study. Claude Code (Anthropic) was used in developing the experimental infrastructure, including data processing pipelines, evaluation scripts, and visualization code. All AI-generated code was reviewed, tested, and validated by the research team before use to ensure correctness and reproducibility. Grammarly was used to assist with writing clarity and sentence restructuring throughout the manuscript preparation process. We maintained full responsibility for the study design, interpretation of results, and all conclusions presented in this paper.



