Taxonomy-Constrained Semantic Alignment For Large-Scale Recruitment Data Structuring

Large-scale recruitment data have emerged as a valuable source for studying labor demand, educational planning, and regional talent dynamics. However, the analytical value of such data hinges on the ability to transform raw job postings into stable, structured records that are suitable for downstream analysis. In practice, recruitment data exhibit substantial noise, redundancy, weak structure, and semantic ambiguity, particularly in the implicit expression of academic-major requirements. The central challenge in many downstream applications is therefore not merely to extract surface terms from job advertisements, but to align implicit and heterogeneous job requirements with standardized academic taxonomies under well-defined constraints. This paper formulates this task as taxonomy-constrained semantic alignment and presents Taxonomy-Constrained Semantic Alignment (TCSA), a three-stage framework that combines education normalization, multi-stage deduplication, and taxonomyconstrained LLM-based alignment. The design embodies a principle of progressive uncertainty reduction: normalization establishes the analytical scope and standardizes heterogeneous representations; deduplication eliminates redundant evidence while preserving legitimate reposting behavior; and semantic alignment maps job descriptions to a closed set of major labels under education-level-specific taxonomies. Experiments conducted on 59 million raw recruitment records demonstrate that the proposed pipeline substantially reduces data redundancy and enables structured major prediction with strong ranking quality. On a new 600-pair education-stratified evaluation at τ = 0.8, the deduplication module achieves population-weighted precision of 0.740, recall of 0.910, and F1 of 0.817, with five unresolved pairs reported separately, while, across five temperature-0.7 runs, the alignment module achieves population-standardized Precision@3 of 0.670±0.009 and MRR of 0.858±0.008 on 485 evaluable records with a multi-annotator majority-vote gold standard (Fleiss’ κ = 0.657 on shared LLM predictions; union-based κ = 0.449). The resulting cleaned dataset further enables educationstratified descriptive analyses that reveal distinct demand patterns across junior college, undergraduate, master, and doctoral levels.

Authors

Publication Details

Journal
International Journal of Software Engineering and Knowledge Engineering
Published
2026-09-26
DOI
https://doi.org/10.1142/s0218194026500907
Primary Topic
Recommender Systems and Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Taxonomy-Constrained Semantic Alignment For Large-Scale Recruitment Data Structuring

Lingling Zhou, Yi Zheng, Jie Luo
International Journal of Software Engineering and Knowledge Engineering
Recommender Systems and Techniques
article

Taxonomy-Constrained Semantic Alignment For Large-Scale Recruitment Data Structuring

Lingling Zhou, Yi Zheng, Jie Luo
article en

Abstract

Large-scale recruitment data have emerged as a valuable source for studying labor demand, educational planning, and regional talent dynamics. However, the analytical value of such data hinges on the ability to transform raw job postings into stable, structured records that are suitable for downstream analysis. In practice, recruitment data exhibit substantial noise, redundancy, weak structure, and semantic ambiguity, particularly in the implicit expression of academic-major requirements. The central challenge in many downstream applications is therefore not merely to extract surface terms from job advertisements, but to align implicit and heterogeneous job requirements with standardized academic taxonomies under well-defined constraints. This paper formulates this task as taxonomy-constrained semantic alignment and presents Taxonomy-Constrained Semantic Alignment (TCSA), a three-stage framework that combines education normalization, multi-stage deduplication, and taxonomyconstrained LLM-based alignment. The design embodies a principle of progressive uncertainty reduction: normalization establishes the analytical scope and standardizes heterogeneous representations; deduplication eliminates redundant evidence while preserving legitimate reposting behavior; and semantic alignment maps job descriptions to a closed set of major labels under education-level-specific taxonomies. Experiments conducted on 59 million raw recruitment records demonstrate that the proposed pipeline substantially reduces data redundancy and enables structured major prediction with strong ranking quality. On a new 600-pair education-stratified evaluation at τ = 0.8, the deduplication module achieves population-weighted precision of 0.740, recall of 0.910, and F1 of 0.817, with five unresolved pairs reported separately, while, across five temperature-0.7 runs, the alignment module achieves population-standardized Precision@3 of 0.670±0.009 and MRR of 0.858±0.008 on 485 evaluable records with a multi-annotator majority-vote gold standard (Fleiss’ κ = 0.657 on shared LLM predictions; union-based κ = 0.449). The resulting cleaned dataset further enables educationstratified descriptive analyses that reveal distinct demand patterns across junior college, undergraduate, master, and doctoral levels.

International Journal of Software Engineering and Knowledge Engineering
Decent work and economic growth
Openalex Percentile: Top 4%
Recommender Systems and Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.