Taxonomy-Constrained Semantic Alignment For Large-Scale Recruitment Data Structuring
Large-scale recruitment data have emerged as a valuable source for studying labor demand, educational planning, and regional talent dynamics. However, the analytical value of such data hinges on the ability to transform raw job postings into stable, structured records that are suitable for downstream analysis. In practice, recruitment data exhibit substantial noise, redundancy, weak structure, and semantic ambiguity, particularly in the implicit expression of academic-major requirements. The central challenge in many downstream applications is therefore not merely to extract surface terms from job advertisements, but to align implicit and heterogeneous job requirements with standardized academic taxonomies under well-defined constraints. This paper formulates this task as taxonomy-constrained semantic alignment and presents Taxonomy-Constrained Semantic Alignment (TCSA), a three-stage framework that combines education normalization, multi-stage deduplication, and taxonomyconstrained LLM-based alignment. The design embodies a principle of progressive uncertainty reduction: normalization establishes the analytical scope and standardizes heterogeneous representations; deduplication eliminates redundant evidence while preserving legitimate reposting behavior; and semantic alignment maps job descriptions to a closed set of major labels under education-level-specific taxonomies. Experiments conducted on 59 million raw recruitment records demonstrate that the proposed pipeline substantially reduces data redundancy and enables structured major prediction with strong ranking quality. On a new 600-pair education-stratified evaluation at τ = 0.8, the deduplication module achieves population-weighted precision of 0.740, recall of 0.910, and F1 of 0.817, with five unresolved pairs reported separately, while, across five temperature-0.7 runs, the alignment module achieves population-standardized Precision@3 of 0.670±0.009 and MRR of 0.858±0.008 on 485 evaluable records with a multi-annotator majority-vote gold standard (Fleiss’ κ = 0.657 on shared LLM predictions; union-based κ = 0.449). The resulting cleaned dataset further enables educationstratified descriptive analyses that reveal distinct demand patterns across junior college, undergraduate, master, and doctoral levels.
Authors
- Lingling Zhou
- Yi Zheng
- Jie Luo (ORCID: https://orcid.org/0009-0003-0224-4535)
Publication Details
- Journal
- International Journal of Software Engineering and Knowledge Engineering
- Published
- 2026-09-26
- DOI
- https://doi.org/10.1142/s0218194026500907
- Primary Topic
- Recommender Systems and Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00