
Yi He · Discover Artificial Intelligence 2026 · 2026
DOI: 10.1007/s44163-026-02232-w
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Aligning job advertisements with occupational and educational taxonomies is normally treated as a single retrieval problem: a piece of vacancy text is encoded and matched against a reference inventory. We show that this uniform treatment is the wrong default, because the three constructs a posting contains differ in how their evidence is distributed. A skill is fixed by a short phrase, an occupation by a title or a whole sentence, and a qualification by an ordered level rather than by any particular string. We therefore propose carl , a construct-aware retrieval and linking layer whose single design principle is that every linking decision should be conditioned on the construct being linked. carl adds three modules to a standard recognizer-plus-encoder pipeline and changes no backbone: Granularity Fusion combines several query widths with weights fitted per construct instead of committing to one width, Boundary Arbitration lets the retrieval score rather than the tagger alone decide where a mention ends, and Ordinal Aggregation replaces flat nearest-neighbour search over qualification strings with pooled evidence over ordered levels together with a fitted rule for abstaining on mentions outside the inventory. Evaluated against the European Qualifications Framework and the ESCO inventory of occupations and skills, carl improves rank-one accuracy by 0.081 on qualifications and 0.040 on skills over the strongest fixed pipeline, both significant under paired testing, while the occupation gain of 0.033 is directional, exactly as our granularity analysis predicts. Ablations show each module contributing where its motivating measurement said it should, and an error budget, a human expert benchmark and repeated runs of five decoder language models place the gains in context: the layer recovers 7 to 14 percent of the addressable error, and the remaining gap is large.
No comments yet — start the discussion below.