Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large language models (LLMs) develop a model of natural language through the large scale consumption of natural language corpora. This thesis explores the development of LLMs that model the language of structured data to enable interoperability with other computational systems. This requires an ability to mediate between structured data and natural language to generate arguments for tool use in JSON format. Existing models with this capability are integrated into a job skill extraction pipeline, as well as a LoRA fine-tuned model developed for this task, Laura from Compliance (LfC). This model was fine-tuned on generations of the model Mistral-7B-Instruct guided by a finite-state machine (FSM). Evaluated on a talent acquisition dataset from Siemens AG, LfC required 83% fewer tokens and half the inference time to achieve performance comparable to the model on which it was fine-tuned. This thesis therefore presents a methodology for the optimisation of an LLM for the use of a singular tool in a data science pipeline.
No comments yet — start the discussion below.