Sanjib Shah, Aman Sheikh, Sunil Nath, Nabin Paudel, Suyog Lamsal · Proceedings of International Conference on Innovation in Computing Science Engineering and Technology 2026 · 2026
DOI: 10.65091/icicset.v3i1.104
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Retrieving data from enterprise relational databasesrequires SQL. Many business users do not have SQL skills,which drives the need for natural language to SQL systems.Most leading systems depends on large, cloud-hosted LLMs,which sits poorly with commercially sensitive ERP data andtheir surrounding retrieval and validation pipelines are rarelyevaluated separately from the generation model itself. This articleinvestigates whether a small locally deployed LLM can competeas a generation model in the same retrieval augmented, rulevalidated agent pipeline as larger models, against which theretrieval and validation pipelines also evaluated with the data.We built an agent for an on-premise Oracle-based Synergy ERPschema, using BGE-M3 embeddings over Qdrant vector databaseto index schema and example retrieval, and a deterministicvalidator with bounded self-correction before execution. We useand compare four interchangeable LLMs: locally run Qwen2.5-Coder-14B (4bit quantization) and cloud run Mistral Large,Minimax M3 and Qwen3.8-Max based on semantic equivalence,execution accuracy, component-level F1, row-count matching,and generation success rate. Qwen3.8-Max performed best onevery metrics, while MiniMax M3 outperformed the largerMistral Large. The much smaller Qwen2.5-Coder-14B matchedor exceeded Mistral Large on several metrics. Schema retrievalwas reliable enough with Hit@5 = 0.93. An ablation comparingbefore and after validation and SQL correction shows the largestaccuracy gains for weaker models, with limited or slightlynegative effect for the strongest model. The locally hosted modelwas also the fastest (2.671 s/query) and cheapest ($0.0011/query)among the four models. These findings suggest model scaledoes not fully determine performance within a fixed pipeline,supporting compact local models as a practical, low-cost, lowlatencyalternative for enterprise deployment.
No comments yet — start the discussion below.