Muhammad Maroof · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23000490
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Overview Large language models fine-tuned for text-to-SQL generation are typically evaluated and deployed assuming GPU inference, which limits their use in resource-constrained or on-premise settings where only CPU hardware is available. This work makes two contributions: Fine-tuning Phi-4-mini-instruct (3.8B params, dense decoder-only Transformer with grouped-query attention) on the QuerySmith-Spider-Bird dataset (ajayk007/querysmith-spider-bird, derived from the Spider and BIRD text-to-SQL benchmarks) to produce a compact, competitive text-to-SQL model. A progression of CPU quantization strategies — culminating in a novel partial-residual quantization scheme — that makes this model practical to run on CPU-only hardware with minimal accuracy loss. Method Fine-tuning: QLoRA (4-bit base weights + LoRA adapters, rank 16, α=32, dropout 0.05) on the query/key/value/output projections, 3 epochs, lr 2×10-4, micro-batch 1, gradient accumulation ×8 across 2 GPUs (effective batch size 16). CPU optimization, evaluated in five stages: Precision sweep: FP32 → FP16 → INT8 → INT4 baselines Quantization strategies: selective (MLP-only), partial (last-k blocks only), residual (INT8 + low-rank SVD error correction), and the combined partial-residual approach Inference engine: conversion to GGUF and benchmarking via llama.cpp at F16/Q8_0/Q4_0 Thread scaling: sweep across {1, 2, 4, 6, 8, 12, 16} CPU threads Post-optimization fine-tuning: re-fine-tuning the optimized model on the same SQL task to recover any remaining accuracy gap Hardware: fine-tuning on a 2-GPU setup; CPU benchmarks on a 14-core Intel Core i5 (13th Gen), 16 GB RAM. Evaluation: 200 held-out examples, scored by exact-match SQL accuracy (primary metric), with execution accuracy against live databases used to confirm exact-match wasn't penalizing semantically-equivalent query variants. Key Results Config Accuracy ms/token Size (MB) FP32 (baseline) 0.713 148.0 14890.5 FP16 0.709 121.4 7445.0 INT8 0.682 89.1 3721.8 INT4 (naive) 0.636 71.1 1859.9 Partial-Residual (ours) 0.707 79.3 3511.2 Partial-Residual + llama.cpp Q4_0 0.707 27.5 — + post-optimization fine-tune 0.743 22.3 3511.2 Naive INT4 cuts model size 87.5% vs. FP32 but costs 7.7 accuracy points. Partial-residual quantization recovers 92.2% of that accuracy loss (0.707 vs. 0.636), landing within 0.006 of the FP32 baseline. Deploying through llama.cpp adds a further 65.3% latency reduction on top of the quantization strategy gains — the two are multiplicative, not redundant. An 8–12 thread operating point is optimal on the test hardware; gains plateau and slightly regress beyond that. Re-fine-tuning the optimized model exceeds the original FP32 baseline (0.743 vs. 0.713, +4.2% relative), while running 6.6× faster and at 4.2× smaller size. Why Partial-Residual Works Selective and partial quantization alone each recover only part of the INT4→FP32 gap; residual quantization alone does better by directly correcting quantization error. Combining the two — confining aggressive quantization to a subset of transformer blocks and correcting the resulting error with a low-rank term — outperforms either alone, consistent with quantization error being concentrated in specific blocks rather than spread uniformly across the network. Limitations The rank-8 residual term was fixed, not swept — larger models or harder tasks may need a higher rank. The low-rank SVD residual adds per-layer compute overhead in raw PyTorch (offsetting some of INT4's raw speed advantage), though this overhead disappears once converted to GGUF and run through llama.cpp. The 200-example evaluation set is modest in size; a larger split would tighten confidence in the reported figures. Code & Reproducibility Benchmark scripts covering all five stages (precision sweep, quantization strategies, GGUF conversion, thread sweep, fine-tuning, and results aggregation) are included in this upload.
No comments yet — start the discussion below.