Joshua Nishanth Tarun A, Joel Ajitesh Varun A · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23144481
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Cue is a compact causal conversation-control model for real-time cascaded voice agents. Unlike conventional voice-activity detection and silence-based endpointing, Cue explicitly models conversational events such as backchannels, accidental sounds, soft interruptions, hard interruptions, takeovers, and turn completion, and maps them to operational actions including CONTINUE, PAUSE, STOP, and RESPOND. Cue v5 uses a shared frozen Whisper-small encoder and a causal Transformer controller, with optional features from the other speaker's audio. The model is trained on approximately 125 hours of synthetic Indian-English recruiter calls and real AMI/ICSI meeting speech using a leakage-aware training protocol designed to prevent the model from exploiting future reply information unavailable to a live voice agent. A key result of the study is the measured synthetic-to-real gap. The earlier synthetic-only Cue v4 model detects only 24.3% of interruptions on held-out AMI meetings under a 10% false-positive-rate constraint. Cue v5 increases interruption recall to 61.4% on AMI and 63.1% on ICSI. Bootstrap confidence intervals are 0.585–0.638 for AMI and 0.574–0.703 for ICSI. End-of-turn prediction on natural meetings remains more difficult, highlighting the difference between controlled synthetic endpointing and real conversational floor management. On the TurnBench development set, without using TurnBench training data, Cue v5 reaches 0.804 interruption recall at 0.079 false-positive rate and 547 ms median delay. A Cue speech-detection configuration combined with a cross-validated silence rule reaches 0.800 end-of-turn recall at 0.070 false-positive rate, while the learned Cue end-of-turn head alone reaches 0.686 recall at 0.081 false-positive rate. The paper also evaluates Cue Tiny, a 6.2M-parameter distilled CPU model, and discusses causal streaming evaluation, acoustic robustness, dual-audio conditioning, synthetic-to-real transfer, deployment trade-offs, and limitations. This record contains the Cue v5 Preprint v1.0. Model weights, runtime code, and reproducibility artifacts are intended to be released separately.
No comments yet — start the discussion below.