Guorui Wang, Jiang Jiang, Jinchen Ma, Dingxiao Liu · Tsinghua Science & Technology 2026 · 2026
DOI: 10.26599/tst.2026.9010081
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Power-system customer service requires answering user queries by jointly interpreting chart images, textual instructions, and high-precision temporal measurements. Existing multimodal large language models are effective at visual-language interaction but remain weak at exact temporal reasoning and reliability-aware decision support. We propose PowerMTQA, a dual-stream multimodal temporal question-answering framework that combines a Qwen-based visual-language backbone with a dedicated multichannel time-series branch, prototype reprogramming, explicit cross-attention fusion, and auxiliary anomaly/risk supervision. Experiments on a unified dataset built from four public electricity and power-related time-series sources show that explicit temporal modeling is critical: dual-stream variants reduce the validation generative loss from 24.5611 for a vision-only baseline to 0.0809 in the single-channel set-ting, while multichannel modeling and LoRA further im-prove the best loss to 0.0721 with 0.9788 token accuracy. On answer-level semantic consistency evaluated by an external judge, PowerMTQA reaches 0.82 accuracy, outperforming the visual-only and text-only baselines by 15 and 39 percentage points, respectively. These results indicate that combining chart understanding with structured temporal reasoning substantially improves factual consistency and practical reliability in power-service question answering.
No comments yet — start the discussion below.