Deepthi R. S., Bindushree U., Amrutha Raghupathy, Shivani N. · International Journal of Innovative Science and Research Technology (IJISRT) 2026 · 2026
DOI: 10.38124/ijisrt/26sep311
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
PictoVoice is an image-to-speech system designed to assist users in understanding visual information through spoken descriptions. The system accepts an image as input and processes it to extract relevant visual features, which are then used to generate a meaningful textual caption. The generated caption is subsequently converted into speech, enabling users to understand the contents of an image through audio output. The system integrates image processing, visual feature extraction, image captioning, and speech generation into a unified pipeline. Multilingual support for languages is currently being implemented to improve accessibility for users from diverse linguistic backgrounds.
No comments yet — start the discussion below.