Yao Hu, Qian Huang · Innovation Discovery 2026 · 2026
DOI: 10.53964/id.2026014
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Modern data are increasingly represented and utilized as interconnected networks, including collaboration graphs, multi-relational user–item interactions, and schema-less knowledge graphs that support retrieval-augmented generation pipelines. While graph analytics has become a key foundation for intelligent data mining, existing approaches often focus on limited aspects of the problem in isolation. For example, similarity-based methods effectively capture local structures but perform poorly on dynamic networks; random-walk-based embeddings model graph topology but overlook rich node semantics; and large-scale graph databases (GDBs) face significant computational challenges in handling relationship-intensive traversals. In this paper, we propose a unified three-tier, graph-driven framework for scalable intelligent data mining that integrates community-aware link prediction, multiplex random-walk-based node embedding, and parallel knowledge reasoning. At the topology level, we formulate link prediction in dynamic networks as a learning task enhanced by community detection to enrich quasi-local structural features. At the representation level, we develop a flexible in-memory random-walk embedding model operating on a multiplex User–Item (MUI) graph, where both intra-layer and inter-layer transitions are unified through a single transition matrix whose size is independent of the number of layers. At the cognition level, we design a parallel GDB-based engine that combines random-walk embeddings with message-passing techniques to enable scalable knowledge reasoning over schema-less, real-world datasets. Extensive experiments conducted on diverse datasets demonstrate the effectiveness of the proposed framework. The results show a high AUCPR for dynamic link prediction, near-perfect node classification accuracy under appropriate parameter settings, and significant parallel speedups for random-walk-based reasoning, highlighting the advantages of a unified graph-driven approach to intelligent data mining. The three tiers are coupled through a shared pipeline, a shared codebase, and a joint evaluation protocol rather than through end-to-end joint training; we state this scope explicitly and discuss what a tighter, jointly optimized coupling across the three tiers would require.
No comments yet — start the discussion below.