Taha Mohseni Ahooyi, Benjamin J. Stear, Yuanchao Zhang, Aditya Lahiri, James Terry, Shiping Zhang, J. Alan Simmons, Ryan Corbett, Patricia Sullivan, Asif Chinwalla, Chris Nemarich, Jo Lynne Rokita, Sharon Diskin, Jonathan C. Silverstein, Deanne M. Taylor · bioRxiv (Cold Spring Harbor Laboratory) 2026 · 2026
DOI: 10.64898/2026.09.28.754397
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Biomedical knowledge graphs can connect information across genes, phenotypes, tissues, pathways, experiments, and clinical resources, but they are difficult to query correctly without detailed knowledge of the graph. Large language models can help write Cypher code for knowledge graphs, yet a query focused on a bioinformatics task that looks reasonable may still use the wrong identifier, relationship direction, source, intermediate node, or output unit if the user is not completely trained on the schema We developed ddkg.skill, an Agent Skill that can be loaded into a compatible LLM session to support querying of the NIH Common Fund Data Ecosystem Data Distillery Knowledge Graph (DDKG), a biomedical property graph integrating more than 180 ontologies, over 40 genomics datasets, and dozens of additional cross-domain biomedical datasets. Rather than utilizing a simple markdown prompt, ddkg.skill contains a controller, a compositional library of DDKG-specific references and structured tables, primary documentation, validated query patterns, a routing table, and a Python script that checks the internal links among these materials. The skill build for the December 2025 DDKG release contains 38 bundled files and 239 routing relationships and is identified by checksum so users can state the exact build used to compose a query. The evaluation reported here was performed separately on an earlier build with 38 files and 227 routing entries. In that orthogonal nine-test evaluation against a live December 2025 DDKG instance, five of seven tests targeting sources absent from the worked examples produced correct executed results. The evaluation also exposed a failed query, a cross-source comparison that was not biologically well posed, and errors in the skill's own reference material that informed later revisions. The skill guides an AI through entity resolution, graph inspection, query construction, and validation for biomedical and bioinformatics queries, while keeping the DDKG itself as the source of returned results. This design provides a portable and versioned method for giving general-purpose AI systems practical knowledge of a complex biomedical knowledge graph to empower complex bioinformatics data integration. Three executed biomedical use cases further demonstrate cross-species phenotype-to-expression querying, genomic overlap between 4DN chromatin-loop anchors and GTEx eQTLs, and integration of congenital heart phenotype associations with ClinVar and GTEx heart evidence.
No comments yet — start the discussion below.