Aziida Nanyonga, Hassan Wasswa, Uğur Turhan, KEITH FRANCIS JOINER, Graham Wild · arXiv (Cornell University) 2026 · 2026
DOI: 10.48550/arxiv.2610.04472
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Aviation accident and incident investigations generate extensive unstructured textual information containing evidence relevant to the causes and contributing factors of safety occurrences. Automatically extracting such information is challenging because causal evidence may be distributed across long and complex investigation narratives. This study proposes a semantic causal-factor inference framework combining natural language processing with a variational autoencoder (VAE) to learn the relationship between aviation investigation narratives and expert-reported probable causes. Investigation narratives and their corresponding probable causes are transformed into numerical representations, after which the encoder maps narrative representations to a probabilistic latent space. The decoder estimates representations of the corresponding probable causes and is trained using an objective that combines Kullback-Leibler divergence with cosine-similarity-based semantic reconstruction. The framework was evaluated using 20,919 finalized U.S. National Transportation Safety Board investigation reports from 2005 to 2020. On the held-out test set, the predicted and expert-reported probable-cause representations achieved a mean cosine similarity of 0.786 (SD = 0.120). The predicted representations also yielded interpretable terms associated with causal information in the reports. The results demonstrate the potential of probabilistic latent representation learning for AI-assisted extraction of causal information from aviation safety narratives while retaining expert investigation as the basis for formal causal determination.
No comments yet — start the discussion below.