Building networks of shared research themes in an academic unit by word co-occurrence fingerprinting
Departments within a university are not only administrative units, but also an effort to gather investigators around common fields of academic study. Identifying shared research interests among members is a pervasive challenge. Here I describe a workflow that uses natural language processing to generate a network connecting 𝑛 = 7 9 members of a university department, or multiple departments within a faculty ( 𝑛 = 2 7 8 ), based on common topics in their research publications. After extracting and processing terms from 𝑛 = 1 6 , 9 4 8 abstracts in the PubMed database, the co-occurrence of terms is encoded in a sparse document-term matrix. Based on the angular distances between the presence-absence vectors for every pair of terms, I use the uniform manifold approximation and projection (UMAP) method to embed the terms into a representation space such that terms that tend to appear in the same documents are closer together. Each author’s corpus defines a probability distribution over terms in this space. Using the Wasserstein distance to quantify the similarity between these distributions, I generate a distance matrix among authors that can be analyzed and visualized as a graph. I demonstrate that the distribution of edges in the graph relating members of a faculty are significantly associated with departmental and research centre affiliations, while identifying untapped connections among members.
Authors
- Art F. Y. Poon (ORCID: https://orcid.org/0000-0003-3779-154X)
Institutions
- Western University (CA)
Publication Details
- Journal
- Journal of Informetrics
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1016/j.joi.2026.101882
- Primary Topic
- scientometrics and bibliometrics research
- Type
- article
- Field-Weighted Citation Impact
- 0.00