An Empirical Study on the Evaluation of Graph‐Based Strategies for Intelligent Modeling Assistants
ABSTRACT Context In the context of model‐driven engineering, the proliferation of intelligent modeling assistants (IMAs) has pushed the boundaries of automation, providing modelers with relevant recommendations based on incomplete modeling artifacts as primary input. Despite their benefits, developing IMAs has also brought a set of challenges that can compromise the quality of the recommendations, spanning from selecting the proper knowledge base to training the most suitable algorithm. In the scope of this paper, we focus on graph‐based techniques that have been proposed to enhance IMAs by exploiting the relationships among modeling artifacts. Althoughthese systems achieve notable accuracy on real‐world datasets, there is a need to understand how their performance differs with respect to different features of the input datasets, for example, the size of the graphs, the number of available artifacts, or their categories. Method We report our experiences in evaluating IMAs by comparing two state‐of‐the‐art graph‐based IMAs. First, we elicit a list of steps and issues that occur in IMAs according to our practical experience. Subsequently, we select two approaches: (i) BORA, a reuse‐oriented assistant exploiting SPARQL queries and N‐gram encoding over RDF graphs, and (ii) MORGAN, a prediction‐oriented assistant leveraging graph kernel similarity. Finally, we evaluate the accompanying tools of the two approaches by using three distinct datasets from different application domains. Results Our findings show that BORA obtains better results when heterogeneous datasets are considered, while homogeneous ones should be used to feed ML‐based techniques such as MORGAN. This finding is directly connected to the nature of the two approaches, as reuse‐oriented approaches can handle heterogeneity more than traditional machine learning approaches. Conclusion We eventually report a list of lessons learned for researchers and practitioners, highlighting the need for a careful data selection process.
Authors
- Davide Di Ruscio (ORCID: https://orcid.org/0000-0002-5077-6793)
- Manuel Wimmer (ORCID: https://orcid.org/0000-0002-1124-7098)
- Ilirian Ibrahimi
- Claudio Di Sipio (ORCID: https://orcid.org/0000-0001-9872-9542)
Institutions
- Johannes Kepler University of Linz (AT)
- University of L'Aquila (IT)
- Deutsches Bergbau-Museum Bochum (DE)
Publication Details
- Journal
- Software Practice and Experience
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1002/spe.70115
- Primary Topic
- Model-Driven Software Engineering Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00