An Empirical Study on the Evaluation of Graph‐Based Strategies for Intelligent Modeling Assistants

ABSTRACT Context In the context of model‐driven engineering, the proliferation of intelligent modeling assistants (IMAs) has pushed the boundaries of automation, providing modelers with relevant recommendations based on incomplete modeling artifacts as primary input. Despite their benefits, developing IMAs has also brought a set of challenges that can compromise the quality of the recommendations, spanning from selecting the proper knowledge base to training the most suitable algorithm. In the scope of this paper, we focus on graph‐based techniques that have been proposed to enhance IMAs by exploiting the relationships among modeling artifacts. Althoughthese systems achieve notable accuracy on real‐world datasets, there is a need to understand how their performance differs with respect to different features of the input datasets, for example, the size of the graphs, the number of available artifacts, or their categories. Method We report our experiences in evaluating IMAs by comparing two state‐of‐the‐art graph‐based IMAs. First, we elicit a list of steps and issues that occur in IMAs according to our practical experience. Subsequently, we select two approaches: (i) BORA, a reuse‐oriented assistant exploiting SPARQL queries and N‐gram encoding over RDF graphs, and (ii) MORGAN, a prediction‐oriented assistant leveraging graph kernel similarity. Finally, we evaluate the accompanying tools of the two approaches by using three distinct datasets from different application domains. Results Our findings show that BORA obtains better results when heterogeneous datasets are considered, while homogeneous ones should be used to feed ML‐based techniques such as MORGAN. This finding is directly connected to the nature of the two approaches, as reuse‐oriented approaches can handle heterogeneity more than traditional machine learning approaches. Conclusion We eventually report a list of lessons learned for researchers and practitioners, highlighting the need for a careful data selection process.

Authors

Institutions

Publication Details

Journal
Software Practice and Experience
Published
2026-10-06
DOI
https://doi.org/10.1002/spe.70115
Primary Topic
Model-Driven Software Engineering Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

An Empirical Study on the Evaluation of Graph‐Based Strategies for Intelligent Modeling Assistants

Davide Di Ruscio, Manuel Wimmer, Ilirian Ibrahimi, Claudio Di Sipio
Software Practice and Experience
Model-Driven Software Engineering Techniques
article

An Empirical Study on the Evaluation of Graph‐Based Strategies for Intelligent Modeling Assistants

Davide Di Ruscio, Manuel Wimmer, Ilirian Ibrahimi, Claudio Di Sipio
article en

Abstract

ABSTRACT Context In the context of model‐driven engineering, the proliferation of intelligent modeling assistants (IMAs) has pushed the boundaries of automation, providing modelers with relevant recommendations based on incomplete modeling artifacts as primary input. Despite their benefits, developing IMAs has also brought a set of challenges that can compromise the quality of the recommendations, spanning from selecting the proper knowledge base to training the most suitable algorithm. In the scope of this paper, we focus on graph‐based techniques that have been proposed to enhance IMAs by exploiting the relationships among modeling artifacts. Althoughthese systems achieve notable accuracy on real‐world datasets, there is a need to understand how their performance differs with respect to different features of the input datasets, for example, the size of the graphs, the number of available artifacts, or their categories. Method We report our experiences in evaluating IMAs by comparing two state‐of‐the‐art graph‐based IMAs. First, we elicit a list of steps and issues that occur in IMAs according to our practical experience. Subsequently, we select two approaches: (i) BORA, a reuse‐oriented assistant exploiting SPARQL queries and N‐gram encoding over RDF graphs, and (ii) MORGAN, a prediction‐oriented assistant leveraging graph kernel similarity. Finally, we evaluate the accompanying tools of the two approaches by using three distinct datasets from different application domains. Results Our findings show that BORA obtains better results when heterogeneous datasets are considered, while homogeneous ones should be used to feed ML‐based techniques such as MORGAN. This finding is directly connected to the nature of the two approaches, as reuse‐oriented approaches can handle heterogeneity more than traditional machine learning approaches. Conclusion We eventually report a list of lessons learned for researchers and practitioners, highlighting the need for a careful data selection process.

Software Practice and Experience
Johannes Kepler University of Linz (AT), University of L'Aquila (IT), Deutsches Bergbau-Museum Bochum (DE)
Openalex Percentile: Top 3%
Model-Driven Software Engineering Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.