A Record Structure Reconstruction Algorithm for Industrial Spreadsheet Data Based on Structural Exception Patterns
Spreadsheets prepared in industrial sites often contain partially altered record structures because repeated values are omitted, distributed, or combined to reduce input effort. This paper defines these structural inconsistencies as a record structure reconstruction problem rather than a structure extraction problem, and proposes a reconstruction algorithm based on structural exception patterns. Seven structural exception types were identified from industrial data. Among them, partial-information-based record generation, split-column dependency, and multi-value combined cells were selected as representative patterns because they directly affect record-level reconstruction. The proposed algorithm extracts cell positions, empty-cell information, delimiters, and key attributes from the initial structuring result, and applies pattern-specific reconstruction rules. Missing attributes are supplemented, distributed values are merged into a single record, and combined values are separated into individual fields. The algorithm was evaluated using 67,582 industrial records against manually prepared ground truth data. The record reconstruction accuracy increased from 37.23% to 73.82%, and structural errors decreased from 42,419 to 17,694, corresponding to a 58.29% reduction. The results show that the proposed algorithm reduces structural errors and generates record-level data from industrial spreadsheet data containing distributed records, omitted attributes, and combined values.
Authors
- WonYeong Song
- HoJin Hwang
Publication Details
- Journal
- Korean Journal of Computational Design and Engineering
- Published
- 2026-09-10
- DOI
- https://doi.org/10.7315/cde.2026.267
- Primary Topic
- Spreadsheets and End-User Computing
- Type
- article
- Field-Weighted Citation Impact
- 0.00