Evaluating Modern Tools and Large Language Models for Digitising Sea-Level Data from Marigrams
Numerous graphical data digitisation tools exist; however, few have been applied to the recovery of historical sea-level marigrams (tidal charts). Using an early twentieth-century marigram from the Clyde Estuary, Scotland, containing one week of sea level observations, we evaluated the performance of 11 digitisation tools and 2 large language models (LLMs) in digitising marigrams. A multi-user testing framework was used to assess the accuracy and speed of each approach. Both LLMs tested, ChatGPT5.5 and Gemini 3.1Pro, were unsuitable for this application under the tested conditions, because they show errors exceeding 10% of the observed tidal range. Conventional digitisation tools performed substantially better, most digitisation tools successfully recovered sea-level data with errors below 1%. Despite this accuracy, these tools remained dependent on substantial human input. Even the fastest approach, DigitGraph, required on average 29 minutes to recover one week of sea-level data. Scaled to a century-long record, this would require more than one working year to complete, making application to large datasets challenging. Consequently, further development of dedicated, automated digitisation algorithms remains necessary.
Authors
- Joanne Williams (ORCID: https://orcid.org/0000-0002-8421-4481)
- Duo Chan
- Christian Kenwright
- Lily Sharp
- Callum Slade
- Patrick Sharpe
- Minnie J. Darby
- Sam E. Pearson-Smith
- Steve McFarland
- Robert J. Nicholls
- Nurul Tazaroh
- Seyedeh Fardis Pourreza Ahmadi
- Molly A. Phillips
- Ivan. D. Haigh
- Marc Becker
Publication Details
- Published
- 2026-09-28
- DOI
- https://doi.org/10.31223/x5p51q
- Primary Topic
- Oceanographic and Atmospheric Processes
- Type
- preprint