Natural-language processing (NLP) and indexing, Part 3. Semantic search
This article explores asymmetric extractive semantic search, using the large language model (LLM) SBERT. It first examines traditional exact, keyword, and fuzzy match search algorithms, discussing their features and limitations. It then explores how semantic search leverages the capacity of an LLM to capture query intent and context beyond literal word matches. An open-source application demonstrates asymmetric extractive semantic search on a public–domain PDF on Ancestral Pueblo ruins in Arizona. Sample queries on site architecture, burial practices, and areal comparisons return relevant passages, illustrating the method’s accuracy and utility to indexers. Appendices detailing the architectures of BERT and SBERT are provided. You shall know a word by the company it keeps. John Rupert Firth, linguist (1957)
Authors
- Donald Howes
Publication Details
- Journal
- The Indexer
- Published
- 2026-09-21
- DOI
- https://doi.org/10.3828/index.2026.30
- Primary Topic
- Language and cultural evolution
- Type
- article
- Field-Weighted Citation Impact
- 0.00