A semantic layer architecture for removing schema identifiers from external LLM prompts, with model-dependent accuracy trade-offs

Abstract The adoption of large language model (LLM)-based AI systems in credit rating agencies introduces a structural tension between operational efficiency and data security. This study examines a Semantic Layer as a technical mechanism for reducing this tension, and shows that whether it removes designated schema identifiers from the LLM-facing prompt depends on an implementation detail invisible to a schema-level analysis: whether the concept-to-SQL mapping is itself withheld from the LLM. An initial implementation (v1) includes this mapping in the prompt and removes none of the designated schema identifiers from the LLM-facing payload; a corrected implementation (v2) withholds it entirely and removes all of them, verified exhaustively over the complete payload. We scope this security claim explicitly: it concerns the removal of designated schema identifiers from the prompt, not a general reduction in sensitive-information exposure, and we specify the associated threat model and its out-of-scope leakage channels. Across two benchmarks (a purpose-built credit-rating benchmark designed by the author alongside the Semantic Layer specification, and the externally-authored BIRD financial benchmark) and three LLM tiers spanning a capability range, we find that accuracy is model-dependent rather than uniformly preserved: on the author-designed credit-rating benchmark v2 improves with model capability and, at the strongest tier, is numerically highest among the three conditions (60.7% vs. 58.7% for v1 and 52.0% for Text-to-SQL), although with $$n=60$$ questions this strongest-tier gap does not reach statistical significance, whereas on the externally-authored benchmark v2 trails both baselines at every tier, with the in-benchmark-versus-external gap widening as capability increases. We characterize this pattern as selective encapsulation , presented as an empirical design lesson rather than a predictive theory. Rather than extending Parnas’s (1972) Information Hiding principle, we empirically demonstrate that it is easily violated at LLM interfaces and requires payload-level verification. We also contribute the first domain-specific NL-to-SQL benchmark for credit-rating contexts, together with an execution-based self-check of its gold queries performed by the benchmark designer.

Authors

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-10-07
DOI
https://doi.org/10.1007/s44163-026-02147-6
Primary Topic
Privacy-Preserving Technologies in Data
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A semantic layer architecture for removing schema identifiers from external LLM prompts, with model-dependent accuracy trade-offs

Munil Yang
Discover Artificial Intelligence
Privacy-Preserving Technologies in Data
article

A semantic layer architecture for removing schema identifiers from external LLM prompts, with model-dependent accuracy trade-offs

Munil Yang
article en

Abstract

Abstract The adoption of large language model (LLM)-based AI systems in credit rating agencies introduces a structural tension between operational efficiency and data security. This study examines a Semantic Layer as a technical mechanism for reducing this tension, and shows that whether it removes designated schema identifiers from the LLM-facing prompt depends on an implementation detail invisible to a schema-level analysis: whether the concept-to-SQL mapping is itself withheld from the LLM. An initial implementation (v1) includes this mapping in the prompt and removes none of the designated schema identifiers from the LLM-facing payload; a corrected implementation (v2) withholds it entirely and removes all of them, verified exhaustively over the complete payload. We scope this security claim explicitly: it concerns the removal of designated schema identifiers from the prompt, not a general reduction in sensitive-information exposure, and we specify the associated threat model and its out-of-scope leakage channels. Across two benchmarks (a purpose-built credit-rating benchmark designed by the author alongside the Semantic Layer specification, and the externally-authored BIRD financial benchmark) and three LLM tiers spanning a capability range, we find that accuracy is model-dependent rather than uniformly preserved: on the author-designed credit-rating benchmark v2 improves with model capability and, at the strongest tier, is numerically highest among the three conditions (60.7% vs. 58.7% for v1 and 52.0% for Text-to-SQL), although with $$n=60$$ questions this strongest-tier gap does not reach statistical significance, whereas on the externally-authored benchmark v2 trails both baselines at every tier, with the in-benchmark-versus-external gap widening as capability increases. We characterize this pattern as selective encapsulation , presented as an empirical design lesson rather than a predictive theory. Rather than extending Parnas’s (1972) Information Hiding principle, we empirically demonstrate that it is easily violated at LLM interfaces and requires payload-level verification. We also contribute the first domain-specific NL-to-SQL benchmark for credit-rating contexts, together with an execution-based self-check of its gold queries performed by the benchmark designer.

Discover Artificial IntelligenceVol. 6(1)
Openalex Percentile: Top 12%
Privacy-Preserving Technologies in Data
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A semantic layer architecture for removing schema identifiers from external LLM prompts, with model-dependent accuracy trade-offs — Munil Yang · Discover Artificial Intelligence (2026) | TGRS Research Map | TGRS