Metadata Leakage in Encrypted Voice
Encrypted real-time voice can reveal information through observable metadata even when the payload cipher remains secure. This paper reformulates the “prosodic shadows” hypothesis as a side-channel problem rather than an encryption null-space theorem. Speech U is encoded and packetized into an encrypted payload C and observable metadata M; payload confidentiality can coexist with I(T;M)>0 for a target T such as language, speech activity, phrase class, or other declared attribute because codec and transport behavior depend on the source signal. The framework defines target-relative metadata leakage, complementary metadata observations, and defense transformations. By the data-processing inequality, any defense that releases a randomized transformation M_prime of M cannot increase information about the target, but the amount removed is empirical and must be reported together with bandwidth, latency, and quality costs. Established encrypted-VoIP studies demonstrate substantial language and phrase leakage for particular VBR codecs, while the magnitude of affective or open-ended lexical leakage remains unestablished. We therefore specify a consent-based, self-generated experimental protocol that separates speaker, utterance, codec, and network confounds and compares metadata-conditioned inference against prior-only baselines. The resulting theory supports rigorous measurement of prosodic side channels without treating encryption as a rank-deficient linear map or presenting unvalidated reconstruction estimates as findings.
Authors
- Cecil Jentges (ORCID: https://orcid.org/0009-0004-9986-8551)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-17
- DOI
- https://doi.org/10.5281/zenodo.22821555
- Primary Topic
- Speech Recognition and Synthesis
- Type
- preprint