BERT-based medical chatbot with application-level cryptographic protection
Abstract Artificial Intelligence technology has been frequently used in the personalized healthcare services domain in recent years. The quality of medical services has been improved and a new era has started for healthcare services with the use of Artificial Intelligence. Simultaneously, deep learning models, which are qualified as a special use of Artificial Intelligence, are preferred depending on the number of users and the amount of data increasing with each passing day. This study proposes a BERT-based medical chatbot architecture that combines contextual natural language processing with an application-level cryptographic protection layer for client-server communication. The cryptographic mechanism is intended as an additional protection layer around the transmitted chatbot payload and is not presented as a replacement for standard TLS or as a mechanism for protecting plaintext data during server-side inference. The NLP component uses a BERT-based Transformer model to process medical queries from the HealthFC and TREC-Health datasets. The experimental implementation was developed using the Hugging Face Transformers library with TensorFlow and spaCy, and the model was evaluated using accuracy, precision, recall, F1-score, and AUC-ROC. Five-fold stratified cross-validation was also employed to examine the consistency of the classification performance across data partitions. For communication security, RSA is used for the distribution of the AES secret key, AES is used to encrypt the transmitted medical payload, and SHA-256-based HMAC is employed to provide message integrity and authentication. The proposed model achieved 99% accuracy, 99% precision, 97% recall, 98% F1-score, and 97% AUC-ROC. In the five-fold evaluation, the accuracy values were 1.00, 0.99, 0.99, 0.98, and 0.99, respectively. Compared with the evaluated LSTM, SVM, and Bi-LSTM baselines, which achieved accuracies of 0.88, 0.91, and 0.94, respectively, the proposed BERT-based model achieved the highest classification performance. These results indicate strong classification performance on the evaluated benchmark datasets, while the cryptographic layer provides an additional application-level mechanism for protecting medical information during client-server communication. The proposed security mechanism should be interpreted within this scope: it complements, rather than replaces, TLS and does not protect plaintext medical data after decryption on the server during inference. The study establishes a BERT-based and security-oriented architecture for medical chatbot applications; however, clinical safety, external validation, paraphrase-aware leakage analysis, end-to-end latency, and multilingual or voice-based capabilities require further experimental evaluation.
Authors
- Ahmet Dogukan Sarıyalçınkaya (ORCID: https://orcid.org/0000-0002-1388-5114)
- Hakan Can Altunay (ORCID: https://orcid.org/0000-0002-0175-239X)
Institutions
- Ondokuz Mayıs University (TR)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1038/s41598-026-72040-w
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00