Systematic evaluation of attention based transformer components for parametric modeling of antenna design

This paper systematically evaluates how individual components of attention based Transformer architectures affect the performance of a data driven surrogate model for predicting the reflection coefficient ( \\(S_{11}\\) ) of a parametric microstrip antenna. The antenna under investigation is a flexible microstrip patch antenna fabricated on a polyimide substrate (dielectric constant = 3.5) with a copper radiating patch, having an overall footprint of \\(36\\,\\textrm{mm} \\times 36\\,\\textrm{mm} \\times 0.1\\,\\textrm{mm}\\) and parameterized by four design variables: the \\(L_{P1}\\) , \\(W_{P2}\\) , \\(L_{G3}\\) , and \\(W_{G1}\\) . A dataset of 306 geometry alongside frequency observations is generated from 17 distinct antenna geometries simulated in Ansys HFSS (High Frequency Structure Simulator) across the frequency range from \\(1\\,\\textrm{GHz}\\) to \\(8\\,\\textrm{GHz}\\) with a uniform step of \\(0.4\\,\\textrm{GHz}\\) , yielding 18 frequency samples per geometry. The input to the model is a five token continuous vector formed by the four geometrical parameters and the operating frequency, and the output is the corresponding reflection coefficient \\(S_{11}\\) expressed in decibels. Controlled ablations examine the effects of attention mechanism, normalization scheme, optimizer, and activation function on regression performance. Specifically, standard multi head attention is compared against linear attention. LayerNorm, RMSNorm, and QKNorm are compared as normalization options. And ReLU, GELU, and SwiGLU are compared as activation functions. Training is performed with the Adam and AdamW optimizers under a fixed 70/15/15 geometry level train/validation/test split, implemented using a grouped shuffle split with a fixed random seed to ensure that all frequency samples from any test geometry are excluded from the training and validation sets, thereby preventing information leakage. The surrogate predictions are validated against HFSS simulation results and against measured \\(S_{11}\\) and gain data from a fabricated prototype obtained with a vector network analyzer. This confirms the practical utility of the proposed approach. Hence, the study provides an empirical design guideline for selecting Transformer components in modeling tasks.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-17
DOI
https://doi.org/10.1038/s41598-026-70021-7
Primary Topic
Antenna Design and Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Systematic evaluation of attention based transformer components for parametric modeling of antenna design

Ravi Prakash Dwivedi, Md. Zuber Khan
Scientific Reports
Antenna Design and Analysis
article

Systematic evaluation of attention based transformer components for parametric modeling of antenna design

Ravi Prakash Dwivedi, Md. Zuber Khan
article en

Abstract

This paper systematically evaluates how individual components of attention based Transformer architectures affect the performance of a data driven surrogate model for predicting the reflection coefficient ( \(S_{11}\) ) of a parametric microstrip antenna. The antenna under investigation is a flexible microstrip patch antenna fabricated on a polyimide substrate (dielectric constant = 3.5) with a copper radiating patch, having an overall footprint of \(36\,\textrm{mm} \times 36\,\textrm{mm} \times 0.1\,\textrm{mm}\) and parameterized by four design variables: the \(L_{P1}\) , \(W_{P2}\) , \(L_{G3}\) , and \(W_{G1}\) . A dataset of 306 geometry alongside frequency observations is generated from 17 distinct antenna geometries simulated in Ansys HFSS (High Frequency Structure Simulator) across the frequency range from \(1\,\textrm{GHz}\) to \(8\,\textrm{GHz}\) with a uniform step of \(0.4\,\textrm{GHz}\) , yielding 18 frequency samples per geometry. The input to the model is a five token continuous vector formed by the four geometrical parameters and the operating frequency, and the output is the corresponding reflection coefficient \(S_{11}\) expressed in decibels. Controlled ablations examine the effects of attention mechanism, normalization scheme, optimizer, and activation function on regression performance. Specifically, standard multi head attention is compared against linear attention. LayerNorm, RMSNorm, and QKNorm are compared as normalization options. And ReLU, GELU, and SwiGLU are compared as activation functions. Training is performed with the Adam and AdamW optimizers under a fixed 70/15/15 geometry level train/validation/test split, implemented using a grouped shuffle split with a fixed random seed to ensure that all frequency samples from any test geometry are excluded from the training and validation sets, thereby preventing information leakage. The surrogate predictions are validated against HFSS simulation results and against measured \(S_{11}\) and gain data from a fabricated prototype obtained with a vector network analyzer. This confirms the practical utility of the proposed approach. Hence, the study provides an empirical design guideline for selecting Transformer components in modeling tasks.

Scientific Reports
Vellore Institute of Technology University (IN)
Openalex Percentile: Top 8%
Antenna Design and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.