The Hot Voice Is a Quiet Voice: Temperature as a Hidden Weight in Logit-Level Persona Mixtures, and Recovering Mixture Shares from Text Alone
Systems that blend several personas or "experts" at the token level commonly give each component its own weight and its own sampling temperature. We show that, under log-linear (product-of-experts) pooling, a per-component temperature is not a stylistic setting but a hidden reweighting: the effective share of component i is (w_i/T_i)/S with S = Σ_j w_j/T_j, and the pooled speaker is silently cooled from temperature T to T/S. On a live six-persona council built from a single base model, a component declared at 25% with temperature 0.10 writes with an effective share of 0.767, and the council speaks at an effective temperature of 0.215. We verify the identity numerically on the running model (maximum discrepancy 4.8×10⁻⁶ nats). We then invert the process: given only a generated text and teacher-forced per-persona log-probabilities, we recover the mixture shares by maximum likelihood on the simplex, with standard errors from the projected Fisher information, and we accept a reading only if it passes a calibration exam whose acceptance bound is declared before measurement. The method declares the voices indistinguishable rather than reporting a number when the standard error exceeds a declared ceiling. We report results on a 32B Arabic model and an 8B English-origin model, and we report two failed earlier exams and five measurement bugs found before the final protocol. Bilingual edition: the English paper is followed by the full Arabic edition in the same file. الملخّص: الأنظمة التي تمزج عدّة شخصيّات أو «خبراء» عند كلّ رمز تعطي عادةً كلّ مكوّن وزنه وحرارة أخذ العيّنات الخاصّة به. نبيّن أنّ حرارة المكوّن في التجميع اللوغاريتميّ الخطّيّ (جداء الخبراء) ليست إعداداً أسلوبيّاً بل إعادة وزن خفيّة: الحصّة الفعليّة للمكوّن هي نسبة وزنه إلى حرارته مقسومة على مجموع هذه النسب، والناطق المجمَّع يبرد بصمت. وعلى مجلس حيّ من ستّ شخصيّات مبنيّ على نموذج أساس واحد، مكوّن معلَن بنسبة خمسة وعشرين بالمئة وحرارة عُشر يكتب بحصّة فعليّة تقارب ثلاثة أرباع المزج. ونتحقّق من الهويّة عدديّاً على النموذج العامل. ثمّ نعكس العمليّة: من النصّ المولَّد ولوغاريتمات احتمال كلّ شخصيّة بالتوجيه القسريّ وحدها، نستردّ حصص المزج بالاحتمال الأعظم على المجسَّم، مع أخطاء معياريّة من معلومات فِشر المُسقَطة، ولا نقبل قراءةً إلّا إذا اجتازت امتحان معايرة أُعلن حدّ قبوله قبل القياس. وتعلن الطريقة أنّ الأصوات غير قابلة للتمييز بدل إصدار رقم حين يتجاوز الخطأ المعياريّ سقفاً معلناً. ونعرض النتائج على نموذج عربيّ بحجم اثنين وثلاثين مليار معامل ونموذج إنجليزيّ الأصل بحجم ثمانية مليارات، ونعرض امتحانين سابقين سقطا وخمسة أخطاء قياس اكتُشفت قبل البروتوكول النهائيّ.
Authors
- Firas Assaf
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23009683
- Primary Topic
- Persona Design and Applications
- Type
- preprint