Beyond the Prompt: Local Optimization and Provider Caching for Cost-Efficient LLM Workloads

{"tonst":[0],"is":[1,108,112],"an":[2,9,98],"MIT-licensed,":[3],"local-first":[4],"wrapper":[5],"that":[6],"sits":[7],"between":[8],"application":[10],"and":[11,31,66,74,82,97,125,154],"any":[12],"paid":[13],"LLM":[14],"API.":[15],"Before":[16],"a":[17,57,69,87,93,118,132],"request":[18,34],"leaves":[19],"the":[20,33,36,79,106,126,140,148],"calling":[21],"machine,":[22],"it":[23],"redacts":[24],"personal":[25],"data":[26],"(PII),":[27],"trims":[28],"redundant":[29],"tokens,":[30],"structures":[32],"so":[35],"provider's":[37],"own":[38,128],"prompt-caching":[39],"mechanism":[40],"can":[41],"discount":[42],"repeat":[43],"calls.":[44],"This":[45],"whitepaper":[46],"reports":[47],"real,":[48,133],"measured":[49],"results":[50],"rather":[51],"than":[52],"simulated":[53],"ones":[54],"wherever":[55],"possible:":[56],"360-iteration":[58],"benchmark":[59,152],"across":[60],"six":[61],"industry":[62],"domains":[63],"(token":[64],"reduction":[65,91],"PII-redaction":[67],"recall),":[68],"200-call":[70],"concurrency":[71],"stress":[72],"test,":[73],"live-billed":[75],"caching":[76],"tests":[77],"against":[78],"real":[80,119],"Anthropic":[81],"Gemini":[83],"APIs":[84],"—":[85,130,144],"including":[86,131],"53%":[88],"net":[89],"cost":[90,100],"on":[92,102],"realistic":[94],"mixed":[95],"workload":[96],"84.9%":[99],"saving":[101],"cache-read":[103],"calls":[104],"once":[105],"cache":[107],"warm.":[109],"Every":[110],"figure":[111],"labeled":[113],"as":[114],"either":[115],"\\"measured\\"":[116],"(from":[117],"run)":[120],"or":[121],"\\"estimate\\"":[122],"(mocked":[123],"pricing),":[124],"paper's":[127],"limitations":[129],"non-zero":[134],"free-text":[135],"PII":[136],"leak":[137],"rate":[138],"under":[139],"current":[141],"redaction":[142],"configuration":[143],"are":[145],"reported":[146],"alongside":[147],"results.":[149],"Source":[150],"code,":[151],"scripts,":[153],"full":[155],"methodology":[156],"notes:":[157],"https://github.com/v-nightwolf/tonst":[158]}

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-14
DOI
https://doi.org/10.5281/zenodo.22745267
Primary Topic
Personal Information Management and User Behavior
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Beyond the Prompt: Local Optimization and Provider Caching for Cost-Efficient LLM Workloads

Gahlout Vinay
Zenodo (CERN European Organization for Nuclear Research)
Personal Information Management and User Behavior
preprint

Beyond the Prompt: Local Optimization and Provider Caching for Cost-Efficient LLM Workloads

Gahlout Vinay
preprint en

Abstract

tonst is an MIT-licensed, local-first wrapper that sits between an application and any paid LLM API. Before a request leaves the calling machine, it redacts personal data (PII), trims redundant tokens, and structures the request so the provider's own prompt-caching mechanism can discount repeat calls. This whitepaper reports real, measured results rather than simulated ones wherever possible: a 360-iteration benchmark across six industry domains (token reduction and PII-redaction recall), a 200-call concurrency stress test, and live-billed caching tests against the real Anthropic and Gemini APIs — including a 53% net cost reduction on a realistic mixed workload and an 84.9% cost saving on cache-read calls once the cache is warm. Every figure is labeled as either "measured" (from a real run) or "estimate" (mocked pricing), and the paper's own limitations — including a real, non-zero free-text PII leak rate under the current redaction configuration — are reported alongside the results. Source code, benchmark scripts, and full methodology notes: https://github.com/v-nightwolf/tonst

Zenodo (CERN European Organization for Nuclear Research)
Industry, innovation and infrastructure
Personal Information Management and User Behavior
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Beyond the Prompt: Local Optimization and Provider Caching for Cost-Efficient LLM Workloads — Gahlout Vinay · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS