One Memory, Many Requests: When and How Much Knowing the Request Is Worth Under a Budget
A memory system that reuses one bounded context across requests must choose its content before it knows which request will come. Under a reading budget, we compare request-blind selection, which chooses one subset for all requests of a history, with request-aware selection, which can choose a subset for the realized request. Both selectors know the workload: the possible requests and their weights, the fragments designated for each, and the fragment costs. With fixed fragments and additive costs, a request is covered when all its designated fragments are selected. The coverage gap between the two optima measures the value of request information. We express this gap through a hypergraph of minimal groups whose designated fragments cannot fit together. We also characterize when the gap is zero. For each optimum, we give the smallest budget that reaches complete coverage. On three workloads (LoCoMo, FairytaleQA, and ContractNLI), peak coverage gaps are 51.27–87.67 percentage points at budgets about 1–17% of their median history lengths. Most FairytaleQA requests need a single fragment. Under the selected optima, these requests contribute 94.30% of FairytaleQA's peak coverage gap. Request-blind selection requires 3.64–9.29 times the smallest uniform per-history budget needed by optimal request-aware selection for complete coverage. For a request-blind extractive policy under the same workload, costs, and budget, the coverage deficit relative to the request-aware optimum separates into the value of request information and a selection shortfall. Better request-blind selection can close the shortfall. Under automatic scoring, coverage-optimal (oracle) request-aware selection gives the LLM readers Qwen and Gemini accuracy advantages of 40.74–55.25 percentage points over request-blind selection at the selected peak-gap budgets on the main evaluation domains. On LoCoMo and FairytaleQA, these advantages remain positive after we correct the automatic scores with a human review of sampled answers. A linear prediction from coverage, calibrated with two control accuracies, tracks the accuracy differences closely on FairytaleQA and less closely on LoCoMo and ContractNLI. This record contains the English preprint and its supplementary material.
Authors
- Hao Fu (ORCID: https://orcid.org/0009-0001-2745-6951)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-01
- DOI
- https://doi.org/10.5281/zenodo.23083934
- Primary Topic
- Personal Information Management and User Behavior
- Type
- preprint