Two-Sided Pricing and Learning with Buyer Choice: A Generalized SGD Method

Learning to Price in Two-Sided Markets Many two-sided platforms set purchase prices for sellers and selling prices for buyers. In “Two-Sided Pricing and Learning with Buyer Choice: A Generalized SGD Method,” Lin and Huh study how platforms can make these decisions when responses on both sides are uncertain and must be learned from limited data. They first show that, when supply and demand models are known, simple fixed-price policies that match expected supply and demand can be asymptotically optimal. Building on this insight, the authors develop a novel generalization of classical stochastic gradient descent (SGD) for learning a fixed-price policy. Unlike standard SGD, their method does not require the construction of a strongly convex objective function. This flexibility allows it to accommodate demand interactions arising from buyer choice. The resulting learning algorithm is near-optimal when the planning horizon is large.

Authors

Institutions

Publication Details

Journal
Operations Research
Published
2026-09-24
DOI
https://doi.org/10.1287/opre.2025.1786
Primary Topic
Advanced Bandit Algorithms Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Two-Sided Pricing and Learning with Buyer Choice: A Generalized SGD Method

Mei-Chun Lin, Woonghee Tim Huh
Operations Research
Advanced Bandit Algorithms Research
article

Two-Sided Pricing and Learning with Buyer Choice: A Generalized SGD Method

Mei-Chun Lin, Woonghee Tim Huh
article en

Abstract

Learning to Price in Two-Sided Markets Many two-sided platforms set purchase prices for sellers and selling prices for buyers. In “Two-Sided Pricing and Learning with Buyer Choice: A Generalized SGD Method,” Lin and Huh study how platforms can make these decisions when responses on both sides are uncertain and must be learned from limited data. They first show that, when supply and demand models are known, simple fixed-price policies that match expected supply and demand can be asymptotically optimal. Building on this insight, the authors develop a novel generalization of classical stochastic gradient descent (SGD) for learning a fixed-price policy. Unlike standard SGD, their method does not require the construction of a strongly convex objective function. This flexibility allows it to accommodate demand interactions arising from buyer choice. The resulting learning algorithm is near-optimal when the planning horizon is large.

Operations Research
University of British Columbia (CA), Singapore Management University (SG)
Openalex Percentile: Top 7%
Advanced Bandit Algorithms Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.