| 研究生: |
林長毅 Lin, Chang-Yi |
|---|---|
| 論文名稱: |
干擾下的自適應 A/B 測試:應用於叫車平台的多臂拉霸機架構 Adaptive A/B Testing under Interference: A Multi-Armed Bandit Framework for Ride-Hailing Platforms |
| 指導教授: |
莊雅棠
Chuang, Ya-Tang |
| 學位類別: |
碩士 Master |
| 系所名稱: |
管理學院 - 工業與資訊管理學系 Department of Industrial and Information Management |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 72 |
| 中文關鍵詞: | A/B測試 、多臂拉霸機 、空間干擾 、叫車平台 |
| 外文關鍵詞: | A/B testing, multi-armed bandits, ride-hailing platforms, spatial interference |
| 相關次數: | 點閱:3 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
在雙邊平台上進行線上 A/B 測試具有挑戰,因為實驗政策可能透過共享供給、地理鄰近性與平台配對機制,在使用者之間產生外溢效果,在叫車服務中,採用固定流量分配的標準實驗通常依賴獨立性與平穩性假設,上述外溢效果可能使這兩項假設不再成立,針對存在干擾的叫車平台定價實驗,本研究提出一套自適應實驗架構,首先在空間網格層級分配定價策略,以降低個體隨機分派所造成的污染,接著透過乾淨曝光條件與 HT-IX 估計器,將乘客回饋轉換為經干擾調整後的手臂層級回饋,為因應非平穩的市場環境,本架構透過 Hedge 混合機制結合 ThompsonSampling 與 EXP3,使平台能在穩定市場中的學習效率與波動環境下的穩健性之間取得平衡。以曼哈頓道路網路與歷史需求資料進行的數值實驗顯示,相較於單獨使用 Thompson Sampling 或 EXP3,此整合方法能提升穩定性、減少過早收斂,並達成較低的累積後悔值。
Online A/B testing on two-sided platforms is challenging because experimental interventions may create spillovers across users through shared supply, geographic proximity, and platform matching. In ride-hailing services, these effects can violate the independence and stationarity assumptions underlying standard experiments with fixed traffic allocation. This thesis develops an adaptive experimentation framework for pricing experiments in ride-hailing platforms under interference. The framework first assigns pricing policies at the spatial grid level to reduce contamination from individual randomization. It then uses a clean exposure condition and an HT-IX estimator to transform rider feedback into arm-level feedback that is adjusted for interference. To address nonstationary market conditions, the framework combines Thompson Sampling and EXP3 through a Hedge mixing mechanism, allowing the platform to balance efficient learning in stable markets with robustness in volatile environments. Numerical experiments based on the Manhattan road network and historical demand data show that the integrated approach improves stability, reduces premature convergence, and achieves lower cumulative regret than using Thompson Sampling or EXP3 alone.
Agarwal, A., Agarwal, A., Masoero, L., and Whitehouse, J. (2024). Mutli-armed bandits with network interference. Advances in Neural Information Processing Systems, 37:36414-36437.
Agrawal, S. and Goyal, N. (2012). Analysis of thompson sampling for the multi-armed bandit problem. In Conference on Learning Theory, pages 39-1. JMLR Workshop and Conference Proceedings.
Aramayo, N., Schiappacasse, M., and Goic, M. (2023). A multiarmed bandit approach for house ads recommendations. Marketing Science, 42(2):271-292.
Aronow, P. M. and Samii, C. (2017). Estimating average causal effects under general interference, with application to a social network experiment. The Annals of Applied Statistics, 11(4):1912-1947.
Athey, S., Eckles, D., and Imbens, G. W. (2018). Exact p-values for network interference. Journal of the American Statistical Association, 113(521):230-240.
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E. (2002). The nonstochastic multiarmed bandit problem. SIAM Journal on Computing, 32(1):48-77.
Auer, P. and Chiang, C.-K. (2016). An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits. In Conference on Learning Theory, pages 116-120. PMLR.
Backstrom, L. and Kleinberg, J. (2011). Network bucket testing. In Proceedings of the 20th International Conference on World Wide Web, pages 615-624.
Banerjee, S., Riquelme, C., and Johari, R. (2015). Pricing in ride-share platforms: A queueing-theoretic approach. Available at SSRN 2568258.
Bubeck, S. and Slivkins, A. (2012). The best of both worlds: Stochastic and adversarial bandits. In Conference on Learning Theory, pages 42-1. JMLR Workshop and Conference Proceedings.
Cachon, G. P., Daniels, K. M., and Lobel, R. (2017). The role of surge pricing on a service platform with self-scheduling capacity. Manufacturing & Service Operations Management, 19(3):368-384.
Cesa-Bianchi, N. and Lugosi, G. (2006). Prediction, learning, and games. Cambridge University Press.
Chamandy, N. (2016). Experimentation in a ridesharing marketplace. https://eng.lyft.com/experimentation-in-a-ridesharing-marketplace-b39db027a66e.
Eckles, D., Karrer, B., and Ugander, J. (2017). Design and analysis of experiments in networks: Reducing bias from interference. Journal of Causal Inference, 5(1):20150021.
Engelhardt, R., Dandl, F., Syed, A.-A., Zhang, Y., Fehn, F., Wolf, F., and Bogenberger, K. (2022). Fleetpy: A modular open-source simulation tool for mobility on-demand services. arXiv Preprint arXiv:2207.14246.
Evans, D. and Schmalensee, R. (2016). The new economics of multi-sided platforms: A guide to the vocabulary. SSRN Electronic Journal.
Feit, E. M. and Berman, R. (2019). Test & roll: Profit-maximizing a/b tests. Marketing Science, 38(6):1038-1058.
Filippas, A., Jagabathula, S., and Sundararajan, A. (2023). The limits of centralized pricing in online marketplaces and the value of user control. Management Science, 69(12):7202-7216.
Freund, Y. and Schapire, R. E. (1997). A decision-theoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences, 55(1):119-139.
Geyik, S. C., Ambler, S., and Kenthapadi, K. (2019). Fairness-aware ranking in search & recommendation systems with application to linkedin talent search. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 2221-2231.
Gomez-Uribe, C. A. and Hunt, N. (2016). The netflix recommender system: Algorithms, business value, and innovation. ACM Trans. Manage. Inf. Syst., 6(4).
Gui, H., Xu, Y., Bhasin, A., and Han, J. (2015). Network a/b testing: From sampling to estimation. In Proceedings of the 24th International Conference on World Wide Web, pages 399-409.
Ha-Thuc, V., Dutta, A., Mao, R., Wood, M., and Liu, Y. (2020). A counterfactual framework for seller-side a/b testing on marketplaces. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2288-2296.
Hagiu, A. and Wright, J. (2015). Multi-sided platforms. International Journal of Industrial Organization, 43:162-174.
Holtz, D., Lobel, F., Lobel, R., Liskovich, I., and Aral, S. (2025). Reducing interference bias in online marketplace experiments using cluster randomization: Evidence from a pricing meta-experiment on Airbnb. Management Science, 71(1):390-406.
Hudgens, M. G. and Halloran, M. E. (2008). Toward causal inference with interference. Journal of the American Statistical Association, 103(482):832-842.
Jia, S., Frazier, P. I., and Kallus, N. (2025). Multi-armed bandits with interference: Bridging causal inference and adversarial bandits. In Forty-second International Conference on Machine Learning.
Jin, S. T., Kong, H., Wu, R., and Sui, D. Z. (2018). Ridesourcing, the sharing economy, and the future of cities. Cities, 76:96-104.
Johari, R., Li, H., Liskovich, I., and Weintraub, G. Y. (2022). Experimental design in two-sided platforms: An analysis of bias. Management Science, 68(10):7069-7089.
Kocák, T., Neu, G., Valko, M., and Munos, R. (2014). Efficient learning by implicit exploration in bandit problems with side observations. In Advances in Neural Information Processing Systems, volume 27, pages 613-621.
Kohavi, R., Longbotham, R., Sommerfield, D., and Henne, R. M. (2009). Controlled experiments on the web: Survey and practical guide. Data Mining and Knowledge Discovery, 18:140-181.
Kohavi, R., Tang, D., and Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to a/b testing. Cambridge University Press.
Leung, M. P. (2022). Rate-optimal cluster-randomized designs for spatial interference. The Annals of Statistics, 50(5):3064-3087.
Lewis, R. A. and Rao, J. M. (2015). The unfavorable economics of measuring the returns to advertising. The Quarterly Journal of Economics, 130(4):1941-1973.
Lian, Z. and Van Ryzin, G. (2021). Optimal growth in two-sided markets. Management Science, 67(11):6862-6879.
Misra, K., Schwartz, E. M., and Abernethy, J. (2019). Dynamic online pricing with incomplete information using multiarmed bandit experiments. Marketing Science, 38(2):226-252.
Neu, G. (2015). Explore no more: Improved high-probability regret bounds for non-stochastic bandits. Advances in Neural Information Processing Systems, 28.
Quin, F., Weyns, D., Galster, M., and Silva, C. C. (2024). A/b testing: A systematic literature review. Journal of Systems and Software, 211:112011.
Rochet, J.-C. and Tirole, J. (2003). Platform competition in two-sided markets. Journal of the European Economic Association, 1(4):990-1029.
Rochet, J.-C. and Tirole, J. (2006). Two-sided markets: A progress report. The RAND Journal of Economics, 37(3):645-667.
Rubin, D. B. (1978). Bayesian inference for causal effects: The role of randomization. The Annals of Statistics, pages 34-58.
Schwartz, E. M., Bradlow, E. T., and Fader, P. S. (2017). Customer acquisition via display advertising using multi-armed bandit experiments. Marketing Science, 36(4):500-522.
Scott, S. L. (2015). Multi-armed bandit experiments in the online service economy. Applied Stochastic Models in Business and Industry, 31(1):37-45.
Seldin, Y. and Slivkins, A. (2014). One practical algorithm for both stochastic and adversarial bandits. In International Conference on Machine Learning, pages 1287-1295. PMLR.
Simchi-Levi, D. and Wang, C. (2025). Multi-armed bandit experimental design: Online decision-making and adaptive inference. Management Science, 71(6):4828-4846.
Tang, D., Agarwal, A., O'Brien, D., and Meyer, M. (2010). Overlapping experiment infrastructure: More, better, faster experimentation. In Proceedings of the 16th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 17-26.
Trovo, F., Paladino, S., Restelli, M., and Gatti, N. (2020). Sliding-window thompson sampling for non-stationary settings. Journal of Artificial Intelligence Research, 68:311-364.
Ugander, J., Karrer, B., Backstrom, L., and Kleinberg, J. (2013). Graph cluster randomization: Network exposure to multiple universes. In Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 329-337.
Ugander, J. and Yin, H. (2023). Randomized graph cluster randomization. Journal of Causal Inference, 11(1):20220014.
Xu, Y., Chen, N., Fernandez, A., Sinno, O., and Bhasin, A. (2015). From infrastructure to culture: A/b testing challenges in large scale social networks. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 2227-2236.
Xu, Y., Lu, W., and Song, R. (2024). Linear contextual bandits with interference. arXiv Preprint arXiv:2409.15682.
Yoon, S. (2018). Designing a/b tests in a collaboration network. The Unofficial Google Data Science Blog.
Yuan, Y., Altenburger, K., and Kooti, F. (2021). Causal network motifs: Identifying heterogeneous spillover effects in a/b tests. In Proceedings of the Web Conference 2021, pages 3359-3370.
Yuan, Y. and Altenburger, K. M. (2025). A two-part machine learning approach to characterizing network interference in a/b testing. Manufacturing & Service Operations Management.
Zhang, Z. and Wang, Z. (2024). Online experimental design with estimation-regret trade-off under network interference. arXiv Preprint arXiv:2412.03727.
Zimmert, J. and Seldin, Y. (2019). An optimal algorithm for stochastic and adversarial bandits. In The 22nd International Conference on Artificial Intelligence and Statistics, pages 467-475. PMLR.
Zimmert, J. and Seldin, Y. (2021). Tsallis-inf: An optimal algorithm for stochastic and adversarial bandits. Journal of Machine Learning Research, 22(28):1-49.