簡易檢索 / 詳目顯示

研究生: 林長毅
Lin, Chang-Yi
論文名稱: 干擾下的自適應 A/B 測試:應用於叫車平台的多臂拉霸機架構
Adaptive A/B Testing under Interference: A Multi-Armed Bandit Framework for Ride-Hailing Platforms
指導教授: 莊雅棠
Chuang, Ya-Tang
學位類別: 碩士
Master
系所名稱: 管理學院 - 工業與資訊管理學系
Department of Industrial and Information Management
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 72
中文關鍵詞: A/B測試多臂拉霸機空間干擾叫車平台
外文關鍵詞: A/B testing, multi-armed bandits, ride-hailing platforms, spatial interference
相關次數: 點閱:3下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 在雙邊平台上進行線上 A/B 測試具有挑戰,因為實驗政策可能透過共享供給、地理鄰近性與平台配對機制,在使用者之間產生外溢效果,在叫車服務中,採用固定流量分配的標準實驗通常依賴獨立性與平穩性假設,上述外溢效果可能使這兩項假設不再成立,針對存在干擾的叫車平台定價實驗,本研究提出一套自適應實驗架構,首先在空間網格層級分配定價策略,以降低個體隨機分派所造成的污染,接著透過乾淨曝光條件與 HT-IX 估計器,將乘客回饋轉換為經干擾調整後的手臂層級回饋,為因應非平穩的市場環境,本架構透過 Hedge 混合機制結合 ThompsonSampling 與 EXP3,使平台能在穩定市場中的學習效率與波動環境下的穩健性之間取得平衡。以曼哈頓道路網路與歷史需求資料進行的數值實驗顯示,相較於單獨使用 Thompson Sampling 或 EXP3,此整合方法能提升穩定性、減少過早收斂,並達成較低的累積後悔值。

    Online A/B testing on two-sided platforms is challenging because experimental interventions may create spillovers across users through shared supply, geographic proximity, and platform matching. In ride-hailing services, these effects can violate the independence and stationarity assumptions underlying standard experiments with fixed traffic allocation. This thesis develops an adaptive experimentation framework for pricing experiments in ride-hailing platforms under interference. The framework first assigns pricing policies at the spatial grid level to reduce contamination from individual randomization. It then uses a clean exposure condition and an HT-IX estimator to transform rider feedback into arm-level feedback that is adjusted for interference. To address nonstationary market conditions, the framework combines Thompson Sampling and EXP3 through a Hedge mixing mechanism, allowing the platform to balance efficient learning in stable markets with robustness in volatile environments. Numerical experiments based on the Manhattan road network and historical demand data show that the integrated approach improves stability, reduces premature convergence, and achieves lower cumulative regret than using Thompson Sampling or EXP3 alone.

    中文摘要 I 英文延伸摘要 II 誌謝 V 目錄 VI 表目錄 VIII 圖目錄 IX 第一章 研究動機與目的 1   1.1 叫車服務平台與線上實驗 1   1.2 研究目的與貢獻 5 第二章 文獻回顧 8   2.1 平台實驗中的干擾與估計偏誤 8   2.2 干擾校正與叢集化實驗設計 9   2.3 自適應與考量干擾效應的多臂拉霸機 10 第三章 叫車服務實驗中的干擾與非平穩性 14 第四章 考量干擾效應的自適應實驗架構 19   4.1 實驗單位、手臂與回饋 20   4.2 全域同手臂反事實與後悔值基準 22   4.3 空間指派與干擾控制 23   4.4 乾淨曝光與 HT-IX 回饋估計 26   4.5 自適應策略分布 28   4.6 湯普森抽樣子策略 31   4.7 EXP3 子策略 33   4.8 小結與整體演算法流程 34 第五章 模擬實驗 36   5.1 實驗設定 37   5.2 數值分析 42   5.2.1 最佳手臂 42   5.2.2 網格大小初始值設定 44   5.2.3 不同演算法之表現比較 44 第六章 結論與未來展望 50   6.1 結論 50   6.2 未來展望 51 參考文獻 53

    Agarwal, A., Agarwal, A., Masoero, L., and Whitehouse, J. (2024). Mutli-armed bandits with network interference. Advances in Neural Information Processing Systems, 37:36414-36437.
    Agrawal, S. and Goyal, N. (2012). Analysis of thompson sampling for the multi-armed bandit problem. In Conference on Learning Theory, pages 39-1. JMLR Workshop and Conference Proceedings.
    Aramayo, N., Schiappacasse, M., and Goic, M. (2023). A multiarmed bandit approach for house ads recommendations. Marketing Science, 42(2):271-292.
    Aronow, P. M. and Samii, C. (2017). Estimating average causal effects under general interference, with application to a social network experiment. The Annals of Applied Statistics, 11(4):1912-1947.
    Athey, S., Eckles, D., and Imbens, G. W. (2018). Exact p-values for network interference. Journal of the American Statistical Association, 113(521):230-240.
    Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E. (2002). The nonstochastic multiarmed bandit problem. SIAM Journal on Computing, 32(1):48-77.
    Auer, P. and Chiang, C.-K. (2016). An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits. In Conference on Learning Theory, pages 116-120. PMLR.
    Backstrom, L. and Kleinberg, J. (2011). Network bucket testing. In Proceedings of the 20th International Conference on World Wide Web, pages 615-624.
    Banerjee, S., Riquelme, C., and Johari, R. (2015). Pricing in ride-share platforms: A queueing-theoretic approach. Available at SSRN 2568258.
    Bubeck, S. and Slivkins, A. (2012). The best of both worlds: Stochastic and adversarial bandits. In Conference on Learning Theory, pages 42-1. JMLR Workshop and Conference Proceedings.
    Cachon, G. P., Daniels, K. M., and Lobel, R. (2017). The role of surge pricing on a service platform with self-scheduling capacity. Manufacturing & Service Operations Management, 19(3):368-384.
    Cesa-Bianchi, N. and Lugosi, G. (2006). Prediction, learning, and games. Cambridge University Press.
    Chamandy, N. (2016). Experimentation in a ridesharing marketplace. https://eng.lyft.com/experimentation-in-a-ridesharing-marketplace-b39db027a66e.
    Eckles, D., Karrer, B., and Ugander, J. (2017). Design and analysis of experiments in networks: Reducing bias from interference. Journal of Causal Inference, 5(1):20150021.
    Engelhardt, R., Dandl, F., Syed, A.-A., Zhang, Y., Fehn, F., Wolf, F., and Bogenberger, K. (2022). Fleetpy: A modular open-source simulation tool for mobility on-demand services. arXiv Preprint arXiv:2207.14246.
    Evans, D. and Schmalensee, R. (2016). The new economics of multi-sided platforms: A guide to the vocabulary. SSRN Electronic Journal.
    Feit, E. M. and Berman, R. (2019). Test & roll: Profit-maximizing a/b tests. Marketing Science, 38(6):1038-1058.
    Filippas, A., Jagabathula, S., and Sundararajan, A. (2023). The limits of centralized pricing in online marketplaces and the value of user control. Management Science, 69(12):7202-7216.
    Freund, Y. and Schapire, R. E. (1997). A decision-theoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences, 55(1):119-139.
    Geyik, S. C., Ambler, S., and Kenthapadi, K. (2019). Fairness-aware ranking in search & recommendation systems with application to linkedin talent search. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 2221-2231.
    Gomez-Uribe, C. A. and Hunt, N. (2016). The netflix recommender system: Algorithms, business value, and innovation. ACM Trans. Manage. Inf. Syst., 6(4).
    Gui, H., Xu, Y., Bhasin, A., and Han, J. (2015). Network a/b testing: From sampling to estimation. In Proceedings of the 24th International Conference on World Wide Web, pages 399-409.
    Ha-Thuc, V., Dutta, A., Mao, R., Wood, M., and Liu, Y. (2020). A counterfactual framework for seller-side a/b testing on marketplaces. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2288-2296.
    Hagiu, A. and Wright, J. (2015). Multi-sided platforms. International Journal of Industrial Organization, 43:162-174.
    Holtz, D., Lobel, F., Lobel, R., Liskovich, I., and Aral, S. (2025). Reducing interference bias in online marketplace experiments using cluster randomization: Evidence from a pricing meta-experiment on Airbnb. Management Science, 71(1):390-406.
    Hudgens, M. G. and Halloran, M. E. (2008). Toward causal inference with interference. Journal of the American Statistical Association, 103(482):832-842.
    Jia, S., Frazier, P. I., and Kallus, N. (2025). Multi-armed bandits with interference: Bridging causal inference and adversarial bandits. In Forty-second International Conference on Machine Learning.
    Jin, S. T., Kong, H., Wu, R., and Sui, D. Z. (2018). Ridesourcing, the sharing economy, and the future of cities. Cities, 76:96-104.
    Johari, R., Li, H., Liskovich, I., and Weintraub, G. Y. (2022). Experimental design in two-sided platforms: An analysis of bias. Management Science, 68(10):7069-7089.
    Kocák, T., Neu, G., Valko, M., and Munos, R. (2014). Efficient learning by implicit exploration in bandit problems with side observations. In Advances in Neural Information Processing Systems, volume 27, pages 613-621.
    Kohavi, R., Longbotham, R., Sommerfield, D., and Henne, R. M. (2009). Controlled experiments on the web: Survey and practical guide. Data Mining and Knowledge Discovery, 18:140-181.
    Kohavi, R., Tang, D., and Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to a/b testing. Cambridge University Press.
    Leung, M. P. (2022). Rate-optimal cluster-randomized designs for spatial interference. The Annals of Statistics, 50(5):3064-3087.
    Lewis, R. A. and Rao, J. M. (2015). The unfavorable economics of measuring the returns to advertising. The Quarterly Journal of Economics, 130(4):1941-1973.
    Lian, Z. and Van Ryzin, G. (2021). Optimal growth in two-sided markets. Management Science, 67(11):6862-6879.
    Misra, K., Schwartz, E. M., and Abernethy, J. (2019). Dynamic online pricing with incomplete information using multiarmed bandit experiments. Marketing Science, 38(2):226-252.
    Neu, G. (2015). Explore no more: Improved high-probability regret bounds for non-stochastic bandits. Advances in Neural Information Processing Systems, 28.
    Quin, F., Weyns, D., Galster, M., and Silva, C. C. (2024). A/b testing: A systematic literature review. Journal of Systems and Software, 211:112011.
    Rochet, J.-C. and Tirole, J. (2003). Platform competition in two-sided markets. Journal of the European Economic Association, 1(4):990-1029.
    Rochet, J.-C. and Tirole, J. (2006). Two-sided markets: A progress report. The RAND Journal of Economics, 37(3):645-667.
    Rubin, D. B. (1978). Bayesian inference for causal effects: The role of randomization. The Annals of Statistics, pages 34-58.
    Schwartz, E. M., Bradlow, E. T., and Fader, P. S. (2017). Customer acquisition via display advertising using multi-armed bandit experiments. Marketing Science, 36(4):500-522.
    Scott, S. L. (2015). Multi-armed bandit experiments in the online service economy. Applied Stochastic Models in Business and Industry, 31(1):37-45.
    Seldin, Y. and Slivkins, A. (2014). One practical algorithm for both stochastic and adversarial bandits. In International Conference on Machine Learning, pages 1287-1295. PMLR.
    Simchi-Levi, D. and Wang, C. (2025). Multi-armed bandit experimental design: Online decision-making and adaptive inference. Management Science, 71(6):4828-4846.
    Tang, D., Agarwal, A., O'Brien, D., and Meyer, M. (2010). Overlapping experiment infrastructure: More, better, faster experimentation. In Proceedings of the 16th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 17-26.
    Trovo, F., Paladino, S., Restelli, M., and Gatti, N. (2020). Sliding-window thompson sampling for non-stationary settings. Journal of Artificial Intelligence Research, 68:311-364.
    Ugander, J., Karrer, B., Backstrom, L., and Kleinberg, J. (2013). Graph cluster randomization: Network exposure to multiple universes. In Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 329-337.
    Ugander, J. and Yin, H. (2023). Randomized graph cluster randomization. Journal of Causal Inference, 11(1):20220014.
    Xu, Y., Chen, N., Fernandez, A., Sinno, O., and Bhasin, A. (2015). From infrastructure to culture: A/b testing challenges in large scale social networks. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 2227-2236.
    Xu, Y., Lu, W., and Song, R. (2024). Linear contextual bandits with interference. arXiv Preprint arXiv:2409.15682.
    Yoon, S. (2018). Designing a/b tests in a collaboration network. The Unofficial Google Data Science Blog.
    Yuan, Y., Altenburger, K., and Kooti, F. (2021). Causal network motifs: Identifying heterogeneous spillover effects in a/b tests. In Proceedings of the Web Conference 2021, pages 3359-3370.
    Yuan, Y. and Altenburger, K. M. (2025). A two-part machine learning approach to characterizing network interference in a/b testing. Manufacturing & Service Operations Management.
    Zhang, Z. and Wang, Z. (2024). Online experimental design with estimation-regret trade-off under network interference. arXiv Preprint arXiv:2412.03727.
    Zimmert, J. and Seldin, Y. (2019). An optimal algorithm for stochastic and adversarial bandits. In The 22nd International Conference on Artificial Intelligence and Statistics, pages 467-475. PMLR.
    Zimmert, J. and Seldin, Y. (2021). Tsallis-inf: An optimal algorithm for stochastic and adversarial bandits. Journal of Machine Learning Research, 22(28):1-49.

    QR CODE