簡易檢索 / 詳目顯示

研究生: 石子弘
Shih, Zih-Hong
論文名稱: 特徵化多臂拉霸機庫存決策模型
Feature-Based Multi-Armed Bandit Inventory Decision Model
指導教授: 莊雅棠
Chuang, Ya-Tang
學位類別: 碩士
Master
系所名稱: 管理學院 - 工業與資訊管理學系
Department of Industrial and Information Management
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 55
中文關鍵詞: 庫存管理多臂拉霸機線上優化特徵學習截斷需求
外文關鍵詞: inventory management, multi-armed bandit, online optimization, feature learning, censored demand
相關次數: 點閱:3下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 庫存決策在需求不確定的環境下,一直是營運管理中的重要問題。傳統報童模型多假設需求分配已知,然而在實務上,決策者無法完整掌握需求資訊,甚至僅能觀測到受庫存限制而形成的實際銷售量,即截斷需求。隨著企業可取得的需求相關特徵日益豐富,如何在需求未知且資訊受限的情況下,結合特徵資訊進行庫存決策,成為值得研究的議題。本研究聚焦於單一產品庫存管理問題,提出一個特徵化多臂拉霸機庫存決策模型。本研究將各個庫存水準視為手臂,假設需求與特徵之間具有線性關係,並以報童成本函數作為決策基礎,將獎勵定義為成本之負值。在截斷需求設定下,決策者於每一期僅能觀測實際銷售量,而無法得知完整需求,因此必須在有限資訊下同時進行需求學習與庫存決策。本研究以平均遺憾值作為績效衡量指標,比較所提方法與全知最佳策略之差異,透過建構一套特徵化的動態決策流程,使演算法能於每一期先觀測特徵資訊,再依據目前參數估計形成需求預測,從有限個庫存水準中選擇決策,並利用可觀測之銷售資訊逐步更新參數。另外,本研究亦討論非線性需求情境,作為模型後續延伸之方向。數值分析顯示,當演算法納入特徵資訊後,其遺憾值會隨決策期數增加而明顯下降,且整體表現優於不含特徵之基準方法。此結果說明,特徵資訊有助於演算法逐步掌握需求與情境間的關係,進而改善庫存決策品質。因此,本研究提出之特徵化多臂拉霸機庫存決策模型,能在截斷需求之環境下,結合特徵資訊進行動態學習與決策,並透過數值分析驗證其可行性。

    Inventory decision-making is an important issue in operations management, particularly under demand uncertainty. Traditional newsvendor models generally assume that the demand distribution is known; however, this assumption is often unrealistic in practice. Firms may only observe realized sales constrained by available inventory, resulting in censored demand. Meanwhile, firms can increasingly access various demand-related features, such as customer characteristics and market conditions. This study proposes a feature-based multi-armed bandit inventory decision model for a single-product problem. Each feasible inventory level is treated as an arm, and demand is assumed to be associated with observable features. The model adopts the newsvendor cost function as the basis for decision-making, with the reward defined as the negative of the corresponding cost. In each period, the decision maker observes the current features, predicts demand based on the estimated parameters, selects an inventory level, and updates the model using the observed sales data. Model performance is evaluated using average regret relative to an omniscient optimal policy with full knowledge of the underlying demand process. In addition to the linear demand setting, this study extends the proposed framework to nonlinear demand relationships by incorporating squared and interaction terms into the feature representation. The numerical results show that, after feature information is incorporated into the algorithm, average regret decreases significantly as the number of decision periods increases, and the proposed method outperforms the benchmark model without features. These findings indicate that feature information enables the algorithm to gradually learn the relationship between demand and the observed context, thereby improving the quality of inventory decisions. Overall, the proposed feature-based multi-armed bandit inventory decision model effectively integrates feature information into dynamic learning and decision-making under censored demand, and its feasibility is demonstrated through numerical analysis.

    摘要 i 英文延伸摘要 ii 致謝 vii 目錄 ix 表目錄 xi 圖目錄 xii 第一章 緒論 1 第一節 研究背景與動機 1 第二節 研究目標 3 第二章 文獻回顧 5 第一節 報童模型 5 第二節 基於特徵的報童模型 6 第三節 多臂拉霸機 8 第四節 情境式多臂拉霸機 10 第五節 研究定位 11 第三章 模型建構 14 第一節 報童問題 14 第一小節 資料驅動的報童問題 15 第二小節 基於特徵的報童問題 16 第二節 多臂拉霸機庫存問題 17 第一小節 多臂拉霸機模型 17 第二小節 建構多臂拉霸機的傳統庫存模型 18 第三節 特徵化多臂拉霸機庫存問題 20 第一小節 建構特徵化多臂拉霸機庫存模型 21 第四節 數值分析 24 第五節 敏感度分析 25 第一小節 需求變異度 26 第二小節 成本參數 27 第四章 非線性模型 29 第一節 非線性需求之特徵化多臂拉霸機庫存模型 29 第二節 數值分析30 第三節 敏感度分析 31 第一小節 需求變異度 31 第二小節 成本參數 33 第五章 結論 35 第一節 研究貢獻 35 第二節 研究假設 36 第三節 未來研究方向 37 參考文獻 38

    Agarwal, A., Foster, D. P., Hsu, D. J., Kakade, S. M., & Rakhlin, A. (2011). Stochastic convex optimization with bandit feedback. Advances in Neural Information Processing Systems, 24.
    Agrawal, R. (1995). Sample mean based index policies by o (log n) regret for the multiarmed bandit problem. Advances in Applied Probability, 27(4), 1054–1078.
    Agrawal, S., & Goyal, N. (2012). Analysis of thompson sampling for the multi-armed bandit problem. In Conference on learning theory, (pp. 39–1). JMLR Workshop and Conference Proceedings.
    Agrawal, S., & Goyal, N. (2013). Thompson sampling for contextual bandits with linear payoffs. In International conference on machine learning, (pp. 127–135). PMLR.
    Agrawal, S., & Jia, R. (2019). Learning in structured mdps with convex cost functions: Improved regret bounds for inventory management. In Proceedings of the 2019 ACM Conference on Economics and Computation, (pp. 743–744).
    Allesiardo, R., F´eraud, R., & Bouneffouf, D. (2014). A neural networks committee for the contextual bandit problem. In International Conference on Neural Information Processing, (pp. 374–381). Springer.
    Amiri, N. H., Udenio, M., & Boute, R. N. (2023). Adaptive multi-armed bandits for non-stationary inventory control. SSRN.
    Arrow, K. J., Karlin, S., Scarf, H. E., Beckmann, M. J., Gessford, J. E., & Muth, R. F. (1958). Studies in the mathematical theory of inventory and production. (No Title).
    Auer, P., Cesa-Bianchi, N., & Fischer, P. (2002). Finite-time analysis of the multiarmed bandit problem. Machine learning, 47(2), 235–256.
    Azoury, K. S. (1985). Bayes solution to dynamic inventory models under unknown demand distribution. Management science, 31(9), 1150–1160.
    Ban, G.-Y., & Rudin, C. (2019). The big data newsvendor: Practical insights from machine learning. Operations Research, 67(1), 90–108.
    Bouneffouf, D., Rish, I., & Aggarwal, C. (2020). Survey on applications of multi-armed and contextual bandits. In 2020 IEEE congress on evolutionary computation (CEC), (pp. 1–8). IEEE.
    Chu, W., Li, L., Reyzin, L., & Schapire, R. (2011). Contextual bandits with linear payoff functions. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, (pp. 208–214). JMLR Workshop and Conference Proceedings.
    Ding, J., Huh, W. T., & Rong, Y. (2024). Feature-based inventory control with censored demand. Manufacturing & Service Operations Management, 26(3), 1157–1172.
    Garivier, A., & Moulines, E. (2008). On upper-confidence bound policies for non-stationary bandit problems (2008). arXiv preprint arXiv:0805.3415.
    Gittins, J. C. (1979). Bandit processes and dynamic allocation indices. Journal of the Royal Statistical Society Series B: Statistical Methodology, 41(2), 148–164.
    Godfrey, G. A., & Powell, W. B. (2001). An adaptive, distribution-free algorithm for the newsvendor problem with censored demands, with applications to inventory and distribution. Management Science, 47(8), 1101–1112.
    Han, J., Hu, M., & Shen, G. (2025). Deep neural newsvendor. Management Science.
    Hannah, L., Powell, W., & Blei, D. (2010). Nonparametric density estimation for stochastic optimization with an observable state variable. Advances in Neural Information Processing Systems, 23.
    Huh, W. T., & Rusmevichientong, P. (2009). A nonparametric asymptotic analysis of inventory planning with censored demand. Mathematics of Operations Research, 34(1), 103–123.
    Kocsis, L., & Szepesv´ari, C. (2006). Discounted ucb. In 2nd PASCAL Challenges Workshop, vol. 2, (pp. 51–134).
    Lai, T. L., & Robbins, H. (1985). Asymptotically efficient adaptive allocation rules. Advances in Applied Mathematics, 6(1), 4–22.
    Langford, J., & Zhang, T. (2007). The epoch-greedy algorithm for multi-armed bandits with side information. Advances in Neural Information Processing Systems, 20.
    Li, L., Chu, W., Langford, J., & Schapire, R. E. (2010). A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th international conference on World wide web, (pp. 661–670).
    Liyanage, L. H., & Shanthikumar, J. G. (2005). A practical inventory control policy using operational statistics. Operations research letters, 33(4), 341–348.
    Oroojlooyjadid, A., Snyder, L. V., & Tak´aˇc, M. (2020). Applying deep learning to the newsvendor problem. Iise Transactions, 52(4), 444–463.
    Robbins, H. (1952). Some aspects of the sequential design of experiments.
    Scarf, H. E., Arrow, K., & Karlin, S. (1957). A min-max solution of an inventory problem. Tech. rep., Rand Corporation Santa Monica.
    Shapiro, A., Dentcheva, D., & Ruszczynski, A. (2021). Lectures on stochastic programming: modeling and theory. SIAM.
    Shi, J. (2022). Application of the model combining demand forecasting and inventory decision in feature based newsvendor problem. Computers & Industrial Engineering, 173, 108709.
    Zhang, L., Yang, J., & Gao, R. (2024). Optimal robust policy for feature-based newsvendor. Management Science, 70(4), 2315–2329.
    Zhao, T., Zhou, W.-X., & Wang, L. (2025). Private optimal inventory policy learning for feature-based newsvendor with unknown demand. Management Science, 71(7), 6092–6111.

    QR CODE