簡易檢索 / 詳目顯示

研究生: 徐偉倫
Xu, Wei-Lun
論文名稱: 基於影像辨識與多模態資料融合之成衣工時預測可行性研究
A Feasibility Study of Garment Sewing Time Prediction Using Image Recognition and Multimodal Data Fusion
指導教授: 李昇暾
Li, Sheng-Tun
學位類別: 碩士
Master
系所名稱: 管理學院 - 工業與資訊管理學系
Department of Industrial and Information Management
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 116
中文關鍵詞: 成衣工時預測多模態資料融合影像辨識階層式預測成衣報價
外文關鍵詞: garment sewing time prediction, multimodal data fusion, image recognition, hierarchical prediction, garment quotation
相關次數: 點閱:56下載:2
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究探討影像辨識與多模態資料融合應用於成衣工時預測之可行性。成衣工時估算直接影響報價,傳統流程多仰賴技術人員依款式圖稿、歷史工時與標準工時資料判斷,較難即時取得穩定且一致的參考工時。因此,如何在低資料門檻下建立工時預測模型,為提升報價效率與成本控管能力之重要課題。
    本研究聚焦於下裝褲類成衣之款式總工時預測,使用某大型成衣製造商之歷史款式資料,整合款式圖稿、部位描述文本、部位工時與款式總工時等欄位,建立多模態預測資料集。模型設計上,以款式圖稿影像與部位描述文本作為主要輸入,透過深度學習模型擷取影像與文字特徵,並結合階層式預測、一致性損失、對數轉換與尾端加權等設計,以提升模型對成衣資料結構與右偏工時分布的適應能力。
    實驗結果顯示,完整多模態模型於測試集上取得 R² = 0.7936 、MAE = 2.47 分鐘、RMSE = 4.56 分鐘與 MAPE = 7.02%,優於本研究資料集下之機器學習基準模型與單模態模型。進一步透過集成模型後,R² 提升至 0.8085,MAE 降至 2.34 分鐘,RMSE 降至 4.39 分鐘,MAPE 降至 6.66%。結果顯示,影像資訊雖不適合作為單獨預測來源,但在與部位描述文本結合後,能提供有助於工時判斷之補充線索。
    高工時樣本分析顯示,模型於主要工時區間具有較佳預測能力,但對右側高工時尾端樣本仍存在低估傾向。因此,本研究不將模型定位為取代樣衣細估、標準工時分析或最終人工審核,而是應用於報價初期之圖式粗估階段,提供初步工時建議,並取代或部分取代第一輪人工初估作業;對於中高工時或高風險款式,則可導入人工覆核流程。整體而言,本研究驗證款式圖稿與部位描述文本於成衣報價初期工時預測之可行性,並說明多模態深度學習方法可作為圖式粗估流程中具實務應用潛力之初步估時方法。

    This study examines the feasibility of applying image recognition and multimodal data fusion to garment sewing time prediction. Sewing time directly affects quotation and cost control, while conventional estimation relies on technicians’ judgment using sketches, historical records, and standard time data. Focusing on lower-body garments, this study uses historical pants-style data from a large garment manufacturer to construct a multimodal dataset containing style sketches, part-level descriptions, part-level sewing time, and total style sewing time.
    The proposed model uses style sketch images and part-level descriptions as inputs, extracts visual and textual features through deep learning, and incorporates hierarchical prediction, consistency loss, logarithmic transformation, and tail weighting to address the structure and right-skewed distribution of sewing time data. The complete multimodal model achieved R² = 0.7936 , MAE = 2.47 minutes, RMSE = 4.56 minutes, and MAPE = 7.02%. After applying an ensemble model, R² increased to 0.8085, MAE decreased to 2.34 minutes, RMSE decreased to 4.39 minutes, and MAPE decreased to 6.66%.
    The results show that image information alone is insufficient for stable prediction but provides supplementary cues when combined with part-level descriptions. High-sewing-time analysis shows better performance in the main sewing time range, while right-tail samples tend to be underestimated. Therefore, the model is positioned for sketch-based rough estimation in early quotation, providing preliminary sewing time suggestions rather than replacing detailed estimation or final manual review, and supporting manual review for high-risk styles.

    摘要 I ABSTRACT II 誌謝 IX 目錄 X 表目錄 XIII 圖目錄 XIV 第一章 緒論 1 1.1 研究背景與動機 1 1.2 研究目的 3 1.3 研究範圍 4 1.4 研究流程 5 第二章 文獻探討 7 2.1 視覺深度學習與成衣影像辨識 7 2.1.1 視覺編碼器架構演進 8 2.1.2 預訓練模型與遷移學習 10 2.1.3 成衣領域應用 11 2.2 工時預測方法 12 2.2.1 標準工時系統 13 2.2.2 機器學習與神經網路工時預測方法 14 2.3 多模態融合與階層式預測 16 2.3.1 多模態模型 16 2.3.2 特徵融合策略 17 2.3.3 階層式預測 20 2.4 損失函數設計與集成模型 21 2.4.1 不平衡迴歸與損失函數 21 2.4.2 尾端樣本之分位數加權 25 2.4.3 集成模型(Ensemble Model) 26 第三章 研究方法 28 3.1 研究架構 28 3.2 資料處理 31 3.2.1 影像與文本處理 32 3.2.2 資料切分方法 33 3.3 模型架構 34 3.3.1 特徵編碼 35 3.3.2 跨模態融合 37 3.3.3 階層式預測 38 3.4 訓練設計 40 3.4.1 損失函數 41 3.4.2 尾端樣本加權 42 3.5 實驗設計 46 3.5.1 消融實驗設計 46 3.5.2 模型集成 48 第四章 實驗結果與分析 49 4.1 資料來源 49 4.1.1 資料結構與款式層級分布 49 4.1.2 部位層級工時分布 54 4.2 實驗設定 56 4.2.1 評估指標 58 4.3 實驗結果比較 60 4.3.1 部位層級預測結果 70 4.4 消融實驗與高工時樣本分析 77 4.4.1 消融實驗結果 77 4.4.2 高工時樣本表現 80 4.5 集成與可行性討論 86 第五章 結論與未來研究方向 90 5.1 研究結論與貢獻 90 5.2 未來研究方向 93 參考文獻 95

    Akter, F., Bushra Akhi, A., Khatun, H., & Sultana, N. (2024). Advanced Deep Learning based Apparel Image Classification System. Proceedings of the 3rd International Conference on Computing Advancements, 779–785. https://doi.org/10.1145/3723178.3723281
    Baltrušaitis, T., Ahuja, C., & Morency, L.-P. (2018). Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2), 423–443. https://doi.org/10.1109/TPAMI.2018.2798607
    Breiman, L. (1996). Stacked Regressions. Machine Learning, 24(1), 49–64. https://doi.org/10.1023/A:1018046112532
    Çakıt, E., & Dağdeviren, M. (2023). Comparative analysis of machine learning algorithms for predicting standard time in a manufacturing environment. Artificial Intelligence for Engineering Design, Analysis and Manufacturing, 37, e2. https://doi.org/10.1017/S0890060422000245
    Cao, H., & Ji, X. (2021). Prediction of Garment Production Cycle Time Based on a Neural Network. Fibres and Textiles in Eastern Europe, 29(1(145)), 8–12. https://doi.org/10.5604/01.3001.0014.5036
    Chen, R., Yu, K., Chen, Y., & Xu, Z. (2025). A segmentation method for virtual clothing effect images. Industria Textila, 76(02), 230–236. https://doi.org/10.35530/IT.076.02.2024111
    Chicco, D., Warrens, M. J., & Jurman, G. (2021). The coefficient of determination R-squared is more informative than SMAPE, MAE, MAPE, MSE and RMSE in regression analysis evaluation. PeerJ Computer Science, 7, e623. https://doi.org/10.7717/peerj-cs.623
    Cho, S., Jeon, J., Kim, M., & Kim, J. (2025). Synergy-CLIP: Extending CLIP With Multi-Modal Integration for Robust Representation Learning. IEEE Access, 13, 65630–65642. https://doi.org/10.1109/ACCESS.2025.3559663
    Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. 2009 IEEE Conference on Computer Vision and Pattern Recognition, 248–255. https://doi.org/10.1109/CVPR.2009.5206848
    Dodge, J., Ilharco, G., Schwartz, R., Farhadi, A., Hajishirzi, H., & Smith, N. (2020). Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping (arXiv:2002.06305). arXiv. https://doi.org/10.48550/arXiv.2002.06305
    Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., ... & Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv. https://arxiv.org/abs/2010.11929
    Dubost, F., Adams, H., Yilmaz, P., Bortsova, G., Tulder, G. V., Ikram, M. A., Niessen, W., Vernooij, M. W., & Bruijne, M. D. (2020). Weakly supervised object detection with 2D and 3D regression neural networks. Medical Image Analysis, 65, 101767. https://doi.org/10.1016/j.media.2020.101767
    Efron, B., & Narasimhan, B. (2020). The Automatic Construction of Bootstrap Confidence Intervals. Journal of Computational and Graphical Statistics, 29(3), 608–619. https://doi.org/10.1080/10618600.2020.1714633
    General Sewing Data Limited. (2001). General Sewing Data (GSD) student manual (Issue 04) [Training manual].
    Gilbreth, F. B., & Kent, R. T. (1911). Motion study: A method for increasing the efficiency of the workman. D. Van Nostrand Company.
    Girshick, R. (2015). Fast r-cnn. In Proceedings of the IEEE international conference on computer vision (pp. 1440-1448).
    He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770–778. https://doi.org/10.1109/CVPR.2016.90
    Huber, P. J. (1964). Robust estimation of a location parameter. The Annals of Mathematical Statistics, 35(1), 73-101. https://doi.org/10.1214/aoms/1177703732
    Hessel, M., Modayil, J., Hasselt, H. van, Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., & Silver, D. (2018). Rainbow: Combining Improvements in Deep Reinforcement Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1). https://doi.org/10.1609/aaai.v32i1.11796
    Hyndman, R. J., Ahmed, R. A., Athanasopoulos, G., & Shang, H. L. (2011). Optimal combination forecasts for hierarchical time series. Computational Statistics & Data Analysis, 55(9), 2579–2589. https://doi.org/10.1016/j.csda.2011.03.006
    Jiang, D., & Ye, M. (2023). Cross-modal implicit relation reasoning and aligning for text-to-image person retrieval. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 2787-2797).
    Ju, C., Bibaut, A., & van der Laan, M. (2018). The relative performance of ensemble methods with deep convolutional neural networks for image classification. Journal of Applied Statistics, 45(15), 2800–2818. https://doi.org/10.1080/02664763.2018.1441383
    Jung, W.-K., Kang, J., Kwon, W., & Kim, H. (2025). StitchingNet and deep transfer learning method for sewing stitch defect detection. Journal of Computational Design and Engineering, 12(4), 140–154. https://doi.org/10.1093/jcde/qwaf037
    Koenker, R. (2017). Quantile regression: 40 years on. Annual review of economics, 9, 155-176. https://doi.org/10.1146/annurev-economics-063016-103651
    Koenker, R., & Bassett, G. (1978). Regression Quantiles. Econometrica, 46(1), 33. https://doi.org/10.2307/1913643
    Kuhn, M., & Johnson, K. (2013). Applied Predictive Modeling. Springer Books. https://ideas.repec.org//b/spr/sprbok/978-1-4614-6849-3.html
    Li, J., Subuda, Daorina, Yao, H., Eerdunbilige, & Zhao, H. (2023). Clothing Image Recognition and Classification Based on Deep Learning. In B. J. Jansen, Q. Zhou, & J. Ye (Eds.), Proceedings of the 2nd International Conference on Cognitive Based Information Processing and Applications (CIPA 2022) (Vol. 156, pp. 377–384). Springer Nature Singapore. https://doi.org/10.1007/978-981-19-9376-3_43
    Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., & Chang, K.-W. (2019). VisualBERT: A Simple and Performant Baseline for Vision and Language (arXiv:1908.03557). arXiv. https://doi.org/10.48550/arXiv.1908.03557
    Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., & Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE/CVF International Conference on Computer Vision, 10012–10022. https://doi.org/10.1109/ICCV48922.2021.00986
    Lu, J., Batra, D., Parikh, D., & Lee, S. (2019). ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks (arXiv:1908.02265). arXiv. https://doi.org/10.48550/arXiv.1908.02265
    Maynard, H. B., Stegemerten, G. J., & Schwab, J. L. (1948). Methods-time measurement (pp. x, 292). McGraw-Hill.
    Mehrani, P., & Tsotsos, J. K. (2023). Self-attention in vision transformers performs perceptual grouping, not attention. Frontiers in Computer Science, 5, 1178450. https://doi.org/10.3389/fcomp.2023.1178450
    Meng, F. (2025). Visual design element recognition of garment based on multi-view image fusion. International Journal for Simulation and Multidisciplinary Design Optimization, 16, 5. https://doi.org/10.1051/smdo/2025003
    Mobinizadeh, H., & Lakizadeh, A. (2025). Improving virtual try on clothes using image depth estimation. Scientific Reports, 15(1), 32192. https://doi.org/10.1038/s41598-025-18107-6
    Montesinos López, O. A., Montesinos López, A., & Crossa, J. (2022). Multivariate Statistical Machine Learning Methods for Genomic Prediction. Springer International Publishing. https://doi.org/10.1007/978-3-030-89010-0
    Pan, S., Wang, P., & Yang, C. (2024). Optimization of automatic classification for women’s pants based on the swin transformer model. Fashion and Textiles, 11(1), 42. https://doi.org/10.1186/s40691-024-00408-5
    Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning Transferable Visual Models From Natural Language Supervision (arXiv:2103.00020). arXiv. https://doi.org/10.48550/arXiv.2103.00020
    Ren, L. (2016). Study on standard time of garment sewing based on GSD. 2016 International Conference on Economics, Social Science, Arts, Education and Management Engineering, 87–92.
    Rico Gómez, R., Lorentz, J., Hartmann, T., Goknil, A., Pal Singh, I., Halaç, T. G., & Boruzanlı Ekinci, G. (2024). An AI pipeline for garment price projection using computer vision. Neural Computing and Applications, 36(25), 15631–15651. https://doi.org/10.1007/s00521-024-09901-w
    Rounaghi, M. M., Jarrar, H., & Dana, L.-P. (2021). Implementation of strategic cost management in manufacturing companies: Overcoming costs stickiness and increasing corporate sustainability. Future Business Journal, 7(1), 31. https://doi.org/10.1186/s43093-021-00079-4
    Shin, S. Y., Jo, G., & Wang, G. (2023). A novel method for fashion clothing image classification based on deep learning. Journal of Information and Communication Technology, 22(1), 127-148. https://doi.org/10.32890/jict2023.22.1.6
    Shen, H., & Ji, X. (2024). Optimization of garment sewing operation standard minute value prediction using an IPSO-BP neural network. AUTEX Research Journal, 24(1), 20230034. https://doi.org/10.1515/aut-2023-0034
    Tan, M., & Le, Q. (2019). EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. Proceedings of the 36th International Conference on Machine Learning, 6105–6114. https://proceedings.mlr.press/v97/tan19a.html
    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
    Vijayaraj, A., Vasanth Raj, P. T., Jebakumar, R., Gururama Senthilvel, P., Kumar, N., Suresh Kumar, R., & Dhanagopal, R. (2022). Deep Learning Image Classification for Fashion Design. Wireless Communications and Mobile Computing, 2022(1), 7549397. https://doi.org/10.1155/2022/7549397
    Wolpert, D. H. (1992). Stacked generalization. Neural Networks, 5(2), 241–259. https://doi.org/10.1016/S0893-6080(05)80023-1
    Xiao, F., Sigal, L., & Lee, Y. J. (2017). Weakly-Supervised Visual Grounding of Phrases with Linguistic Structures. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 5253–5262. https://doi.org/10.1109/CVPR.2017.558
    Yang, A., Pan, J., Lin, J., Men, R., Zhang, Y., Zhou, J., & Zhou, C. (2023). Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese (arXiv:2211.01335). arXiv. https://doi.org/10.48550/arXiv.2211.01335
    Yang, Y., Zha, K., Chen, Y., Wang, H., & Katabi, D. (2021, July). Delving into deep imbalanced regression. In International conference on machine learning (pp. 11842-11851). PMLR. https://proceedings.mlr.press/v139/yang21m.html
    Yosinski, J., Clune, J., Bengio, Y., & Lipson, H. (2014). How transferable are features in deep neural networks? Advances in Neural Information Processing Systems, 27. https://proceedings.neurips.cc/paper_files/paper/2014/hash/532a2f85b6977104bc93f8580abbb330-Abstract.html
    Zeiler, M. D., & Fergus, R. (2014). Visualizing and Understanding Convolutional Networks. In D. Fleet, T. Pajdla, B. Schiele, & T. Tuytelaars (Eds.), Computer Vision – ECCV 2014 (Vol. 8689, pp. 818–833). Springer International Publishing. https://doi.org/10.1007/978-3-319-10590-1_53

    下載圖示
    校外:立即公開
    QR CODE