| 研究生: |
劉奕廷 Liu, Yi-Ting |
|---|---|
| 論文名稱: |
以大型語言情境工程探索建築平面圖辨識 Architectural Floor Plan Recognition through LLM Context Engineering |
| 指導教授: |
簡聖芬
Chien, Sheng-Fen |
| 學位類別: |
碩士 Master |
| 系所名稱: |
規劃與設計學院 - 建築學系 Department of Architecture |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 90 |
| 中文關鍵詞: | 多代理協作 、視覺語言模型 、建築知識 |
| 外文關鍵詞: | Multi-Agent Collaboration, Vision-Language Models, Architectural Knowledge |
| 相關次數: | 點閱:85 下載:6 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究將建築平面圖之解讀任務視為一個將二維像素點陣影像轉譯為向量化結構表示的問題。在建築設計前期階段,案例分析高度仰賴人工重複描圖以使用點陣形式的圖面資源,這成為數位工作流與案例知識再利用的實務阻礙。雖然通用多模態大模型在開放詞彙語義理解上具備優勢,但在面對特定專業領域的建築圖學任務時,容易產生空間定位落差與幾何拓樸幻覺。基於此,本研究採用原型開發與定量評估法,提出一個將建築專業能力中的讀圖任務以情境工程落實為運作流程的向量化轉譯架構。本研究提出的研究假設,主張在不微調與不重新訓練通用大模型的前提下,透過情境工程的策略,將平面圖辨識的任務以多代理的階段性編排,從而有效抑制模型的定位幻覺,並使辨識任務的產出成果逐步收斂,從而支持通用模型進行專業讀圖工作的可能性。
文獻回顧確立平面圖資訊重點為連續拓樸元素、離散符號元件與文字尺寸標註等三大類;檢視既有平面圖解讀的規則式與學習式解讀技術的發展脈絡與侷限。並探討視覺語言模型於空間與位置任務中展現的語意優勢與定位落差。整理歸納情境工程與視覺增強機制,作為彌補上述落差的關鍵技術。
本研究提出建築平面圖情境工程轉譯架構,以建築專業者的讀圖思維為基礎,將視覺語言模型視為需要被任務焦點與注意力安排所引導的通用工具,建構包含初探解析、情境擴充、交叉審查與產出裁定四種階段代理單元的共通推論骨架。並依循分類處理策略,針對牆體等連續拓樸元素借鑒稀疏點提示與標記集合機制,將連續像素預測轉變為標記代號的識別與拓樸連接的推理;針對門窗等離散符號元件則提出視覺檢索增強機制,挪用文本分塊的檢索增強生成,執行影像滑動視窗裁切的特徵匹配;對於文字尺寸標註則建立字元與空間幾何的映射關係。再安排子流程在特定推論節點進行跨流程的座標疊合、幾何交集與數值調整,建立一致的整體轉譯成果。
上述架構以住宅平面圖資料集為基礎進行實作檢驗,前處理步驟透過節點標註與座標標準化建立一致的相對數值空間。在四階段代理單元驅動下,依序執行端點初探配對、元件檢索過濾、數值與拓樸性質的交叉審查以及語義的裁定。量化結果顯示,主案例之評估指標自初探階段的58.3%逐步收斂至最終的94.1%,二十個相異案例之平均成效亦從77.2%提升至89.5%,支持提出的架構具有實際成效。最後,系統將轉譯數據透過應用程式介面導入數位編輯環境,自動生成具備參數化編輯能力之建築資訊模型。
本研究之核心價值在於將建築專業的讀圖認知與空間概念,轉化為順利運作且具成效的情境工程工作流程。單扇門開口任務之實證結果,支持通用的大型語言模型在適當情境與幾何規則的強化下,執行專業任務的可行性。然而,實作採用的人工前置標註與檢索品質有其技術侷限,建議未來研究可結合電腦視覺自動生成中介表示、串接牆體轉角以及全面整併其他家具設備子流程等以提升完備性。
This study proposes a context-engineering architecture that translates architectural floor plans from raster images into vectorized structures. Case study analysis has long relied on manual, repetitive tracing, a bottleneck for digital workflows and knowledge reuse. While general multimodal LLMs offer strong semantic understanding, they produce spatial and geometric hallucinations on architectural tasks. Without fine-tuning, this study decomposes floor-plan recognition into a four-stage multi-agent pipeline — Exploration, Augmentation, Verification, and Resolution — treating walls, doors and windows, and text or dimensions with tailored strategies: sparse point prompts and marker sets, retrieval-augmented feature matching, and character-geometry mapping. Tested on a residential dataset, the pipeline raised recognition accuracy from 58.3% to 94.1% on the primary case, and from 77.2% to 89.5% across twenty cases, then auto-generated an editable BIM via API. These results support the feasibility of general LLMs for professional floor-plan reading when guided by appropriate context and geometric rules, though manual pre-annotation and retrieval quality remain current limitations.
Ahmed, S., Liwicki, M., Weber, M., & Dengel, A. (2012). Automatic Room Detection and Room Labeling from Architectural Floor Plans. 2012 10th IAPR International Workshop on Document Analysis Systems, 339–343. https://doi.org/10.1109/DAS.2012.22
Ah-Soon, C., & Tombre, K. (2001). Architectural symbol recognition using a network of constraints. Pattern Recognition Letters, 22(2), 231–248. https://doi.org/10.1016/S0167-8655(00)00091-X
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J. L., Borgeaud, S., … Simonyan, K. (2022). Flamingo: A Visual Language Model for Few-Shot Learning. Advances in Neural Information Processing Systems 35, 35, 23716–23736. https://doi.org/10.52202/068431-1723
Bahng, H., Jahanian, A., Sankaranarayanan, S., & Isola, P. (2022). Exploring Visual Prompts for Adapting Large-Scale Models. arXiv. http://arxiv.org/abs/2203.17274
Baltrusaitis, T., Ahuja, C., & Morency, L.-P. (2019). Multimodal Machine Learning: A Survey and Taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2), 423–443. https://doi.org/10.1109/TPAMI.2018.2798607
Chiang, T.-R., Robinson, J., Yu, X. V., & Yogatama, D. (2024). LocateBench: Evaluating the Locating Ability of Vision Language Models. arXiv. http://arxiv.org/abs/2410.19808
Ching, F. D. K. . (2007). Architecture: Form, Space, & Order (3rd ed.). John Wiley & Sons.
Dodge, S., Xu, J., & Stenger, B. (2017). Parsing floor plan images. 2017 Fifteenth IAPR International Conference on Machine Vision Applications (MVA), 358–361. https://doi.org/10.23919/MVA.2017.7986875
Dosch, P., Tombre, K., Ah-Soon, C., & Masini, G. (2000). A complete system for the analysis of architectural drawings. International Journal on Document Analysis and Recognition, 3(2), 102–116. https://doi.org/10.1007/PL00010901
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. International Conference on Learning Representations. https://openreview.net/forum?id=YicbFdNTTy
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., & Mordatch, I. (2024). Improving Factuality and Reasoning in Language Models through Multiagent Debate. Proceedings of the 41st International Conference on Machine Learning (Vol. 235, pp. 11733–11763). PMLR. https://proceedings.mlr.press/v235/du24e.html
Eastman, C., Teicholz, P., Sacks, R., & Liston, K. (2008). BIM Handbook: A Guide to Building Information Modeling for Owners, Managers, Designers, Engineers, and Contractors. Wiley. https://doi.org/10.1002/9780470261309
Fan, Z., Zhu, L., Li, H., Chen, X., Zhu, S., & Tan, P. (2021). FloorPlanCAD: A Large-Scale CAD Drawing Dataset for Panoptic Symbol Spotting. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 10128–10137.
Gimenez, L., Robert, S., Suard, F., & Zreik, K. (2016). Automatic reconstruction of 3D building models from scanned 2D floor plans. Automation in Construction, 63, 48–56. https://doi.org/10.1016/j.autcon.2015.12.008
Hua, Q., Ye, L., Fu, D., Xiao, Y., Cai, X., Wu, Y., Lin, J., Wang, J., & Liu, P. (2025). Context Engineering 2.0: The Context of Context Engineering. arXiv. http://arxiv.org/abs/2510.26493
Johnson, J., Douze, M., & Jegou, H. (2021). Billion-Scale Similarity Search with GPUs. IEEE Transactions on Big Data, 7(3), 535–547. https://doi.org/10.1109/TBDATA.2019.2921572
Kalervo, A., Ylioinas, J., Häikiö, M., Karhu, A., & Kannala, J. (2019). CubiCasa5K: A Dataset and an Improved Multi-task Model for Floorplan Image Analysis. Image Analysis. SCIA 2019, 28–40. https://doi.org/10.1007/978-3-030-20205-7_3
Kandpal, N., Deng, H., Roberts, A., Wallace, E., & Raffel, C. (2023). Large Language Models Struggle to Learn Long-Tail Knowledge. Proceedings of the 40th International Conference on Machine Learning (Vol. 202, pp. 15696–15707). PMLR. https://proceedings.mlr.press/v202/kandpal23a.html
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., & Girshick, R. (2023). Segment Anything. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), 3992–4003. https://doi.org/10.1109/ICCV51070.2023.00371
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Advances in Neural Information Processing Systems (Vol. 33, pp. 9459–9474). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf
Li, K., Meng, Z., Lin, H., Luo, Z., Tian, Y., Ma, J., Huang, Z., & Chua, T.-S. (2025). ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use. Proceedings of the 33rd ACM International Conference on Multimedia, 8778–8786. https://doi.org/10.1145/3746027.3755688
Li, L. H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.-N., Chang, K.-W., & Gao, J. (2022). Grounded Language-Image Pre-training. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10955–10965. https://doi.org/10.1109/CVPR52688.2022.01069
Li, Z., Yang, B., Liu, Q., Ma, Z., Zhang, S., Yang, J., Sun, Y., Liu, Y., & Bai, X. (2024). Monkey: Image Resolution and Text Label are Important Things for Large Multi-Modal Models. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 26753–26763. https://doi.org/10.1109/CVPR52733.2024.02527
Liu, C., Wu, J., Kohli, P., & Furukawa, Y. (2017). Raster-to-Vector: Revisiting Floorplan Transformation. 2017 IEEE International Conference on Computer Vision (ICCV), 2214–2222. https://doi.org/10.1109/ICCV.2017.241
Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2023). Visual Instruction Tuning. Advances in Neural Information Processing Systems 36, 36, 34892–34916. https://doi.org/10.52202/075280-1516
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., & Neubig, G. (2023). Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Computing Surveys, 55(9), 1–35. https://doi.org/10.1145/3560815
Lu, Y., Yang, J., Shen, Y., & Awadallah, A. (2024). OmniParser for Pure Vision Based GUI Agent. arXiv. http://arxiv.org/abs/2408.00203
Mei, L., Yao, J., Ge, Y., Wang, Y., Bi, B., Cai, Y., Liu, J., Li, M., Li, Z.-Z., Zhang, D., Zhou, C., Mao, J., Xia, T., Guo, J., & Liu, S. (2025). A Survey of Context Engineering for Large Language Models. arXiv. http://arxiv.org/abs/2507.13334
Pizarro, P. N., Hitschfeld, N., Sipiran, I., & Saavedra, J. M. (2022). Automatic floor plan analysis and recognition. Automation in Construction, 140, 104348. https://doi.org/10.1016/j.autcon.2022.104348
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning Transferable Visual Models From Natural Language Supervision. Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 8748–8763). PMLR. https://proceedings.mlr.press/v139/radford21a.html
Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. Advances in Neural Information Processing Systems (Vol. 36, pp. 68539–68551). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2023/file/d842425e4bf79ba039352da0f658a906-Paper-Conference.pdf
Si, C., Zhang, Y., Li, R., Yang, Z., Liu, R., & Yang, D. (2025). Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3956–3974. https://doi.org/10.18653/v1/2025.naacl-long.199
Tombre, K., Ah-Soon, C., Dosch, P., Masini, G., & Tabbone, S. (2000). Stable and Robust Vectorization: How to Make the Right Choices. Graphics Recognition Recent Advances, 3–18. https://doi.org/10.1007/3-540-40953-X_1
Wang, W., Jing, Y., Ding, L., Wang, Y., Shen, L., Luo, Y., Du, B., & Tao, D. (2025). Retrieval-Augmented Perception: High-resolution Image Perception Meets Visual RAG. Proceedings of the 42nd International Conference on Machine Learning (Vol. 267, pp. 63290–63307). PMLR. https://proceedings.mlr.press/v267/wang25at.html
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. v, & Zhou, D. (2022). Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models. Advances in Neural Information Processing Systems 35, 24824–24837. https://doi.org/10.52202/068431-1800
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., & Wang, C. (2024). AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. ICLR 2024 Workshop on Large Language Model (LLM) Agents. https://openreview.net/forum?id=uAjxFFing2
Yamasaki, T., Zhang, J., & Takada, Y. (2018). Apartment Structure Estimation Using Fully Convolutional Networks and Graph Model. Proceedings of the 2018 ACM Workshop on Multimedia for Real Estate Tech, 1–6. https://doi.org/10.1145/3210499.3210528
Yang, J., Zhang, H., Li, F., Zou, X., Li, C., & Gao, J. (2023). Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V. arXiv. http://arxiv.org/abs/2310.11441
Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., & Chen, E. (2024). A survey on multimodal large language models. National Science Review, 11(12). https://doi.org/10.1093/nsr/nwae403
Yin, X., Wonka, P., & Razdan, A. (2009). Generating 3D Building Models from Architectural Drawings: A Survey. IEEE Computer Graphics and Applications, 29(1), 20–30. https://doi.org/10.1109/MCG.2009.9
Zeng, Z., Li, X., Yu, Y. K., & Fu, C.-W. (2019). Deep Floor Plan Recognition Using a Multi-Task Network With Room-Boundary-Guided Attention. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 9095–9103. https://doi.org/10.1109/ICCV.2019.00919
Zhao, W. X., Zhou, K., Li, J., Tang, T., Dong, Z., Hou, Y., Zhang, B., Min, Y., Zhang, J., Liu, P., Wang, X., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., … Wen, J.-R. (2026). A Survey of Large Language Models. Frontiers of Computer Science, 20(12), 2012627. https://doi.org/10.1007/s11704-026-60308-3