簡易檢索 / 詳目顯示

研究生: 劉奕廷
Liu, Yi-Ting
論文名稱: 以大型語言情境工程探索建築平面圖辨識
Architectural Floor Plan Recognition through LLM Context Engineering
指導教授: 簡聖芬
Chien, Sheng-Fen
學位類別: 碩士
Master
系所名稱: 規劃與設計學院 - 建築學系
Department of Architecture
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 90
中文關鍵詞: 多代理協作 、視覺語言模型 、建築知識
外文關鍵詞: Multi-Agent Collaboration, Vision-Language Models, Architectural Knowledge
相關次數: 點閱:85  下載:6 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究將建築平面圖之解讀任務視為一個將二維像素點陣影像轉譯為向量化結構表示的問題。在建築設計前期階段,案例分析高度仰賴人工重複描圖以使用點陣形式的圖面資源,這成為數位工作流與案例知識再利用的實務阻礙。雖然通用多模態大模型在開放詞彙語義理解上具備優勢,但在面對特定專業領域的建築圖學任務時,容易產生空間定位落差與幾何拓樸幻覺。基於此,本研究採用原型開發與定量評估法,提出一個將建築專業能力中的讀圖任務以情境工程落實為運作流程的向量化轉譯架構。本研究提出的研究假設,主張在不微調與不重新訓練通用大模型的前提下,透過情境工程的策略,將平面圖辨識的任務以多代理的階段性編排,從而有效抑制模型的定位幻覺,並使辨識任務的產出成果逐步收斂,從而支持通用模型進行專業讀圖工作的可能性。
    文獻回顧確立平面圖資訊重點為連續拓樸元素、離散符號元件與文字尺寸標註等三大類;檢視既有平面圖解讀的規則式與學習式解讀技術的發展脈絡與侷限。並探討視覺語言模型於空間與位置任務中展現的語意優勢與定位落差。整理歸納情境工程與視覺增強機制,作為彌補上述落差的關鍵技術。
    本研究提出建築平面圖情境工程轉譯架構,以建築專業者的讀圖思維為基礎,將視覺語言模型視為需要被任務焦點與注意力安排所引導的通用工具,建構包含初探解析、情境擴充、交叉審查與產出裁定四種階段代理單元的共通推論骨架。並依循分類處理策略,針對牆體等連續拓樸元素借鑒稀疏點提示與標記集合機制,將連續像素預測轉變為標記代號的識別與拓樸連接的推理;針對門窗等離散符號元件則提出視覺檢索增強機制,挪用文本分塊的檢索增強生成,執行影像滑動視窗裁切的特徵匹配;對於文字尺寸標註則建立字元與空間幾何的映射關係。再安排子流程在特定推論節點進行跨流程的座標疊合、幾何交集與數值調整,建立一致的整體轉譯成果。
    上述架構以住宅平面圖資料集為基礎進行實作檢驗,前處理步驟透過節點標註與座標標準化建立一致的相對數值空間。在四階段代理單元驅動下,依序執行端點初探配對、元件檢索過濾、數值與拓樸性質的交叉審查以及語義的裁定。量化結果顯示,主案例之評估指標自初探階段的58.3%逐步收斂至最終的94.1%,二十個相異案例之平均成效亦從77.2%提升至89.5%,支持提出的架構具有實際成效。最後,系統將轉譯數據透過應用程式介面導入數位編輯環境,自動生成具備參數化編輯能力之建築資訊模型。
    本研究之核心價值在於將建築專業的讀圖認知與空間概念,轉化為順利運作且具成效的情境工程工作流程。單扇門開口任務之實證結果,支持通用的大型語言模型在適當情境與幾何規則的強化下,執行專業任務的可行性。然而,實作採用的人工前置標註與檢索品質有其技術侷限,建議未來研究可結合電腦視覺自動生成中介表示、串接牆體轉角以及全面整併其他家具設備子流程等以提升完備性。

    This study proposes a context-engineering architecture that translates architectural floor plans from raster images into vectorized structures. Case study analysis has long relied on manual, repetitive tracing, a bottleneck for digital workflows and knowledge reuse. While general multimodal LLMs offer strong semantic understanding, they produce spatial and geometric hallucinations on architectural tasks. Without fine-tuning, this study decomposes floor-plan recognition into a four-stage multi-agent pipeline — Exploration, Augmentation, Verification, and Resolution — treating walls, doors and windows, and text or dimensions with tailored strategies: sparse point prompts and marker sets, retrieval-augmented feature matching, and character-geometry mapping. Tested on a residential dataset, the pipeline raised recognition accuracy from 58.3% to 94.1% on the primary case, and from 77.2% to 89.5% across twenty cases, then auto-generated an editable BIM via API. These results support the feasibility of general LLMs for professional floor-plan reading when guided by appropriate context and geometric rules, though manual pre-annotation and retrieval quality remain current limitations.

    摘要 i Abstract iii 致謝 vii 目錄 viii 表目錄 xi 圖目錄 xii 第1章 緒論 1 1.1 研究背景與動機 1 1.2 多模態預訓練大型語言模型 3 1.3 情境工程 4 1.4 研究課題與目標 6 1.5 研究方法 7 1.6 論文架構 8 1.7 AI工具於撰寫論文時的使用情形 9 第2章 文獻回顧 10 2.1 建築平面圖的資訊構成 10 2.1.1 連續拓樸元素與空間邊界 11 2.1.2 符號元件與功能圖塊 11 2.1.3 尺寸約束與類別標註 12 2.2 自動化解讀技術 13 2.3 視覺語言模型於空間任務之應用與限制 14 2.4 視覺任務中的情境工程與增強機制 15 2.5 小結 16 第3章 建築平面圖的情境工程轉譯架構 17 3.1 整體架構 17 3.1.1 建築圖面內容解析與專業識圖知識 18 3.1.2 建築專業知識之情境工程化編排 19 3.2 階段代理架構 21 3.2.1 初探解析 23 3.2.2 情境擴充 23 3.2.3 交叉審查 24 3.2.4 產出裁定 24 3.3 分類處理策略 24 3.3.1 連續拓樸元素之處理策略 25 3.3.2 離散符號元件之處理策略 25 3.3.3 文字與尺寸標註資訊之處理策略 26 3.3.4 情境工程方法與中介成果互通 27 第4章 實作 28 4.1 資料來源與實作範圍 28 4.2 前處理:手動標註與座標標準化 29 4.3 單扇門開口判讀子流程之實作 31 4.3.1 初探解析:端點之初步配對 32 4.3.2 情境擴充:引入門元件進行過濾 32 4.3.3 交叉審查:以幾何與拓樸規則升降級 33 4.3.4 產出裁定:交還VLM進行語義終審 34 4.4 案例展示與分析 35 4.4.1 主案例之中介成果推演 36 4.4.2 代理單元階段效能 40 4.4.3 情境擴充之視覺檢索品質影響與瓶頸 43 4.5 向量轉譯的檔案整合與產出 45 第5章 結論 47 5.1 研究貢獻 47 5.2 研究反思 48 參考文獻 50 附錄一 階段提示詞 53 附錄二 step01a_attempt.py 56 附錄三 step01b_filter.py 59 附錄四 step01c_compute.py 64 附錄五 step01d_decide.py 72 附錄六 二十個任務影像的各階段F1-score 76

    Ahmed, S., Liwicki, M., Weber, M., & Dengel, A. (2012). Automatic Room Detection and Room Labeling from Architectural Floor Plans. 2012 10th IAPR International Workshop on Document Analysis Systems, 339–343. https://doi.org/10.1109/DAS.2012.22
    Ah-Soon, C., & Tombre, K. (2001). Architectural symbol recognition using a network of constraints. Pattern Recognition Letters, 22(2), 231–248. https://doi.org/10.1016/S0167-8655(00)00091-X
    Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J. L., Borgeaud, S., … Simonyan, K. (2022). Flamingo: A Visual Language Model for Few-Shot Learning. Advances in Neural Information Processing Systems 35, 35, 23716–23736. https://doi.org/10.52202/068431-1723
    Bahng, H., Jahanian, A., Sankaranarayanan, S., & Isola, P. (2022). Exploring Visual Prompts for Adapting Large-Scale Models. arXiv. http://arxiv.org/abs/2203.17274
    Baltrusaitis, T., Ahuja, C., & Morency, L.-P. (2019). Multimodal Machine Learning: A Survey and Taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2), 423–443. https://doi.org/10.1109/TPAMI.2018.2798607
    Chiang, T.-R., Robinson, J., Yu, X. V., & Yogatama, D. (2024). LocateBench: Evaluating the Locating Ability of Vision Language Models. arXiv. http://arxiv.org/abs/2410.19808
    Ching, F. D. K. . (2007). Architecture: Form, Space, & Order (3rd ed.). John Wiley & Sons.
    Dodge, S., Xu, J., & Stenger, B. (2017). Parsing floor plan images. 2017 Fifteenth IAPR International Conference on Machine Vision Applications (MVA), 358–361. https://doi.org/10.23919/MVA.2017.7986875
    Dosch, P., Tombre, K., Ah-Soon, C., & Masini, G. (2000). A complete system for the analysis of architectural drawings. International Journal on Document Analysis and Recognition, 3(2), 102–116. https://doi.org/10.1007/PL00010901
    Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. International Conference on Learning Representations. https://openreview.net/forum?id=YicbFdNTTy
    Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., & Mordatch, I. (2024). Improving Factuality and Reasoning in Language Models through Multiagent Debate. Proceedings of the 41st International Conference on Machine Learning (Vol. 235, pp. 11733–11763). PMLR. https://proceedings.mlr.press/v235/du24e.html
    Eastman, C., Teicholz, P., Sacks, R., & Liston, K. (2008). BIM Handbook: A Guide to Building Information Modeling for Owners, Managers, Designers, Engineers, and Contractors. Wiley. https://doi.org/10.1002/9780470261309
    Fan, Z., Zhu, L., Li, H., Chen, X., Zhu, S., & Tan, P. (2021). FloorPlanCAD: A Large-Scale CAD Drawing Dataset for Panoptic Symbol Spotting. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 10128–10137.
    Gimenez, L., Robert, S., Suard, F., & Zreik, K. (2016). Automatic reconstruction of 3D building models from scanned 2D floor plans. Automation in Construction, 63, 48–56. https://doi.org/10.1016/j.autcon.2015.12.008
    Hua, Q., Ye, L., Fu, D., Xiao, Y., Cai, X., Wu, Y., Lin, J., Wang, J., & Liu, P. (2025). Context Engineering 2.0: The Context of Context Engineering. arXiv. http://arxiv.org/abs/2510.26493
    Johnson, J., Douze, M., & Jegou, H. (2021). Billion-Scale Similarity Search with GPUs. IEEE Transactions on Big Data, 7(3), 535–547. https://doi.org/10.1109/TBDATA.2019.2921572
    Kalervo, A., Ylioinas, J., Häikiö, M., Karhu, A., & Kannala, J. (2019). CubiCasa5K: A Dataset and an Improved Multi-task Model for Floorplan Image Analysis. Image Analysis. SCIA 2019, 28–40. https://doi.org/10.1007/978-3-030-20205-7_3
    Kandpal, N., Deng, H., Roberts, A., Wallace, E., & Raffel, C. (2023). Large Language Models Struggle to Learn Long-Tail Knowledge. Proceedings of the 40th International Conference on Machine Learning (Vol. 202, pp. 15696–15707). PMLR. https://proceedings.mlr.press/v202/kandpal23a.html
    Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., & Girshick, R. (2023). Segment Anything. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), 3992–4003. https://doi.org/10.1109/ICCV51070.2023.00371
    Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Advances in Neural Information Processing Systems (Vol. 33, pp. 9459–9474). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf
    Li, K., Meng, Z., Lin, H., Luo, Z., Tian, Y., Ma, J., Huang, Z., & Chua, T.-S. (2025). ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use. Proceedings of the 33rd ACM International Conference on Multimedia, 8778–8786. https://doi.org/10.1145/3746027.3755688
    Li, L. H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.-N., Chang, K.-W., & Gao, J. (2022). Grounded Language-Image Pre-training. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10955–10965. https://doi.org/10.1109/CVPR52688.2022.01069
    Li, Z., Yang, B., Liu, Q., Ma, Z., Zhang, S., Yang, J., Sun, Y., Liu, Y., & Bai, X. (2024). Monkey: Image Resolution and Text Label are Important Things for Large Multi-Modal Models. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 26753–26763. https://doi.org/10.1109/CVPR52733.2024.02527
    Liu, C., Wu, J., Kohli, P., & Furukawa, Y. (2017). Raster-to-Vector: Revisiting Floorplan Transformation. 2017 IEEE International Conference on Computer Vision (ICCV), 2214–2222. https://doi.org/10.1109/ICCV.2017.241
    Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2023). Visual Instruction Tuning. Advances in Neural Information Processing Systems 36, 36, 34892–34916. https://doi.org/10.52202/075280-1516
    Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., & Neubig, G. (2023). Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Computing Surveys, 55(9), 1–35. https://doi.org/10.1145/3560815
    Lu, Y., Yang, J., Shen, Y., & Awadallah, A. (2024). OmniParser for Pure Vision Based GUI Agent. arXiv. http://arxiv.org/abs/2408.00203
    Mei, L., Yao, J., Ge, Y., Wang, Y., Bi, B., Cai, Y., Liu, J., Li, M., Li, Z.-Z., Zhang, D., Zhou, C., Mao, J., Xia, T., Guo, J., & Liu, S. (2025). A Survey of Context Engineering for Large Language Models. arXiv. http://arxiv.org/abs/2507.13334
    Pizarro, P. N., Hitschfeld, N., Sipiran, I., & Saavedra, J. M. (2022). Automatic floor plan analysis and recognition. Automation in Construction, 140, 104348. https://doi.org/10.1016/j.autcon.2022.104348
    Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning Transferable Visual Models From Natural Language Supervision. Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 8748–8763). PMLR. https://proceedings.mlr.press/v139/radford21a.html
    Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. Advances in Neural Information Processing Systems (Vol. 36, pp. 68539–68551). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2023/file/d842425e4bf79ba039352da0f658a906-Paper-Conference.pdf
    Si, C., Zhang, Y., Li, R., Yang, Z., Liu, R., & Yang, D. (2025). Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3956–3974. https://doi.org/10.18653/v1/2025.naacl-long.199
    Tombre, K., Ah-Soon, C., Dosch, P., Masini, G., & Tabbone, S. (2000). Stable and Robust Vectorization: How to Make the Right Choices. Graphics Recognition Recent Advances, 3–18. https://doi.org/10.1007/3-540-40953-X_1
    Wang, W., Jing, Y., Ding, L., Wang, Y., Shen, L., Luo, Y., Du, B., & Tao, D. (2025). Retrieval-Augmented Perception: High-resolution Image Perception Meets Visual RAG. Proceedings of the 42nd International Conference on Machine Learning (Vol. 267, pp. 63290–63307). PMLR. https://proceedings.mlr.press/v267/wang25at.html
    Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. v, & Zhou, D. (2022). Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models. Advances in Neural Information Processing Systems 35, 24824–24837. https://doi.org/10.52202/068431-1800
    Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., & Wang, C. (2024). AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. ICLR 2024 Workshop on Large Language Model (LLM) Agents. https://openreview.net/forum?id=uAjxFFing2
    Yamasaki, T., Zhang, J., & Takada, Y. (2018). Apartment Structure Estimation Using Fully Convolutional Networks and Graph Model. Proceedings of the 2018 ACM Workshop on Multimedia for Real Estate Tech, 1–6. https://doi.org/10.1145/3210499.3210528
    Yang, J., Zhang, H., Li, F., Zou, X., Li, C., & Gao, J. (2023). Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V. arXiv. http://arxiv.org/abs/2310.11441
    Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., & Chen, E. (2024). A survey on multimodal large language models. National Science Review, 11(12). https://doi.org/10.1093/nsr/nwae403
    Yin, X., Wonka, P., & Razdan, A. (2009). Generating 3D Building Models from Architectural Drawings: A Survey. IEEE Computer Graphics and Applications, 29(1), 20–30. https://doi.org/10.1109/MCG.2009.9
    Zeng, Z., Li, X., Yu, Y. K., & Fu, C.-W. (2019). Deep Floor Plan Recognition Using a Multi-Task Network With Room-Boundary-Guided Attention. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 9095–9103. https://doi.org/10.1109/ICCV.2019.00919
    Zhao, W. X., Zhou, K., Li, J., Tang, T., Dong, Z., Hou, Y., Zhang, B., Min, Y., Zhang, J., Liu, P., Wang, X., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., … Wen, J.-R. (2026). A Survey of Large Language Models. Frontiers of Computer Science, 20(12), 2012627. https://doi.org/10.1007/s11704-026-60308-3

    下載圖示
    校外:立即公開
    QR CODE