| 研究生: |
譚丞岡 Tan, Cheng-Kang |
|---|---|
| 論文名稱: |
基於常識與脈絡線索強化人與物體之間的互動偵測 Human Object Interaction Detection Enhanced by Common Sense and Contextual Cues |
| 指導教授: |
朱威達
Chu, Wei-Ta |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2025 |
| 畢業學年度: | 113 |
| 語文別: | 英文 |
| 論文頁數: | 45 |
| 中文關鍵詞: | 人與物體互動偵測 、常識 、脈絡線索 、大型語言模型 |
| 外文關鍵詞: | Human-object interaction detection, Common sense, Contextual cues, Large language model |
| 相關次數: | 點閱:83 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
偵測人與物體的互動(Human-Object Interaction, HOI)旨在定位人與物體的配對關係,並辨識其互動類型。近年來,許多 HOI 偵測框架透過視覺語言模型(Visual Language Models, VLMs)或外部資料集來獲取更多語意資訊,展現了亮眼的表現。這些方法主要專注於有效提取 VLMs 中的知識,並常仰賴大量外部資料。然而我們主張並非所有潛在的互動對於每一組人與物體而言都是同等重要的。人與特定物體之間發生某種互動的可能性,取決於該物體的性質。例如「騎」這個動作比較可能發生在腳踏車上,而不是書本上。像「騎書本」這樣的情境,我們會自然判斷為不合理,這類常識已隱含在我們對圖像的理解能力之中。
此外脈絡資訊亦是推論互動類型時不可忽視的關鍵因素。即使是相同的人與物體組合,例如人與汽車,其相對位置的不同也可能導致截然不同的語意判斷。當人位於車內時,較可能對應到「開車」的互動;反之,若人與車保持距離,則更可能表示其正在「檢查車輛」或「觀看汽車」。因此理解人與物體之間的語意關係,不僅需考量物體的類別與屬性,亦應融合其空間配置與常識知識,以更準確、合理地進行互動推論。
受到此觀察的啟發,我們提出一種新方法,結合大型語言模型(Large Language Models, LLMs)以生成常識知識與額外的脈絡線索,進而為不同的人-物-互動組合賦予適當的權重。我們的方法雖然簡單,卻具高度效能,且能無縫整合至現有 HOI偵測框架中以提升整體表現。我們在 HICO-DET 與 V-COCO 資料集上的實驗結果顯示,與既有方法相比,此策略能顯著提升 HOI 偵測的準確率。
Detecting human-object interactions (HOI) involves localizing human-object pairs and identifying their interactions. Recently, several HOI detection frameworks leveraging Visual Language Models (VLMs) or external datasets to get more information have demonstrated promising performance. These approaches focus on efficiently distilling knowledge from pre-trained VLMs and often rely heavily on external datasets. We argue that not all possible interactions are equally relevant to every human-object pair. How likely an interaction happens between a human and a specific object depends on the types of objects. For example, a human more likely ride a bike but unlikely ride a book. This is implicitly embedded into the common sense for us the understand an image.
Moreover, contextual information also plays a crucial role in determining the interaction type. Even when the human and the object categories are the same, the relative spatial configuration may imply different semantics. For example, if a person is inside a car, the interaction is likely to be interpreted as driving a car; whereas if the person is standing away from the car, it could indicate inspecting the car or watching the car. These differences highlight the importance of considering both spatial layout and common sense when inferring HOIs.
Inspired by this observation, we propose leveraging a Large Language Model (LLM) to output common sense and additional contextual information, which can then be used to weight different human-object-interaction tuples. The proposed simple yet powerful idea can be incorporated with existing HOI detection frameworks to make performance improvements. Our experiments on the HICO-DET and the V-COCO datasets demonstrate that this approach significantly improves performance compared to previous methods.
[1] Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In Proceedings of the European Conference on Computer Vision, 2020.
[2] Yu-Wei Chao, Yunfan Liu, Xieyang Liu, Huayi Zeng, and Jia Deng. Learning to detect human-object interactions. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision, 2018.
[3] Mu Chen, Minghan Chen, and Yi Yang. UAHOI: Uncertainty-aware robust interaction learning for HOI detection. Computer Vision and Image Understanding, 247:104091, 2024.
[4] Shuman Fang, Zhiwen Lin, Ke Yan, Jie Li, Xianming Lin, and Rongrong Ji. HODN: Disentangling human-object feature for HOI detection. IEEE Transactions on Multimedia, 26:3125–3136, 2023.
[5] Jiayi Gao, Kongming Liang, Tao Wei, Wei Chen, Zhanyu Ma, and Jun Guo. Dual-prior augmented decoding network for long tail distribution in HOI detection. In Proceedings of the AAAI Conference on Artificial Intelligence, 2024.
[6] Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2014.
[7] Georgia Gkioxari, Ross Girshick, Piotr Dollár, and Kaiming He. Detecting and recognizing human-object interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018.
[8] Yixin Guo, Yu Liu, Jianghao Li, Weimin Wang, and Qi Jia. Unseen No More: Unlocking the potential of CLIP for generative zero-shot HOI detection. In Proceedings of the ACM International Conference on Multimedia, 2024.
[9] Saurabh Gupta and Jitendra Malik. Visual semantic role labeling. arXiv preprint arXiv:1505.04474, 2015.
[10] Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision, 2017.
[11] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2016.
[12] Zhi Hou, Baosheng Yu, Yu Qiao, Xiaojiang Peng, and Dacheng Tao. Affordance transfer learning for human-object interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
[13] Zhi Hou, Baosheng Yu, Yu Qiao, Xiaojiang Peng, and Dacheng Tao. Detecting human-object interaction via fabricated compositional learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
[14] Mingda Jia, Liming Zhao, Ge Li, and Yun Zheng. ContextHOI: Spatial context learning for human-object interaction detection. In Proceedings of the AAAI Conference on Artificial Intelligence, 2025.
[15] Weibo Jiang, Weihong Ren, Jiandong Tian, Liangqiong Qu, Zhiyong Wang, and Honghai Liu. Exploring self-and cross-triplet correlations for human-object interaction detection. In Proceedings of the AAAI Conference on Artificial Intelligence, 2024.
[16] Bumsoo Kim, Junhyun Lee, Jaewoo Kang, Eun-Sol Kim, and Hyunwoo J Kim. Humanobject interaction detection via contextual information distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
[17] Sanghyun Kim, Deunsol Jung, and Minsu Cho. Relational context learning for humanobject interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023.
[18] Ting Lei, Fabian Caba, Qingchao Chen, Hailin Jin, Yuxin Peng, and Yang Liu. Efficient adaptive human-object interaction detection with concept-guided memory. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023.
[19] Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. BLIP: Bootstrapping languageimage pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning, 2022.
[20] Zhuolong Li, Xingao Li, Changxing Ding, and Xiangmin Xu. Disentangled pre-training for human-object interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024.
[21] Yue Liao, Si Liu, Fei Wang, Yanjie Chen, Chen Qian, and Jiashi Feng. PPDM: Parallel point detection and matching for real-time human object interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020.
[22] Yue Liao, Aixi Zhang, Miao Lu, Yongliang Wang, Xiaobo Li, and Si Liu. GEN-VLKT: Simplify association and enhance interaction understanding for HOI detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
[23] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In Proceedings of the European Conference on Computer Vision, 2014.
[24] Ye Liu, Junsong Yuan, and Chang Wen Chen. ConsNet: Learning consistency graph for zero-shot human-object interaction detection. In Proceedings of the ACM International Conference on Multimedia, 2020.
[25] Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
[26] Shuailei Ma, Yuefeng Wang, Shanze Wang, and Ying Wei. FGAHOI: Fine-grained anchors for human-object interaction detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46:2415–2429, 2023.
[27] Yunyao Mao, Jiajun Deng, Wengang Zhou, Li Li, Yao Fang, and Houqiang Li. CLIP4HOI: Towards adapting CLIP for practical zero-shot HOI detection. In Proceedings of the Advances in Neural Information Processing Systems, 2023.
[28] Shan Ning, Longtian Qiu, Yongfei Liu, and Xuming He. HOICLIP: Efficient knowledge transfer for HOI detection with vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023.
[29] Jeeseung Park, Jin-Woo Park, and Jong-Seok Lee. ViPLO: Vision transformer based pose-conditioned self-loop graph for human-object interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023.
[30] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, 2021.
[31] Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: Towards realtime object detection with region proposal networks. In Proceedings of the Advances in Neural Information Processing Systems, 2015.
[32] Masato Tamura, Hiroki Ohashi, and Tomoaki Yoshinaga. QPIC: Query-based pairwise human-object interaction detection with image-wide contextual information. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
[33] Cheng-Kang Tan and Wei-Ta Chu. CS-HOI: Human-object interaction detection enhanced by common sense. In Proceedings of the ACM International Conference on Multimedia in Asia, 2024.
[34] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, 2017.
[35] Bo Wan, Desen Zhou, Yongfei Liu, Rongjie Li, and Xuming He. Pose-aware multilevel feature network for human-object interaction detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019.
[36] Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In Proceedings of the European Conference on Computer Vision, 2018.
[37] Chi Xie, Fangao Zeng, Yue Hu, Shuang Liang, and Yichen Wei. Category query learning for human-object interaction classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023.
[38] Weiying Xue, Qi Liu, Yuxiao Wang, Zhenao Wei, Xiaofen Xing, and Xiangmin Xu. Towards zero-shot human-object interaction detection via vision-language integration. arXiv preprint arXiv:2403.07246, 2024.
[39] Sijie Yan, Yuanjun Xiong, and Dahua Lin. Spatial temporal graph convolutional networks for skeleton-based action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, 2018.
[40] Jie Yang, Bingliang Li, Ailing Zeng, Lei Zhang, and Ruimao Zhang. Open-world human-object interaction detection via multi-modal prompts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024.
[41] Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo. Image captioning with semantic attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2016.
[42] Aixi Zhang, Yue Liao, Si Liu, Miao Lu, Yongliang Wang, Chen Gao, and Xiaobo Li. Mining the benefits of two-stage and one-stage HOI detection. In Proceedings of the Advances in Neural Information Processing Systems, 2021.
[43] Frederic Z Zhang, Dylan Campbell, and Stephen Gould. Efficient two-stage detection of human-object interactions with a novel unary pairwise transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
[44] Yong Zhang, Yingwei Pan, Ting Yao, Rui Huang, Tao Mei, and Chang-Wen Chen. Exploring structure-aware transformer over interaction proposals for human-object interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
[45] Sipeng Zheng, Boshen Xu, and Qin Jin. Open-category human-object interaction pretraining via language modeling framework. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023.
[46] Xubin Zhong, Xian Qu, Changxing Ding, and Dacheng Tao. Glance and gaze: Inferring action-aware points for one-stage human-object interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
[47] Cheng Zou, Bohan Wang, Yue Hu, Junqi Liu, Qian Wu, Yu Zhao, Boxun Li, Chenguang Zhang, Chi Zhang, Yichen Wei, and Jian Sun. End-to end human object interaction detection with HOI transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.