| 研究生: |
鄭敬恒 Cheng, Ching-Heng |
|---|---|
| 論文名稱: |
結合擴散式影像修復與可控散景合成之兩階段單張影像重聚焦框架 A Two-Stage Framework for Single-Image Refocusing with Diffusion-Based Restoration and Controllable Bokeh Synthesis |
| 指導教授: |
許志仲
Hsu, Chih-Chung 鄭順林 Jeng, Shuen-Lin |
| 學位類別: |
碩士 Master |
| 系所名稱: |
管理學院 - 數據科學研究所 Institute of Data Science |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 112 |
| 中文關鍵詞: | 影像重聚焦 、失焦去模糊 、散景渲染 |
| 外文關鍵詞: | Image Refocusing, Defocus Deblurring, Bokeh Rendering |
| ORCID: | 0009-0003-0887-9734 |
| 相關次數: | 點閱:5 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究提出一種結合潛在擴散模型與輕量化渲染之兩階段影像重聚焦框架,旨在同時提升失焦去模糊與散景生成之視覺品質。在去模糊階段,鑑於失焦去模糊屬於高度病態的逆問題,傳統像素空間模型(如 Restormer)雖在像素級指標表現優異,但在處理嚴重失焦模糊時,容易產生過於平滑且缺乏細節的視覺結果。為解決此問題,本研究利用預訓練擴散模型之強大生成先驗,並結合影像導引與高頻資訊,以在維持輸入影像結構一致性的同時,重建高保真度且具豐富細節的清晰影像。
在散景生成階段,針對擴散模型在處理高解析度影像時效率較低且容易產生像素偏移之限制,本研究採用輕量化之像素空間渲染模型。透過使用者指定之焦點位置與對應的失焦地圖(Defocus Map)作為導引,此階段能在維持高運算效率與像素級對齊的同時,生成自然且擬真的散景效果。實驗結果顯示,本框架能有效恢復受失焦模糊影響之影像細節,並根據指定焦平面生成自然且連續的景深效果。整體而言,本系統能產出兼具清晰結構細節與擬真散景表現的高品質重聚焦影像,展現其於計算攝影與影像後製應用中的潛力。
This thesis presents a two-stage framework for single-image refocusing that combines diffusion-based restoration with lightweight bokeh rendering. The first stage addresses defocus deblurring as an ill-posed restoration problem: when blur is severe, pixel-space regression models can favor conservative, over-smoothed predictions. To recover perceptually plausible details while retaining the input layout, the proposed deblurring stage adapts a pre-trained latent diffusion model with image-based guidance and high-frequency decoder refinement. The second stage uses a pixel-space renderer conditioned on defocus maps estimated from user-specified focal positions. This design avoids applying a heavy generative model to bokeh synthesis, where preserving image alignment and focused structures is more important than regenerating content. The experiments evaluate the two stages separately and as a complete refocusing pipeline, showing that the framework improves perceptual defocus restoration, supports controllable depth-of-field synthesis, and provides an efficient route for post-capture refocusing.
[1] A. Abuolaim and M. S. Brown, “Defocus deblurring using dual-pixel data,” in European conference on computer vision. Springer, 2020, pp. 111–126.
[2] A. Abuolaim, M. Delbracio, D. Kelly, M. S. Brown, and P. Milanfar, “Learning to reduce defocus blur by realistically modeling dual-pixel data,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 2289–2298.
[3] E. Agustsson and R. Timofte, “Ntire 2017 challenge on single image superresolution: Dataset and study,” in Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 2017, pp. 126–135.
[4] S. K. Aithal, P. Maini, Z. C. Lipton, and J. Z. Kolter, “Understanding hallucinations in diffusion models through mode interpolation,” Advances in neural information processing systems, vol. 37, pp. 134 614–134 644, 2024.
[5] H. Alzayer, A. Abuolaim, L. C. Chan, Y. Yang, Y. C. Lou, J.-B. Huang, and A. Kar, “Dc2: Dual-camera defocus control by learning to refocus,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 21 488–21 497.
[6] M. R. Banham and A. K. Katsaggelos, “Digital image restoration,” IEEE signal processing magazine, vol. 14, no. 2, pp. 24–41, 2002.
[7] S. A. Biyouki and H. Hwangbo, “A comprehensive survey on deep neural image deblurring,” arXiv preprint arXiv:2310.04719, 2023.
[8] Y. Blau and T. Michaeli, “The perception-distortion tradeoff,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6228– 6237.
[9] B. Busam, M. Hog, S. McDonagh, and G. Slabaugh, “Sterefo: Efficient image refocusing with stereo vision,” in Proceedings of the IEEE/CVF international conference on computer vision workshops, 2019, pp. 0–0.
[10] K. Chen and Y. Liu, “Efficient defocus deblurring networks based on diffusion models,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5.
[11] L. Chen, X. Chu, X. Zhang, and J. Sun, “Simple baselines for image restoration,” in European conference on computer vision. Springer, 2022, pp. 17–33.
[12] L. Chen, X. Lu, J. Zhang, X. Chu, and C. Chen, “Hinet: Half instance normalization network for image restoration,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 182–192.
[13] L. Chesser, T. Carbone, and A. Zahid, “The unsplash dataset,” Available at unsplash.com/data, 2026. [Online]. Available: https://unsplash.com/data
[14] F.-A. Croitoru, V. Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,” IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 9, pp. 10 850–10 869, 2023.
[15] K. Ding, K. Ma, S. Wang, and E. P. Simoncelli, “Image quality assessment: Unifying structure and texture similarity,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 12 544–12 553.
[16] G. D. Evangelidis and E. Z. Psarakis, “Parametric image alignment using enhanced correlation coefficient maximization,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 30, no. 10, pp. 1858–1865, 2008.
[17] H. Feng, H. Zhou, T. Ye, S. Chen, and L. Zhu, “Residual diffusion deblurring model for single image defocus deblurring,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 3, 2025, pp. 2960–2968.
[18] M. A. Fischler and R. C. Bolles, “Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography,” Communications of the ACM, vol. 24, no. 6, pp. 381–395, 1981.
[19] A. Fortes, T. Wei, S. Zhou, and X. Pan, “Bokeh diffusion: Defocus blur control in text-to-image diffusion models,” in Proceedings of the SIGGRAPH Asia 2025 Conference Papers, 2025, pp. 1–11.
[20] T. Garber and T. Tirer, “Image restoration by denoising diffusion models with iteratively preconditioned guidance,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 25 245–25 254.
[21] C. He, Y. Shen, C. Fang, F. Xiao, L. Tang, Y. Zhang, W. Zuo, Z. Guo, and X. Li, “Diffusion models in low-level vision: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.
[22] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in Advances in Neural Information Processing Systems, vol. 30, 2017.
[23] J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems, vol. 33, pp. 6840–6851, 2020.
[24] A. Ignatov, J. Patel, and R. Timofte, “Rendering natural camera bokeh effect with deep learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 418–419.
[25] A. Ignatov, R. Timofte, M. Qian, C. Qiao, J. Lin, Z. Guo, C. Li, C. Leng, J. Cheng, J. Peng et al., “Aim 2020 challenge on rendering realistic bokeh,” in European Conference on Computer Vision. Springer, 2020, pp. 213–228.
[26] A. Karaali and C. R. Jung, “Edge-based defocus blur estimation with adaptive scale selection,” IEEE Transactions on Image Processing, vol. 27, no. 3, pp. 1126– 1137, 2017.
[27] J. Ke, Q. Wang, Y. Wang, P. Milanfar, and F. Yang, “Musiq: Multi-scale image quality transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 5148–5157.
[28] C. Kolb, D. Mitchell, and P. Hanrahan, “A realistic camera model for computer graphics,” in Proceedings of the 22nd Annual Conference on Computer Graphics and Interactive Techniques, ser. SIGGRAPH ’95. New York, NY, USA: Association for Computing Machinery, 1995, p. 317–324.
[29] L. Kong, J. Dong, J. Tang, M.-H. Yang, and J. Pan, “Efficient visual state space model for image deblurring,” in Proceedings of the computer vision and pattern recognition conference, 2025, pp. 12 710–12 719.
[30] L. Kong, J. Zhang, D. Zou, F. L. Wang, J. S. Ren, X. Wu, J. Dong, and J. Pan, “Deblurdiff: Real-world image deblurring with generative diffusion models,” Advances in Neural Information Processing Systems, vol. 38, pp. 154 855–154 873, 2026.
[31] M. Kraus and M. Strengert, “Depth-of-field rendering by pyramidal image processing,” in Computer graphics forum, vol. 26, no. 3. Wiley Online Library, 2007, pp. 645–654.
[32] B. F. Labs, S. Batifol, A. Blattmann, F. Boesel, S. Consul, C. Diagne, T. Dockhorn, J. English, Z. English, P. Esser, S. Kulal, K. Lacey, Y. Levi, C. Li, D. Lorenz, J. M¨uller, D. Podell, R. Rombach, H. Saini, A. Sauer, and L. Smith, “Flux.1 kontext: Flow matching for in-context image generation and editing in latent space,” 2025. [Online]. Available: https://arxiv.org/abs/2506.15742
[33] C. Ledig, L. Theis, F. Husz´ar, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang et al., “Photo-realistic single image super-resolution using a generative adversarial network,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 4681–4690.
[34] J. Lee, H. Son, J. Rim, S. Cho, and S. Lee, “Iterative filter adaptive network for single image defocus deblurring,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 2034–2042.
[35] S. Lee, E. Eisemann, and H.-P. Seidel, “Depth-of-field rendering with multiview synthesis,” ACM Transactions on Graphics (TOG), vol. 28, no. 5, pp. 1–6, 2009.
[36] A. Levin, R. Fergus, F. Durand, and W. T. Freeman, “Image and depth from a conventional camera with a coded aperture,” ACM transactions on graphics (TOG), vol. 26, no. 3, pp. 70–es, 2007.
[37] X. Li, S. Wu, Q. Zhu, S. Xie, and S. Agaian, “Single image defocus deblurring via multimodal-guided diffusion and depth-aware fusion,” Pattern Recognition, p. 112133, 2025.
[38] X. Li, Y. Ren, X. Jin, C. Lan, X. Wang, W. Zeng, X. Wang, and Z. Chen, “Diffusion models for image restoration and enhancement: A comprehensive survey,” International Journal of Computer Vision, vol. 133, no. 11, pp. 8078–8108, 2025.
[39] Y. Li, H. Fang, X. Lei, Q. Wang, G. Hu, J. Dong, Z. Li, J. Lin, Q. Liu, and X. Song, “Real-world defocus deblurring via score-based diffusion models,” Scientific Reports, vol. 15, no. 1, p. 22942, 2025.
[40] H. Liang, S. Chai, X. Zhao, and J. Kan, “Swin-diff: a single defocus image deblurring network based on diffusion model,” Complex & Intelligent Systems, vol. 11, no. 3, p. 170, 2025.
[41] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte, “Swinir: Image restoration using swin transformer,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 1833–1844.
[42] P. Liang, J. Jiang, X. Liu, and J. Ma, “Decoupling image deblurring into twofold: A hierarchical model for defocus deblurring,” IEEE Transactions on Computational Imaging, vol. 10, pp. 1207–1220, 2024.
[43] X. Lin, J. He, Z. Chen, Z. Lyu, B. Dai, F. Yu, Y. Qiao, W. Ouyang, and C. Dong, “Diffbir: Toward blind image restoration with generative diffusion prior,” in European conference on computer vision. Springer, 2024, pp. 430–448.
[44] D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International journal of computer vision, vol. 60, no. 2, pp. 91–110, 2004.
[45] X. Luo, J. Peng, K. Xian, Z. Wu, and Z. Cao, “Defocus to focus: Photo-realistic bokeh rendering by fusing defocus and radiance priors,” Information Fusion, vol. 89, pp. 320–335, 2023.
[46] Z. Luo, F. K. Gustafsson, Z. Zhao, J. Sj¨olund, and T. B. Sch¨on, “Image restoration with mean-reverting stochastic differential equations,” arXiv preprint arXiv:2301.11699, 2023.
[47] ——, “Refusion: Enabling large-size realistic image restoration with latent-space diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 1680–1691.
[48] A. Mittal, A. K. Moorthy, and A. C. Bovik, “No-reference image quality assessment in the spatial domain,” IEEE Transactions on Image Processing, vol. 21, no. 12, pp. 4695–4708, 2012.
[49] A. Mittal, R. Soundararajan, and A. C. Bovik, “Making a completely blind image quality analyzer,” IEEE Signal Processing Letters, vol. 20, no. 3, pp. 209–212, 2013.
[50] C.-W. T. Mu, C.-D. Fan, J.-B. Huang, and Y.-L. Liu, “Generative refocusing: Flexible defocus control from a single image,” arXiv preprint arXiv:2512.16923, 2025.
[51] H. Nasse, “Depth of field and bokeh,” Carl Zeiss camera lens division report, vol. 1, no. 4, p. 316, 2010.
[52] R. Ng, M. Levoy, M. Br´edif, G. Duval, M. Horowitz, and P. Hanrahan, “Light field photography with a hand-held plenopic camera,” Technical Report CTSR 2005-02, vol. CTSR, 01 2005.
[53] J. Peng, Z. Cao, X. Luo, H. Lu, K. Xian, and J. Zhang, “Bokehme: When neural rendering meets classical rendering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 16 283–16 292.
[54] J. Peng, J. Zhang, X. Luo, H. Lu, K. Xian, and Z. Cao, “Mpib: An mpi-based bokeh rendering framework for realistic partial occlusion effects,” in European Conference on Computer Vision. Springer, 2022, pp. 590–607.
[55] M. Pharr, W. Jakob, and G. Humphreys, Physically Based Rendering: From Theory to Implementation, 3rd ed. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 2016.
[56] M. Potmesil and I. Chakravarty, “A lens and aperture camera model for synthetic image generation,” ACM SIGGRAPH Computer Graphics, vol. 15, no. 3, pp. 297– 305, 1981.
[57] Y. Quan, Z. Wu, and H. Ji, “Gaussian kernel mixture network for single image defocus deblurring,” Advances in Neural Information Processing Systems, vol. 34, pp. 20 812–20 824, 2021.
[58] Y. Quan, X. Yao, and H. Ji, “Single image defocus deblurring via implicit neural inverse kernels,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 12 600–12 610.
[59] M. Ren, M. Delbracio, H. Talebi, G. Gerig, and P. Milanfar, “Multiscale structure guided diffusion for image deblurring,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 10 721–10 733.
[60] M. Roberts, J. Ramapuram, A. Ranjan, A. Kumar, M. A. Bautista, N. Paczan, R. Webb, and J. M. Susskind, “Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10 912–10 922.
[61] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10 684–10 695.
[62] L. Ruan, B. Chen, J. Li, and M.-L. Lam, “Aifnet: All-in-focus image restoration network using a light field-based dataset,” IEEE Transactions on Computational Imaging, vol. 7, pp. 675–688, 2021.
[63] L. Ruan, B. Chen, J. Li, and M. Lam, “Learning to deblur using light field generated and real defocus images,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 16 304–16 313.
[64] M. S. Sajjadi, B. Scholkopf, and M. Hirsch, “Enhancenet: Single image superresolution through automated texture synthesis,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 4491–4500.
[65] P. Sakurikar, I. Mehta, V. N. Balasubramanian, and P. Narayanan, “Refocusgan: Scene refocusing using a single image,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 497–512.
[66] T. Seizinger, F.-A. Vasluianu, J. Chen, Z. Zhou, Z. Wu, R. Timofte, D. Zhang, Y. Lin, Q. Yan, J. Chen et al., “The first controllable bokeh rendering challenge at ntire 2026,” arXiv preprint arXiv:2605.05510, 2026.
[67] T. Seizinger, F.-A. Vasluianu, M. V. Conde, Z. Wu, and R. Timofte, “Bokehlicious: Photorealistic bokeh rendering with controllable apertures,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 8908–8917.
[68] Y. Sheng, Z. Yu, L. Ling, Z. Cao, X. Zhang, X. Lu, K. Xian, H. Lin, and B. Benes, “Dr. bokeh: Differentiable occlusion-aware bokeh rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 4515–4525.
[69] L. Shi and S. Chen, “Bokeh rendering via diffusion models without depth priors,” in 2025 International Conference on Machine Learning, Computational Intelligence and Pattern Recognition (MLCIPR). IEEE, 2025, pp. 117–120.
[70] H. Son, J. Lee, S. Cho, and S. Lee, “Single image defocus deblurring using kernelsharing parallel atrous convolutions,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 2642–2650.
[71] M. Subbarao and G. Surya, “Depth from defocus: A spatial domain approach,” International Journal of computer vision, vol. 13, no. 3, pp. 271–294, 1994.
[72] N. Wadhwa, R. Garg, D. E. Jacobs, B. E. Feldman, N. Kanazawa, R. Carroll, Y. Movshovitz-Attias, J. T. Barron, Y. Pritch, and M. Levoy, “Synthetic depth-offield with a single-camera mobile phone,” ACM Transactions on Graphics (ToG), vol. 37, no. 4, pp. 1–13, 2018.
[73] J. Wang, K. C. K. Chan, and C. C. Loy, “Exploring clip for assessing the look and feel of images,” in Proceedings of the AAAI Conference on Artificial Intelligence, 2023.
[74] L. Wang, X. Shen, J. Zhang, O. Wang, Z. Lin, C.-Y. Hsieh, S. Kong, and H. Lu, “Deeplens: Shallow depth of field from a single image,” arXiv preprint arXiv:1810.08100, 2018.
[75] Y. Wang, X. Chen, X. Xu, Y. Liu, and H. Zhao, “Diffcamera: Arbitrary refocusing on images,” in Proceedings of the SIGGRAPH Asia 2025 Conference Papers, 2025, pp. 1–10.
[76] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,” IEEE Transactions on Image Processing, vol. 13, no. 4, pp. 600–612, 2004.
[77] J. Whang, M. Delbracio, H. Talebi, C. Saharia, A. G. Dimakis, and P. Milanfar, “Deblurring via stochastic refinement,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 16 293–16 303.
[78] R. Wu, T. Yang, L. Sun, Z. Zhang, S. Li, and L. Zhang, “Seesr: Towards semanticsaware real-world image super-resolution,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 25 456–25 467.
[79] S. Xin, N. Wadhwa, T. Xue, J. T. Barron, P. P. Srinivasan, J. Chen, I. Gkioulekas, and R. Garg, “Defocus map estimation and deblurring from a single dual-pixel image,” IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
[80] L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” Advances in Neural Information Processing Systems, vol. 37, pp. 21 875–21 911, 2024.
[81] S. Yang, T. Wu, S. Shi, S. Lao, Y. Gong, M. Cao, J. Wang, and Y. Yang, “Maniqa: Multi-dimension attention network for no-reference image quality assessment,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2022, pp. 1191–1200.
[82] T. Yang, R. Wu, P. Ren, X. Xie, and L. Zhang, “Pixel-aware stable diffusion for realistic image super-resolution and personalized stylization,” in European conference on computer vision. Springer, 2024, pp. 74–91.
[83] Y. Yang, H. Lin, Z. Yu, S. Paris, and J. Yu, “Virtual dslr: High quality dynamic depth-of-field synthesis on mobile platforms,” Electronic Imaging, vol. 28, pp. 1–9, 2016.
[84] S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M.-H. Yang, “Restormer: Efficient transformer for high-resolution image restoration,” in CVPR, 2022.
[85] L. Zhai, Y. Wang, S. Cui, and Y. Zhou, “A comprehensive review of deep learningbased real-world image restoration,” Ieee Access, vol. 11, pp. 21 049–21 067, 2023.
[86] K. Zhang, W. Ren, W. Luo, W.-S. Lai, B. Stenger, M.-H. Yang, and H. Li, “Deep image deblurring: A survey,” International Journal of Computer Vision, vol. 130, no. 9, pp. 2103–2130, 2022.
[87] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 586–595.
[88] X. Zhang, R. Wang, X. Jiang, W. Wang, and W. Gao, “Spatially variant defocus blur map estimation and deblurring from a single image,” Journal of Visual Communication and Image Representation, vol. 35, pp. 257–264, 2016.
[89] Y. Zhang, X. Huang, J. Ma, Z. Li, Z. Luo, Y. Xie, Y. Qin, T. Luo, Y. Li, S. Liu et al., “Recognize anything: A strong image tagging model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 1724–1732.
[90] Z. Zhang, L. G. Foo, H. Rahmani, J. Liu, and D. W. Soh, “Performing defocus deblurring by modeling its formation process,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2025, pp. 5791– 5801.
[91] C. Zhu, Q. Fan, Q. Zhang, J. Chen, H. Zhang, C. Xu, and B. Shi, “Bokehdiff: Neural lens blur with one-step diffusion,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 9508–9518.
[92] S. Zhuo and T. Sim, “Defocus map estimation from a single image,” Pattern Recognition, vol. 44, no. 9, pp. 1852–1858, 2011.