| 研究生: |
王星馳 Wang, Sing-Chih |
|---|---|
| 論文名稱: |
特徵選擇後基於大型語言模型之 MDS-UPDRS 手指敲擊嚴重度評分 LLM-Based MDS-UPDRS Finger Tapping Severity After Feature Selection |
| 指導教授: |
吳馬丁
Nordling, Torbjörn |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 機械工程學系 Department of Mechanical Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 241 |
| 中文關鍵詞: | MDS-UPDRS手指敲擊 、運動遲緩 、大型語言模型 、參數化知識 、閉卷評測 、提示消融 、少樣本範例 、共識特徵選擇 、標記高效臨床評分 |
| 外文關鍵詞: | MDS-UPDRS finger tapping, bradykinesia, large language models, parametric knowledge, closed-book evaluation, prompt ablation, few-shot exemplars, consensus feature selection, token-efficient clinical scoring |
| 相關次數: | 點閱:125 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
研究介紹: 巴金森氏症(PD)是一種進行性神經退化性疾病,也是全球神經失能負擔的主要來源之一;根據 2021 年全球疾病負擔研究(Global Burden of Disease Study 2021),全球約有 1180 萬人受此疾病影響。臨床上以運動障礙協會統一巴金森氏症評定量表(MDS-UPDRS)評定運動嚴重程度,其中手指敲擊(FT)項目透過反覆拇指與食指敲擊的速度、振幅與節奏來評估運動遲緩。然而,這些評分須由受過訓練的神經科醫師人工給定,此過程耗時、難以規模化,且存在評分者間(inter-rater)變異。
研究目標: 本論文探討:在一套臨床對齊的運動學特徵引導下,大型語言模型(LLMs)能否重現專家對 MDS-UPDRS 手指敲擊項目的評分,以及此一結果究竟建立在該流程的哪一個環節之上。
研究方法: 本研究依序處理三個問題。首先確定應提供哪些特徵。本研究招募 16 名巴金森氏症受試者,使用智慧型手機(720p、240 fps)錄製 53 段手指敲擊測試影片,由兩名運動障礙專科神經科醫師獨立評分,並以兩位評分的平均作為參考標準。以 CoTracker 提取拇指與食指軌跡,彙整為含 148 個運動學特徵的特徵池,其中 39 個特徵明確對應 MDS-UPDRS 3.4 條目所評估的臨床建構,如振幅遞減、開合相不對稱、停頓彙總與任務完成失敗;惟其中兩項建構在本世代群體上無法成立——符合評分準則字面定義的中斷次數在 53 段影片中全為零,而遞減起始指標在其中 51 段上無定義。為避免僅信任單一選擇器,本研究以自助重抽樣(bootstrap)下的跨方法選取頻率,彙整 28 種 filter、wrapper 與 embedded 特徵選擇方法,形成 17 特徵共識池,再經詳盡的條件數(conditioning)分析,縮減為一個良置(well-conditioned)的五特徵子集 S5。其次檢驗模型是否真的具備該量表的知識,為此設計一項閉卷(closed-book)知識探查:在提示中不提供任何評分準則的情況下,詢問 11 個開源權重模型(參數量自 7B 至 1.1T,涵蓋五個模型家族)與 5 個前沿模型三個問題——MDS-UPDRS 是什麼、第三部分運動檢查包含哪些條目,以及條目 3.4 自 0 至 4 的確切評分準則;每個開源權重模型各回答五次,全部 165 份回答皆以一套固定的決定性(deterministic)評分準則評分,其中條目 3.4 拆解為九項事實(三個核心維度 × 三個嚴重度層級)。第三則檢驗提示中究竟是哪一個元素在決定分數:以一項十臂消融實驗,在探查證實確實具備該量表知識的前沿模型上,依序移除外部檢索的提供、書面理由的要求、逐字評分準則、臨床框架、特徵名稱、特徵定義,以及四個示範範例;每段影片抽樣五次,並以配對的 session 層級自助重抽樣與置換檢定搭配 Holm-Bonferroni 校正進行分析。
研究結果: 知識探查顯示,條目 3.4 的操作型內容在開源權重模型中幾乎完全缺席:在全部 55 次開源權重模型的作答中,三個中斷次數門檻與三個振幅遞減起始點——亦即區隔各嚴重度層級的關鍵條款——的回收次數為零,且參數量增加約 15 倍亦未帶回其中任何一項;相對地,每一個前沿模型皆回收了九項事實中的八或九項。在 17 特徵共識池上,三個彼此獨立的訓練式讀出器——多層感知器、隨機森林與比例勝算序位模型——在留一交叉驗證下皆與神經科醫師共識一致,最佳者達二次加權 Cohen's κw = 0.86,顯示該一致度來自子集本身而非任何單一估計器。在縮減後的 S5 與四個示範範例之下,前沿模型對兩位醫師平均分數的一致度最高達 κw = 0.743,高於兩位醫師彼此之間的 0.364。在被消融的七個提示元素中,只有示範範例造成的分數變動超過重複執行本身的離散度,也只有該對照通過多重比較校正:即使移除臨床框架、評分準則、特徵名稱與特徵定義,只要保留範例,一致度仍有 κw = 0.644;反之,給足上述所有語意線索但不給範例,一致度掉到 0.371。範例所提供的並非該量表的知識,而是一個族群層級的參考框架——因為特徵是在每段錄影內部正規化的,缺少錨點時模型會把正規化後的範圍當成絕對值,於是把一群病人讀成幾近健康。以預先萃取的特徵取代原始軌跡,將每次查詢的輸入長度自約 24,700 個標記縮減至約 1,900 個——降幅達 92%——這正是十個實驗臂的重複抽樣得以在計算上可行的原因。
研究結論: 以共識為基礎的特徵選擇流程,將含 148 個運動學特徵的特徵池萃取為精簡且具臨床可解釋性的子集,使大型語言模型在少樣本提示下,於 MDS-UPDRS 手指敲擊評分上達到優於臨床醫師彼此的一致度。消融實驗精確定位了提示的貢獻:在提示所包含的一切之中,只有示範範例會移動分數;而被注入的評分準則——一如所有此類流程的設計所預設,本應是關鍵所在——可以整份移除而不產生可偵測的變化。其餘則由特徵承擔:三個訓練式讀出器在完全沒有提示、也沒有語言模型的情況下,於共識池上達到二次加權 κw 最高 0.86。由於知識探查顯示該量表本身並不存在於開源權重模型的權重之中,MDS-UPDRS 評分任務的基礎模型選擇應先經知識探查驗證,方能提出任何評分主張;而在 S5 上的開源權重模型評分讀出,仍是尚待補上的量測。
Parkinson's disease (PD) is a progressive neurodegenerative disorder and a leading contributor to the global burden of neurological disability, affecting roughly 11.8 million people worldwide as of 2021 . Motor severity is graded clinically with the Movement Disorder Society-sponsored revision of the Unified Parkinson's Disease Rating Scale (MDS-UPDRS), whose finger-tapping (FT) item probes bradykinesia through the speed, amplitude, and rhythm of repeated thumb–index tapping. Because these scores are assigned by trained neurologists—a process that is labour-intensive, difficult to scale, and subject to inter-rater variability—this thesis investigates whether large language models (LLMs), guided by a clinically grounded set of kinematic features, can reproduce expert FT scoring, and which part of such a pipeline the result actually depends on.
Three questions are addressed in sequence. Which features to supply is settled first. Sixteen subjects with PD contributed 53 smartphone-recorded FT sessions (720p, 240 fps) that were independently scored by two movement-disorder specialists, with the mean of their two ratings taken as the reference standard. Thumb and index-finger trajectories were extracted with CoTracker and summarised as a 148-feature kinematic pool, of which 39 features are explicitly aligned with MDS-UPDRS 3.4 constructs such as amplitude decrement, opening/closing-phase asymmetry, halt aggregates, and task-completion failure—an alignment that two of those constructs cannot sustain on this cohort, where the rubric-literal interruption count is zero for all 53 recordings and the decrement-onset index is undefined in 51 of them. Rather than trusting any single selector, 28 filter, wrapper, and embedded feature-selection methods were aggregated by their cross-method selection frequency under bootstrap resampling, and the resulting 17-feature consensus pool was reduced by an exhaustive conditioning analysis to a well-conditioned five-feature subset S5. Whether the models hold the scale at all is tested second, by a closed-book knowledge probe: eleven open-weight models (7B–1.1T parameters, spanning five families) and five frontier models were asked what the MDS-UPDRS is, which items constitute the Part III motor examination, and the exact 0–4 criteria of item 3.4, with no rubric in context; each open-weight model answered five times, and all 165 answers were graded by a fixed deterministic rubric that scores item 3.4 as nine facts (three cardinal dimensions × three severity levels). Which part of the prompt carries the score is tested third, by a ten-arm ablation on a frontier model that the probe shows does hold the criteria, removing in turn an offer of external retrieval, the request for a written motivation, the verbatim rubric, the clinical framing, the feature names, the feature definitions, and the four worked exemplars, with five stochastic draws per recording and paired session-level bootstrap and permutation tests under Holm-Bonferroni correction.
The probe found the operational content of item 3.4 to be effectively absent from open-weight models: across all 55 open-weight runs the three interruption-count thresholds and the three amplitude-decrement onsets—the clauses separating one severity level from the next—were recovered zero times, and a ~15-fold increase in parameters added none of them, whereas every frontier model recovered eight or nine of the nine graded facts. On the 17-feature consensus pool, three independent trained readouts—a multilayer perceptron, a random forest, and an ordinal logistic model—agreed with the neurologist consensus under leave-one-out cross-validation, the best reaching quadratic-weighted Cohen's κw = 0.86, so the agreement is a property of the subset rather than of any single estimator. Given the reduced subset S5 and four worked examples, a frontier model reached κw up to 0.743 against the two-rater mean, above the 0.364 at which the two clinicians agreed with each other. Of the seven prompt components ablated, only the worked exemplars changed the score by more than the run-to-run spread, and only that contrast survived multiplicity correction: stripped of the clinical framing, the rubric, the feature names and the feature definitions but given exemplars, the readout still reached κw 0.644, while given every one of those cues and no exemplars it fell to 0.371. What the exemplars supply is a population reference frame rather than knowledge of the scale, since the features are normalised within each recording, so without anchors the model reads the normalised range as absolute and a patient cohort as very nearly healthy. Supplying pre-extracted features instead of raw trajectories shortened the per-query input from ~24,700 tokens to ~1,900 tokens—a 92% reduction—which is what made repeated sampling across ten arms computationally feasible.
In sum, a consensus feature-selection pipeline distils a 148-feature kinematic pool into a compact, clinically interpretable subset on which an LLM reaches better-than-clinician agreement on MDS-UPDRS finger-tapping scoring under few-shot prompting. The ablation locates the prompt's contribution precisely: of everything the prompt contains, only the worked examples move the score, and the injected rubric—which the design of every such pipeline assumes to be doing the work—can be removed without a detectable change. The features carry the rest: three trained readouts reach quadratic-weighted κw up to 0.86 on the consensus pool with no prompt and no language model at all. Because the probe shows the scale itself is absent from open-weight weights, base-model selection for MDS-UPDRS scoring should be validated by a knowledge probe before any scoring claim is made; an open-weight readout on S5 remains the outstanding measurement.
Abdo, W. F., van de Warrenburg, B. P. C., Burn, D. J., Quinn, N. P., and Bloem, B. R. (2010). The clinical approach to movement disorders. Nature Reviews Neurology, 6(1):29–37.
Alzubaidi, S. and Soori, P. K. (2012). Energy efficient lighting system design for hospitals diagnostic and treatment room—a case study. Journal of Light & Visual Environment, 36(1):23–31.
Ashyani, A., Lin, C.-L., Román Catafau, E., Yeh, T., Kuo, T., Tsai, W.-F., Lin, Y., Tu, R., Su, A., Wang, C.-C., Tan, C.-H., and Nordling, T. E. M. (2022). A protocol for digitization of UPDRS upper limb motor examinations towards automated quantification of symptoms of Parkinson's disease. Manuscript in preparation.
Bank, P. J., Marinus, J., Meskers, C. G., de Groot, J. H., and van Hilten, J. J. (2017). Optical hand tracking: A novel technique for the assessment of bradykinesia in parkinson's disease. Movement Disorders Clinical Practice, 4(6):875–883.
Belcak, P., Heinrich, G., Diao, S., Fu, Y., Dong, X., Muralidharan, S., Lin, Y. C., and Molchanov, P. (2025). Small language models are the future of agentic AI. arXiv preprint arXiv:2506.02153.
Belsley, D. A., Kuh, E., and Welsch, R. E. (1980). Regression Diagnostics: Identifying Influential Data and Sources of Collinearity. John Wiley & Sons.
Berardelli, A., Wenning, G. K., Antonini, A., Berg, D., Bloem, B. R., Bonifati, V., Brooks, D., Burn, D. J., Colosimo, C., Fanciulli, A., Ferreira, J., Gasser, T., Grandas, F., Kanovsky, P., Kostic, V., Kulisevsky, J., Oertel, W., Poewe, W., Reese, J.-P., Relja, M., Ruzicka, E., Schrag, A., Seppi, K., Taba, P., and Vidailhet, M. (2013). EFNS/MDS-ES recommendations for the diagnosis of Parkinson's disease. European journal of neurology, 20(1):16–34.
Bologna, M., Leodori, G., Stirpe, P., Paparella, G., Colella, D., Belvisi, D., Fasano, A., Fabbrini, G., and Berardelli, A. (2016). Bradykinesia in early and advanced Parkinson's disease. Journal of the Neurological Sciences, 369:286–291.
Bolón-Canedo, V. and Alonso-Betanzos, A. (2019). Ensembles for feature selection: A review and future trends. Information Fusion, 52:1–12.
Breiman, L. (2001). Random forests. Machine Learning, 45(1):5–32.
Brown, T. B., Mann, B., Ryder, N., and Subbiah, M. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems 33, volume 33, pages 1877–1901.
Buckley, T. A., Crowe, B., Abdulnour, R.-E. E., Rodman, A., and Manrai, A. K. (2025). Comparison of frontier open-source and proprietary large language models for complex diagnoses. JAMA Health Forum, 6(3):e250040.
Burke, H. B., Hoang, A., Lopreiato, J. O., King, H., Hemmer, P., Montgomery, M., and Gagarin, V. (2024). Assessing the ability of a large language model to score free-text medical student clinical notes: Quantitative study. JMIR Medical Education, 10:e56342–e56342.
Butt, A. H., Rovini, E., Dolciotti, C., De Petris, G., Bongioanni, P., Carboncini, M. C., and Cavallo, F. (2018). Objective and automatic classification of Parkinson disease with Leap Motion controller. Biomedical engineering online, 17(1):1–21.
Cao, Z., Simon, T., Wei, S.-E., and Sheikh, Y. (2017). Realtime multi-person 2d pose estimation using part affinity fields. In Proceedings of the IEEE conference on computer vision and pattern, pages 1302–1310. IEEE.
Caruccio, L., Cirillo, S., Polese, G., Solimando, G., Sundaramurthy, S., and Tortora, G. (2024). Can chatgpt provide intelligent diagnoses? a comparative study between predictive models and chatgpt to define a new medical diagnostic bot. Expert Systems with Applications, 235:121186.
Chen, T. and Guestrin, C. (2016). XGBoost: a scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on, pages 785–794. ACM.
Chen, Y.-D. (2025). MDS-UPDRS finger tapping test evaluation comparing ChatGPT-4 and two neurologists. Master's thesis, National Cheng Kung University, Tainan, Taiwan.
Chiu, Y.-H., Liu, T.-K., Yang, C.-P., Chien, C.-F., Liou, L.-M., Lin, L.-C., Dong, H.-P., and Ouyang, C.-S. (2025). An automatic rating approach using machine learning and feature selection for finger tapping in MDS-UPDRS part III. In 2025 19th International Conference on Machine Vision and Applications (MVA), pages 1–4, Kyoto, Japan. IEEE.
Cohen, J. (1968). Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. Psychological bulletin, 70(4):213–220.
de Raadt, A., Warrens, M. J., Bosker, R. J., and Kiers, H. A. L. (2021). A comparison of reliability coefficients for ordinal rating scales. Journal of Classification, 38(3):519–543.
Deng, D., Ostrem, J. L., Nguyen, V., Cummins, D. D., Sun, J., Pathak, A., Little, S., and Abbasi-Asl, R. (2024). Interpretable video-based tracking and quantification of parkinsonism clinical motor states. npj Parkinson's Disease, 10(1):122.
Djurić-Jovičić, M., Petrović, I., Ječmenica-Lukić, M., Radovanović, S., Dragašević-Mišković, N., Belić, M., Miler-Jerković, V., Popović, M. B., and Kostić, V. S. (2016). Finger tapping analysis in patients with parkinson’s disease and atypical parkinsonism. Journal of Clinical Neuroscience, 30:49–55.
Dong, Q., Li, L., Dai, D., Zheng, C., Ma, J., Li, R., Xia, H., Xu, J., Wu, Z., Chang, B., Sun, X., Li, L., and Sui, Z. (2024). A survey on in-context learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1107–1128. Association for Computational Linguistics.
Drucker, H., Burges, C. J. C., Kaufman, L., Smola, A. J., and Vapnik, V. (1997). Support vector regression machines. In Advances in Neural Information Processing Systems (NIPS), volume 9, pages 155–161.
Efron, B., Hastie, T., Johnstone, I., and Tibshirani, R. (2004). Least angle regression. The Annals of Statistics, 32(2):407–451.
Espay, A. J., Giuffrida, J. P., Chen, R., Payne, M., Mazzella, F., Dunn, E., Vaughan, J. E., Duker, A. P., Sahay, A., Kim, S. J., Revilla, F. J., and Heldman, D. A. (2011). Differential response of speed, amplitude, and rhythm to dopaminergic medications in Parkinson's disease. Movement Disorders, 26(14):2504–2508.
Evers, L. J. W., Krijthe, J. H., Meinders, M. J., Bloem, B. R., and Heskes, T. M. (2019). Measuring Parkinson's disease over time: The real-world within-subject reliability of the MDS-UPDRS. Movement Disorders, 34(10):1480–1487.
Fahn, S., Elton, R. L., and Members of the UPDRS Development Committee (1987). Unified Parkinson's Disease Rating Scale. In Fahn, S., Marsden, C. D., Goldstein, M., and Calne, D. B., editors, Recent Developments in Parkinson's Disease, Volume II, pages 153–163. Macmillan Healthcare Information, Florham Park, NJ. vol. 2.
Fisher, A., Rudin, C., and Dominici, F. (2019). All models are wrong, but many are useful: Learning a variable's importance by studying an entire class of prediction models simultaneously. Journal of Machine Learning Research, 20(177):1–81.
Friedman, J. H. (1991). Multivariate adaptive regression splines. The Annals of Statistics, 19(1):1–67.
Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5):1189–1232.
Future AGI (2026). GPT-4 Turbo (2024-04-09) pricing — OpenAI.
Gao, Y., Xiong, Y., Gao, X., and Jia, K. a. (2023). Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 [cs.CL].
Geurts, P., Ernst, D., and Wehenkel, L. (2006). Extremely randomized trees. Machine Learning, 63(1):3–42.
Goetz, C. G., Poewe, W., Rascol, O., Sampaio, C., Stebbins, G. T., Counsell, C., Giladi, N., Holloway, R. G., Moore, C. G., Wenning, G. K., Yahr, M. D., and Seidl, L. (2004). Movement Disorder Society Task Force report on the Hoehn and Yahr staging scale: Status and recommendations. the Movement Disorder Society Task Force on rating scales for Parkinson's disease. Movement Disorders, 19(9):1020–1028.
Goetz, C. G., Stebbins, G. T., Chmura, T. A., Fahn, S., Poewe, W., and Tanner, C. M. (2010). Teaching program for the movement disorder society-sponsored revision of the unified parkinson's disease rating scale: (mds-updrs). Movement disorders, 25(9):1190–1194.
Goetz, C. G., Tilley, B. C., Shaftman, S. R., Stebbins, G. T., Fahn, S., Martinez-Martin, P., Poewe, W., Sampaio, C., Stern, M. B., Dodel, R., Dubois, B., Holloway, R., Jankovic, J., Kulisevsky, J., Lang, A. E., Lees, A., Leurgans, S., LeWitt, P. A., Nyenhuis, D., Olanow, C. W., Rascol, O., Schrag, A., Teresi, J. A., van Hilten, J. J., and LaPelle, N. (2008). Movement Disorder Society-sponsored revision of the Unified Parkinson's Disease Rating Scale (MDS-UPDRS): Scale presentation and clinimetric testing results. Movement Disorders, 23(15):2129–2170.
Gretton, A., Bousquet, O., Smola, A., and Schölkopf, B. (2005). Measuring statistical dependence with Hilbert-Schmidt norms. In Proceedings of the 16th International Conference on Algorithmic, pages 63–77. Springer Berlin Heidelberg.
Gu, J., Jiang, X., Shi, Z., Tan, H., Zhai, X., Xu, C., Li, W., Shen, Y., Ma, S., Liu, H., et al. (2024). A survey on LLM-as-a-Judge. arXiv preprint arXiv:2411.15594.
Gu, Q., Li, Z., and Han, J. (2011). Generalized Fisher score for feature selection. In Proceedings of the Twenty-Seventh Conference on Uncertainty in, pages 266–273.
Guarín, D. L., Wolfe, J. G., Kane, S., Lange, F., and Wong, J. K. (2026). Automated video analysis for early detection of bradykinesia in Parkinson's disease. Journal of NeuroEngineering and Rehabilitation, 23(1):95.
Guarín, D. L., Wong, J. K., McFarland, N. R., and Ramirez-Zamora, A. (2024). Characterizing disease progression in Parkinson's disease from videos of the finger tapping test. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 32:2293–2301.
Guerra-Manzanares, A., Lechuga Lopez, L. J., Maniatakos, M., and Shamout, F. E. (2023). Privacy-preserving machine learning for healthcare: Open challenges and future perspectives. In Trustworthy Machine Learning for Healthcare (TML4H), volume 13932 of Lecture Notes in Computer Science, pages 25–40. Springer.
Guo, D., Yang, D., Zhang, H., Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, R., Ma, S., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., Xue, B., Wang, B., Wu, B., Feng, B., Lu, C., Zhao, C., Deng, C., Ruan, C., Dai, D., Chen, D., Ji, D., Li, E., Lin, F., Dai, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Xu, H., Ding, H., Gao, H., Qu, H., Li, H., Guo, J., Li, J., Chen, J., Yuan, J., Tu, J., Qiu, J., Li, J., Cai, J. L., Ni, J., Liang, J., Chen, J., Dong, K., Hu, K., You, K., Gao, K., Guan, K., Huang, K., Yu, K., Wang, L., Zhang, L., Zhao, L., Wang, L., Zhang, L., Xu, L., Xia, L., Zhang, M., Zhang, M., Tang, M., Zhou, M., Li, M., Wang, M., Li, M., Tian, N., Huang, P., Zhang, P., Wang, Q., Chen, Q., Du, Q., Ge, R., Zhang, R., Pan, R., Wang, R., Chen, R. J., Jin, R. L., Chen, R., Lu, S., Zhou, S., Chen, S., Ye, S., Wang, S., Yu, S., Zhou, S., Pan, S., Li, S. S., Zhou, S., Wu, S., Yun, T., Pei, T., Sun, T., Wang, T., Zeng, W., Liu, W., Liang, W., Gao, W., Yu, W., Zhang, W., Xiao, W. L., An, W., Liu, X., Wang, X., Chen, X., Nie, X., Cheng, X., Liu, X., Xie, X., Liu, X., Yang, X., Li, X., Su, X., Lin, X., Li, X. Q., Jin, X., Shen, X., Chen, X., Sun, X., Wang, X., Song, X., Zhou, X., Wang, X., Shan, X., Li, Y. K., Wang, Y. Q., Wei, Y. X., Zhang, Y., Xu, Y., Li, Y., Zhao, Y., Sun, Y., Wang, Y., Yu, Y., Zhang, Y., Shi, Y., Xiong, Y., He, Y., Piao, Y., Wang, Y., Tan, Y., Ma, Y., Liu, Y., Guo, Y., Ou, Y., Wang, Y., Gong, Y., Zou, Y., He, Y., Xiong, Y., Luo, Y., You, Y., Liu, Y., Zhou, Y., Zhu, Y. X., Huang, Y., Li, Y., Zheng, Y., Zhu, Y., Ma, Y., Tang, Y., Zha, Y., Yan, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Xu, Z., Xie, Z., Zhang, Z., Hao, Z., Ma, Z., Yan, Z., Wu, Z., Gu, Z., Zhu, Z., Liu, Z., Li, Z., Xie, Z., Song, Z., Pan, Z., Huang, Z., Xu, Z., Zhang, Z., and Zhang, Z. (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature, 645(8081):633–638.
Guyon, I. and Elisseeff, A. (2003). An introduction to variable and feature selection. Journal of Machine Learning Research, 3:1157–1182.
Guyon, I., Weston, J., Barnhill, S., and Vapnik, V. (2002). Gene selection for cancer classification using support vector machines. Machine Learning, 46(1-3):389–422.
Heye, K., Li, R., Bai, Q., St George, R. J., Rudd, K., Huang, G., Meinders, M. J., Bloem, B. R., and Alty, J. E. (2024). Validation of computer vision technology for analyzing bradykinesia in outpatient clinic videos of people with parkinson's disease. Journal of the Neurological Sciences, 466:123271.
Hoerl, A. E. and Kennard, R. W. (1970). Ridge regression: Biased estimation for nonorthogonal problems. Technometrics, 12(1):55–67.
Hsieh, C.-Y., Li, C.-L., Yeh, C.-K., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C.-Y., and Pfister, T. (2023). Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes. In Findings of the Association for Computational Linguistics: ACL 2023, pages 8003–8017. Association for Computational Linguistics.
Hsu, Y.-C., Su, Y.-H., Cheng, B.-R., Sung, S.-F., Liu, J.-X., Hsu, H.-C., and Hsiung, P.-A. (2024). Movement disorder evaluation of parkinson's disease severity based on deep neural network models. IEEE Access, 12:143413–143433.
Islam, M. S., Rahman, W., Abdelkader, A., Lee, S., Yang, P. T., Purks, J. L., Adams, J. L., Schneider, R. B., Dorsey, E. R., and Hoque, E. (2023). Using AI to measure Parkinson's disease severity at home. npj Digital Medicine, 6(1):156.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Le Scao, T., Lavril, T., Wang, T., Lacroix, T., and El Sayed, W. (2023). Mistral 7B. arXiv preprint arXiv:2310.06825 [cs.CL].
Kaliosis, P., Ganesan, A. V., Kjell, O. N. E., Ringwald, W., Feltman, S., Carr, M. A., Samaras, D., Ruggero, C., Luft, B. J., Kotov, R., and Schwartz, A. H. (2026). A systematic evaluation of large language models for PTSD severity estimation: The role of contextual knowledge and modeling strategies.
Kalousis, A., Prados, J., and Hilario, M. (2007). Stability of feature selection algorithms: a study on high-dimensional spaces. Knowledge and Information Systems, 12(1):95–116.
Karaev, N., Rocco, I., Graham, B., Neverova, N., Vedaldi, A., and Rupprecht, C. (2024). CoTracker: It is better to track together. In Computer Vision -- ECCV 2024, volume 15120 of Lecture Notes in Computer Science, pages 18–35, Cham. Springer Nature Switzerland.
Kendall, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1/2):81–93.
Kenny, L., Azizi, Z., Moore, K., Alcock, M., Heywood, S., Jonsson, A., McGrath, K., Foley, M. J., Sweeney, B., O’Sullivan, S., Barton, J., Tedesco, S., Sica, M., Crowe, C., and Timmons, S. (2024). Inter-rater reliability of hand motor function assessment in Parkinson’s disease: Impact of clinician training. Clinical Parkinsonism & Related Disorders, 11:100278.
Kira, K. and Rendell, L. A. (1992). The feature selection problem: traditional methods and a new algorithm. In Proceedings of the Tenth National Conference on Artificial Intelligence (AAAI-92), pages 129–134. AAAI Press.
Kononenko, I. (1994). Estimating attributes: analysis and extensions of RELIEF. In Machine Learning: ECML-94, volume 784 of Lecture Notes in Computer Science, pages 171–182. Springer.
Kuncheva, L. I. (2007). A stability index for feature selection. In Proceedings of the 25th IASTED International Multi-Conference: Artificial Intelligence and Applications (AIAP'07), pages 390–395, Innsbruck, Austria. ACTA Press.
Kursa, M. B. and Rudnicki, W. R. (2010). Feature selection with the Boruta package. Journal of Statistical Software, 36(11):1–13.
Landis, J. R. and Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1):159.
Li, H., Leung, J., and Shen, Z. (2024). Towards goal-oriented prompt engineering for large language models: A survey.
Li, H., Shao, X., Zhang, C., and Qian, X. (2021). Automated assessment of Parkinsonian finger-tapping tests through a vision-based fine-grained classification model. Neurocomputing, 441:260–271.
Li, M., Ye, X., Huang, Z., Ye, L., and Chen, C. (2025). Global burden of Parkinson's disease from 1990 to 2021: a population-based study. BMJ Open, 15(4):e095610.
Li, Z., Lu, K., Cai, M., Liu, X., Wang, Y., and Yang, J. (2022). An automatic evaluation method for parkinson's dyskinesia using finger tapping video for small samples. Journal of Medical and Biological Engineering, 42(3):351–363.
Ling, H., Massey, L. A., Lees, A. J., Brown, P., and Day, B. L. (2012). Hypokinesia without decrement distinguishes progressive supranuclear palsy from Parkinson's disease. Brain, 135(4):1141–1153.
Liu, Y., Chen, J., Hu, C., Ma, Y., Ge, D., Miao, S., Xue, Y., and Li, L. (2019). Vision-based method for automatic quantification of parkinsonian bradykinesia. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 27(10):1952–1961.
Liu, Y., Yang, Z., Cai, M., Wang, Y., Liu, X., Tong, H., Peng, Y., Lou, Y., and Li, Z. (2024). ATST-Net: A method to identify early symptoms in the upper and lower extremities of PD. Medical Engineering & Physics, 128(1):104171.
Lu, M., Poston, K., Pfefferbaum, A., Sullivan, E. V., Fei-Fei, L., Pohl, K. M., Niebles, J. C., and Adeli, E. (2020). Vision-based estimation of MDS-UPDRS gait scores for assessing Parkinson's disease motor severity. In Medical Image Computing and Computer Assisted Intervention (MICCAI), pages 637–647.
Lugaresi, C., Tang, J., Nash, H., McClanahan, C., et al. (2019). MediaPipe: A framework for building perception pipelines. arXiv preprint arXiv:1906.08172 [cs.CV].
Ma, Y., Tang, W., Feng, C., and Tu, X. M. (2008). Inference for kappas for longitudinal study data: Applications to sexual health research. Biometrics, 64(3):781–789.
Marill, T. and Green, D. M. (1963). On the effectiveness of receptors in recognition systems. IEEE Transactions on Information Theory, 9(1):11–17.
Marsili, L., Abanto, J., Mahajan, A., Duque, K. R., Chinchihualpa Paredes, N. O., Deraz, H. A., Espay, A. J., and Bologna, M. (2024). Dysrhythmia as a prominent feature of Parkinson's disease: An app-based tapping test. Journal of the Neurological Sciences, 463:123144.
Martinez-Manzanera, O., Roosma, E., Beudel, M., Borgemeester, R. W. K., van Laar, T., and Maurits, N. M. (2016). A method for automatic and objective scoring of bradykinesia using orientation sensors and classification algorithms. IEEE Transactions on Biomedical Engineering, 63(5):1016–1024.
Mehta, D., Asif, U., Hao, T., Bilal, E., Von Cavallar, S., Harrer, S., and Rogers, J. (2021). Towards automated and marker-less Parkinson disease assessment: Predicting UPDRS scores using sit-stand videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW).
Meinshausen, N. and Bühlmann, P. (2010). Stability selection. Journal of the Royal Statistical Society: Series B, 72(4):417–473.
Minaee, S., Mikolov, T., Nikzad, N., Chenaghlu, M., Socher, R., Amatriain, X., and Gao, J. (2024). Large language models: A survey. arXiv preprint arXiv:2402.06196 [cs.CL].
Muennighoff, N., Yang, Z., Shi, W., Li, X. L., Fei-Fei, L., Hajishirzi, H., Zettlemoyer, L., Liang, P., Candès, E., and Hashimoto, T. (2025). s1: Simple test-time scaling. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 20275–20321. Association for Computational Linguistics.
Mukherjee, S., Mitra, A., Jawahar, G., Agarwal, S., Palangi, H., and Awadallah, A. (2023). Orca: Progressive learning from complex explanation traces of GPT-4. arXiv preprint arXiv:2306.02707.
Nogueira, S., Sechidis, K., and Brown, G. (2018). On the stability of feature selection algorithms. Journal of Machine Learning Research, 18(174):1–54.
Nori, H., Lee, Y. T., Zhang, S., and Carignan, D. (2023). Can generalist foundation models outcompete special-purpose tuning? arXiv preprint arXiv:2311.16452 [cs.LG].
Omberg, L., Chaibub Neto, E., Perumal, T. M., Pratap, A., Tediarjo, A., Adams, J., Bloem, B. R., Bot, B. M., Elson, M., Goldman, S. M., et al. (2022). Remote smartphone monitoring of Parkinson's disease and individual response to therapy. Nature Biotechnology, 40(4):480–487.
OpenAI (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774 [cs.CL].
Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., McGuinness, L. A., Stewart, L. A., Thomas, J., Tricco, A. C., Welch, V. A., Whiting, P., and Moher, D. (2021). The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ, 372:n71.
Parker, R. I., Vannest, K. J., and Davis, J. L. (2013). Reliability of multi-category rating scales. Journal of School Psychology, 51(2):217–229.
Patil, A., Tao, S., and Gedhu, A. (2025). Evaluating reasoning LLMs for suicide screening with the Columbia-Suicide Severity Rating Scale. In 2025 IEEE International Conference on Future Machine Learning and Data Science (FMLDS), pages 245–254. IEEE.
Pearson, K. (1895). VII. Note on regression and inheritance in the case of two parents. Proceedings of the Royal Society of London, 58(347-352):240–242.
Pedregosa, F., Varoquaux, G., and Gramfort, A. (2011). Scikit-learn: Machine learning in Python. Journal of machine learning research, 12(October):2825–2830.
Peng, H., Long, F., and Ding, C. (2005). Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy. IEEE transactions on pattern analysis and machine intelligence, 27(8):1226–1238.
Post, B., Merkus, M. P., de Bie, R. M., de Haan, R. J., and Speelman, J. D. (2005). Unified parkinson's disease rating scale motor examination: Are ratings of nurses, residents in neurology, and movement disorders specialists interchangeable? Movement Disorders, 20(12):1577–1584.
Postuma, R. B., Berg, D., Stern, M., Poewe, W., Olanow, C. W., Oertel, W., Obeso, J., Marek, K., Litvan, I., Lang, A. E., Halliday, G., Goetz, C. G., Gasser, T., Dubois, B., Chan, P., Bloem, B. R., Adler, C. H., and Deuschl, G. (2015). Mds clinical diagnostic criteria for parkinson's disease. Movement Disorders, 30(12):1591–1601.
Qin, Z., Wu, J., Shen, J., Liu, T., and Wang, X. (2024). LAMPO: Large language models as preference machines for few-shot ordinal classification.
Robnik-Šikonja, M. and Kononenko, I. (2003). Theoretical and empirical analysis of ReliefF and RReliefF. Machine Learning, 53(1-2):23–69.
Saeys, Y., Abeel, T., and Van de Peer, Y. (2008). Robust feature selection using ensemble feature selection techniques. In Machine Learning and Knowledge Discovery in Databases (ECML PKDD 2008), volume 5212 of Lecture Notes in Computer Science, pages 313–325. Springer Berlin Heidelberg.
Saeys, Y., Inza, I. n., and Larra naga, P. (2007). A review of feature selection techniques in bioinformatics. Bioinformatics, 23(19):2507–2517.
Sarker, I. H. (2022). Ai-based modeling: Techniques, applications and research issues towards automation, intelligent and smart systems. SN Computer Science, 3(2):158.
Schapira, A. H., Chaudhuri, K. R., and Jenner, P. (2017). Non-motor features of parkinson disease. Nature Reviews Neuroscience, 18(7):435–450.
Schober, P., Boer, C., and Schwarte, L. A. (2018). Correlation coefficients: appropriate use and interpretation. Anesthesia & Analgesia, 126(5):1763–1768.
Sengupta, A., Jin, F., Zhang, R., and Cao, S. (2020). mm-pose: Real-time human skeletal posture estimation using mmwave radars and cnns. IEEE Sensors Journal, 20(17):10032–10044.
Shah, R. D. and Samworth, R. J. (2013). Variable selection with error control: another look at stability selection. Journal of the Royal Statistical Society Series B: Statistical Methodology, 75(1):55–80.
Shin, J., Matsumoto, M., Maniruzzaman, M., Hasan, M. A. M., Hirooka, K., Hagihara, Y., Kotsuki, N., Inomata-Terada, S., Terao, Y., and Kobayashi, S. (2024). Classification of hand-movement disabilities in parkinson's disease using a motion-capture device and machine learning. IEEE Access, 12:52466–52479.
Shrout, P. E. and Fleiss, J. L. (1979). Intraclass correlations: uses in assessing rater reliability. Psychological bulletin, 86(2):420–428.
Sibley, K., Girges, C., Candelario, J., Milabo, C., Salazar, M., Esperida, J. O., Dushin, Y., Limousin, P., and Foltynie, T. (2022). An evaluation of KELVIN, an artificial intelligence platform, as an objective assessment of the MDS-UPDRS part III. Journal of Parkinson's Disease, 12(7):2223–2233.
Singh, M., Prakash, P., Kaur, R., Sowers, R., Brašić, J. R., and Hernandez, M. E. (2023). A deep learning approach for automatic and objective grading of the motor impairment severity in parkinson's disease for use in tele-assessments. Sensors, 23(21):9004.
Spearman, C. (1904). The proof and measurement of association between two things. The American Journal of Psychology, 15(1):72–101.
Steinmetz, J. D., Seeher, K. M., Schiess, N., et al. (2024). Global, regional, and national burden of disorders affecting the nervous system, 1990–2021: a systematic analysis for the Global Burden of Disease Study 2021. The Lancet Neurology, 23(4):344–381.
Székely, G. J., Rizzo, M. L., and Bakirov, N. K. (2007). Measuring and testing dependence by correlation of distances. The Annals of Statistics, 35(6):2769–2794.
Tibshirani, R. (1996). Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society Series B: Statistical Methodology, 58(1):267–288.
Toloşi, L. and Lengauer, T. (2011). Classification with correlated features: unreliability of feature ranking and solutions. Bioinformatics, 27(14):1986–1994.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. (2023). Llama: Open and efficient foundation language models.
Tsanas, A., Little, M. A., McSharry, P. E., Spielman, J., and Ramig, L. O. (2012). Novel speech signal processing algorithms for high-accuracy classification of Parkinson's disease. IEEE Transactions on Biomedical Engineering, 59(5):1264–1271.
Vanbelle, S. (2016). A new interpretation of the weighted kappa coefficients. Psychometrika, 81(2):399–410.
Vignoud, G., Desjardins, C., Salardaine, Q., Mongin, M., Garcin, B., Venance, L., and Degos, B. (2022). Video-based automated assessment of movement parameters consistent with MDS-UPDRS III in Parkinson's disease. Journal of Parkinson's Disease, 12(7):2211–2222.
Villani, C. (2009). The Wasserstein distances. In Optimal Transport: Old and New, volume 338 of Grundlehren der mathematischen Wissenschaften, pages 93–111. Springer.
Wang, S. C. and Nordling, T. E. M. (2026a). Consensus feature selection across filter, embedded, and wrapper families for MDS-UPDRS finger tapping. In Rojas, I., Ortu no, F., Rojas Ruiz, F., Herrera, L. J., Valenzuela, O., and Escobar, J. J., editors, Bioinformatics and Biomedical Engineering (IWBBIO 2026), Lecture Notes in Bioinformatics. Springer. Forthcoming.
Wang, S. C. and Nordling, T. E. M. (2026b). Do large language models know the MDS-UPDRS? a closed-book probe of parametric knowledge as a precondition for automated finger-tapping scoring. Manuscript in preparation.
Wang, S. C. and Nordling, T. E. M. (2026c). Worked examples, not injected knowledge: A ten-arm paired ablation of prompt content in LLM MDS-UPDRS finger-tapping scoring. Manuscript in preparation.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., and Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems 35, pages 24824–24837. Neural Information Processing Systems Foundation, Inc. (NeurIPS).
Whitney, A. W. (1971). A direct method of nonparametric measurement selection. IEEE Transactions on Computers, C-20(9):1100–1103.
Williams, S., Wong, D., Alty, J. E., and Relton, S. D. (2023). Parkinsonian hand or clinician’s eye? finger tap bradykinesia interrater reliability for 21 movement disorder experts. Journal of Parkinson's Disease, 13(4):525–536.
Xian, L., Ni, J., and Wang, M. (2025). Leveraging large language models for cost-effective, multilingual depression detection and severity assessment.
Yang, A., Li, A., Yang, B., et al. (2025). Qwen3 technical report.
Yang, A., Yang, B., Zhang, B., et al. (2024a). Qwen2.5 technical report.
Yang, J., Williams, S., Hogg, D. C., Alty, J. E., and Relton, S. D. (2024b). Deep learning of Parkinson's movement from video, without human-defined measures. Journal of the Neurological Sciences, 463:123089.
Yang, N., Liu, D.-F., Liu, T., Han, T., Zhang, P., Xu, X., Lou, S., Liu, H.-G., Yang, A.-C., Dong, C., Vai, M. I., Pun, S. H., and Zhang, J.-G. (2022). Automatic detection pipeline for accessing the motor severity of parkinson’s disease in finger tapping and postural stability. IEEE ACCESS, 10:66961–66973.
Yang, Y.-Y., Ho, M.-Y., Tai, C.-H., Wu, R.-M., Kuo, M.-C., and Tseng, Y. J. (2024c). FastEval Parkinsonism: an instant deep learning-assisted video-based online system for parkinsonian motor symptom evaluation. npj Digital Medicine, 7(1):31.
Yu, L. and Liu, H. (2003). Feature selection for high-dimensional data: A fast correlation-based filter solution. In Proceedings of the Twentieth International Conference on Machine Learning (ICML-03), pages 856–863. AAAI Press.
Yu, T., Park, K. W., McKeown, M. J., and Wang, Z. J. (2023). Clinically informed automated assessment of finger tapping videos in Parkinson's disease. Sensors, 23(22):9149.
Zarrat Ehsan, T., Tangermann, M., Güçlütürk, Y., Shin, S., Ho, K. C., Bloem, B. R., and Evers, L. J. W. (2026). Interpretable and granular video-based quantification of motor characteristics from the finger-tapping test in Parkinson's disease. npj Parkinson's Disease, 12(1):101.
Zhang, E., Goto, R., Sagan, N., Mutter, J., Phillips, N., Alizadeh, A., Lee, K., Blanchet, J., Pilanci, M., and Tibshirani, R. (2025). LLM-Lasso: A robust framework for domain-informed feature selection and regularization. arXiv:2502.10648v2.
Zhao, W. X., Zhou, K., Li, J., and Tang, T. a. (2023). A survey of large language models. arXiv preprint arXiv:2303.18223 [cs.CL].
Zhou, H., Liu, F., Gu, B., and Zou, X. a. (2023). A survey of large language models in medicine: Progress, application, and challenge. arXiv preprint arXiv:2311.05112 [cs.CL].
Zou, H. (2006). The adaptive lasso and its oracle properties. Journal of the American Statistical Association, 101(476):1418–1429.
Zou, H. and Hastie, T. (2005). Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society Series B: Statistical Methodology, 67(2):301–320.