| 研究生: |
林宜瑩 LIN, I-Ying |
|---|---|
| 論文名稱: |
量尺污染模式對DIF 偵測結果之影響:以DSP與DFTD量尺淨化程序為例 Effects of Scale Contamination on DIF Detection: Evidence from the DSP and DFTD Scale Purification Procedures |
| 指導教授: |
鄭中平
Cheng, Chung-Ping |
| 學位類別: |
碩士 Master |
| 系所名稱: |
社會科學院 - 心理學系 Department of Psychology |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 155 |
| 中文關鍵詞: | 差異試題功能 、量尺污染 、量尺淨化 、DFTD 、DSP |
| 外文關鍵詞: | differential item functioning, scale contamination, scale purification, DFTD, DSP |
| 相關次數: | 點閱:5 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究旨在探討不同量尺污染模式對差異試題功能(Differential Item Functioning, DIF)偵測結果之影響,並比較DIF-Free-Then-DIF(DFTD)與 Dual-Scale Purification(DSP)兩種量尺淨化程序在不同污染情境下之表現。過去DIF研究多假設非目標DIF試題完全符合測量不變性,然而實際測驗中,非目標DIF試題仍可能存在幅度較小之群體參數偏移。此類微小偏移雖未必達DIF 偵測標準,但可能於共同量尺建立過程中累積,進而影響參照基準與DIF 偵測結果。
本研究採用 Rasch 模型進行模擬資料生成,測驗長度固定為 32 題,參照組與焦點組樣本數分別為 2027 與 2008。研究操弄DIF試題數量、模擬情境、群體能力差異、DIF強度、定錨試題數量及量尺淨化方法等因素,並以 MIMIC 模型進行DIF 偵測。研究以檢定力(power)與第一類錯誤率(Type I error rate)作為 DIF 偵測表現之主要評估指標,同時結合圖形型態分析,觀察不同污染情境下試題分布、參照基準及參照差異分布的變化。研究結果顯示,檢定力主要受到 DIF 強度與模擬情境影響。當 DIF 強度增加時,目標 DIF 試題與其他試題之差異較為明顯,因此 DFTD 與 DSP 皆呈現較佳的 DIF 偵測能力。另一方面,第一類錯誤率主要受到模擬情境影響,顯示量尺背景中不同形式的微小偏移可能改變非目標 DIF 試題遭錯誤判定的情形。相較於未加入額外擾動的情境,當微小偏移具有一致方向或呈現較分散的雙向變動時,參照基準與試題相對位置可能受到影響,進而增加 DIF 判定的不穩定性。圖形型態分析結果進一步顯示,不同模擬情境下之參照差異分布可歸納為 P1 量尺穩定型、P2-A 局部偏移型、P2-B 擴散偏移型及 P3 量尺失穩型。整體而言,DFTD與DSP在類型歸納上呈現相近趨勢,但兩者在特定情境下之參照基準與誤判表現仍有所差異。研究結果指出,量尺污染對DIF 偵測之影響不僅表現在偵測比例,也反映於參照差異分布與量尺穩定性之改變。因此,進行DIF 偵測時,除考量DIF試題本身的不公平強度外,亦應重視微小偏移之擾動形式及其對共同量尺建立之影響。
This study examined how different patterns of scale contamination affect differential item functioning (DIF) detection and compared two scale purification procedures, DIF-Free-Then-DIF (DFTD) and Dual-Scale Purification (DSP). Unlike conventional simulation studies that assume non-target DIF items are fully invariant across groups, the present study allowed small between-group parameter shifts among non-target items to represent more complex measurement conditions. Data were generated under the Rasch model for a 32-item test with 2,027 examinees in the reference group and 2,008 in the focal group. Six factors were manipulated: number of DIF items, simulation condition, group ability difference, DIF magnitude, number of anchor items, and scale purification procedure. DIF was detected using a MIMIC model, and performance was evaluated using statistical power and Type I error rate. Graphical pattern analysis was also used to examine changes in item distributions, reference baselines, and anchor difference distributions. Results showed that statistical power was mainly influenced by DIF magnitude and simulation condition, whereas Type I error was mainly influenced by simulation condition. Four graphical patterns were identified: P1 scale-stable, P2-A local-shift, P2-B diffuse-shift, and P3 scale-unstable. Overall, the findings indicate that the direction and distribution of small shifts can affect common-scale construction and the stability of DIF classification.
1. Belzak, W. C. M., & Bauer, D. J. (2020). Improving the assessment of measurement invariance: Using regularization to select anchor items and identify Differential item functioning. Psychological Methods, 25(6), 673–690.
2. Chen, C.-T., & Hwu, B.-S. (2018). Improving the assessment of Differential item functioning in large-scale programs with dual-scale purification of Rasch models: The PISA example. Applied Psychological Measurement, 42(3), 206–220.
3. Chun, S., Stark, S., Kim, E. S., & Chernyshenko, O. S. (2016). MIMIC methods for detecting DIF among multiple groups: Exploring a new sequential-free baseline procedure. Applied Psychological Measurement, 40(7), 486–499.
4. De Boeck, P. (2008). Random item IRT models. Psychometrika, 73(4), 533–559.
5. De Boeck, P., Cho, S.-J., & Wilson, M. (2011). Explanatory secondary dimension modeling of latent Differential item functioning. Applied Psychological Measurement, 35(8), 583–603.
6. De Jong, M. G., Steenkamp, J.-B. E. M., & Fox, J.-P. (2007). Relaxing measurement invariance in cross-national consumer research using a hierarchical IRT model. Journal of Consumer Research, 34(2), 260–278.
7. Finch, W. H. (2016). Detection of Differential item functioning for more than two groups: A Monte Carlo comparison of methods. Applied Measurement in Education, 29(1), 30–45.
8. Hartig, J., Köhler, C., & Naumann, A. (2020). Using a multilevel random item Rasch model to examine item Difficulty variance between random groups. Psychological Test and Assessment Modeling, 62(1), 11–27.
9. Huelmann, T., Debelak, R., & Strobl, C. (2020). A comparison of aggregation rules for selecting anchor items in multi group DIF analysis. Journal of Educational Measurement, 57(2), 185–215.
10. Kopf, J., Zeileis, A., & Strobl, C. (2015a). A framework for anchor methods and an iterative forward approach for DIF detection. Applied Psychological Measurement, 39(2), 83–103.
11. Kopf, J., Zeileis, A., & Strobl, C. (2015b). Anchor selection strategies for DIF analysis: Review, assessment, and new approaches. Educational and Psychological Measurement, 75(1), 22–56.
12. Lek, K., Oberski, D., Davidov, E., Cieciuch, J., Seddig, D., & Schmidt, P. (2018). Approximate measurement invariance. In T. P. Johnson, B.-E. Pennell, I. A. L. Stoop, & B. Dorer (Eds.), Advances in comparative survey methods. Wiley.
13. Magis, D., Tuerlinckx, F., & De Boeck, P. (2015). Detection ofDifferential item functioning using the lasso approach. Journal of Educational and Behavioral Statistics, 40(2), 111–135.
14. Muthén, B. (2017). Recent methods for the study of measurement invariance with many groups: Alignment and random effects. Sociological Methods & Research, 47(4), 637–664.
15. Shealy, R., & Stout, W. (1993). A model-based standardization approach that separates true bias/DIF from group ability differences and detects test bias/DTF as well as item bias/DIF. Psychometrika, 58(2), 159–194.
16. Shih, C.-L., & Wang, W.-C. (2009). Differential item functioning detection using the multiple indicators, multiple causes method with a pure short anchor. Applied Psychological Measurement, 33(3), 184–199.
17. Wang, W.-C., & Shih, C.-L. (2010). MIMIC methods for assessing Differential item functioning in polytomous items. Applied Psychological Measurement, 34(3), 166–180.
18. Wang, W.-C., Shih, C.-L., & Sun, G.-W. (2012). The DIF-free-then-DIF strategy for the assessment of Differential item functioning. Educational and Psychological Measurement, 72(4), 687–708.
19. Wang, W.-C., Shih, C.-L., & Yang, C.-C. (2009). The MIMIC method with scale purification for detecting Differential item functioning. Educational and Psychological Measurement, 69(5), 713–731.