<?xml version="1.0" encoding="ISO-8859-1"?><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
<front>
<journal-meta>
<journal-id>0212-9728</journal-id>
<journal-title><![CDATA[Anales de Psicología]]></journal-title>
<abbrev-journal-title><![CDATA[Anal. Psicol.]]></abbrev-journal-title>
<issn>0212-9728</issn>
<publisher>
<publisher-name><![CDATA[Universidad de Murcia]]></publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id>S0212-97282017000300010</article-id>
<article-id pub-id-type="doi">10.6018/analesps.33.3.238621</article-id>
<title-group>
<article-title xml:lang="en"><![CDATA[An investigation of enhancement of ability evaluation by using a nested logit model for multiple-choice items]]></article-title>
<article-title xml:lang="es"><![CDATA[Una investigación de la mejora de la capacidad de evaluación mediante el uso de un modelo logit anidado para items de elección múltiple]]></article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname><![CDATA[Tour]]></surname>
<given-names><![CDATA[Liu]]></given-names>
</name>
<xref ref-type="aff" rid="A01"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname><![CDATA[Mengcheng]]></surname>
<given-names><![CDATA[Wang]]></given-names>
</name>
<xref ref-type="aff" rid="A02"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname><![CDATA[Tao]]></surname>
<given-names><![CDATA[Xin]]></given-names>
</name>
<xref ref-type="aff" rid="A03"/>
</contrib>
</contrib-group>
<aff id="A01">
<institution><![CDATA[,Tianjin Normal University School of Education Science ]]></institution>
<addr-line><![CDATA[Tianjin ]]></addr-line>
<country>China</country>
</aff>
<aff id="A02">
<institution><![CDATA[,Guangzhou University Center for Psychometric and Latent Variable Modeling ]]></institution>
<addr-line><![CDATA[Guangzhou ]]></addr-line>
<country>China</country>
</aff>
<aff id="A03">
<institution><![CDATA[,Bejing Normal University Collaborative Innovation Center of Assessment toward Basic Education Quality (CICA-BEQ) ]]></institution>
<addr-line><![CDATA[Beijing ]]></addr-line>
<country>China</country>
</aff>
<pub-date pub-type="pub">
<day>00</day>
<month>10</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="epub">
<day>00</day>
<month>10</month>
<year>2017</year>
</pub-date>
<volume>33</volume>
<numero>3</numero>
<fpage>530</fpage>
<lpage>537</lpage>
<copyright-statement/>
<copyright-year/>
<self-uri xlink:href="http://scielo.isciii.es/scielo.php?script=sci_arttext&amp;pid=S0212-97282017000300010&amp;lng=en&amp;nrm=iso"></self-uri><self-uri xlink:href="http://scielo.isciii.es/scielo.php?script=sci_abstract&amp;pid=S0212-97282017000300010&amp;lng=en&amp;nrm=iso"></self-uri><self-uri xlink:href="http://scielo.isciii.es/scielo.php?script=sci_pdf&amp;pid=S0212-97282017000300010&amp;lng=en&amp;nrm=iso"></self-uri><abstract abstract-type="short" xml:lang="en"><p><![CDATA[Multiple-choice items are wildly used in psychological and educational test. The present study investigated that if a multiple-choice item have an advantage over a dichotomous item on ability or latent trait evaluation. An item response model, 2-parameter logistic nested logit model (2PL-NLM), was used to fit the multiple-choice data. Both simulation study and empirical study indicated that the accuracy and the stability of ability estimation were enhanced by using multiple-choice model rather than dichotomous model, because more information was included in multiple-choice items' distractors. But the accuracy of ability estimation showed little differences in four-choice items, five-choice items and six-choice items. Moreover, 2PL-NLM could extract more information from low-level respondents than from high-level ones, because they had more distractor chosen behaviors. In the empirical study, respondents at different trait levels would be attracted by different distractors from the Chinese Vocabulary Test for Grade 1 by using the changing traces of distractor probabilities calculated from 2PL-NLM. It is suggested that the responses of students at different levels could reflect the students' vocabulary development process.]]></p></abstract>
<abstract abstract-type="short" xml:lang="es"><p><![CDATA[Los items de elección múltiple se han usado ampliamente en tests psicológicos y educativos. Este estudio investiga si los items de elección múltiple tiene ventajas sobre los items dicotómicos o sobre la evaluación de rasgo latente. Un modelo de respuesta al item, con un modelo logit anidado, logístico 2-parámetros (2PL-NKM), fue usado para ajustar los datos de elección múltiple. Los estudios de simulación y empíricos indicaron que la precisión y la estabilidad de la estimación de capacidad mejoró usando el modelo de elección múltiple en contraposición al modelo dicotómico, debido a la mayor información incluida en los items distractores de la elección múltiple. Pero la precisión y la capacidad de estimación mostró pequeñas diferencias en items de cuatro elecciones, cinco y seis elecciones. Además, el modelo 2PL-NLM puede extraer más información respondientes de bajo nivel que de los de alto nivel, debido a que tienen conductas de elección con más distractores. En el estudio empírico, los respondientes en diferentes niveles de rasgo fueron atraídos por diferentes distractrores del Test de Vocabulario chino en el primer grado, usando trazos cambiantes en la probabilidad de distractor a partir de 2PL-NLM. Esto sugiere que las respuestas de los estudiantes a diferentes niveles puede reflejar un proceso evolutivo de vocabulario en los estudiantes.]]></p></abstract>
<kwd-group>
<kwd lng="en"><![CDATA[multiple-choice item]]></kwd>
<kwd lng="en"><![CDATA[nested logit model]]></kwd>
<kwd lng="en"><![CDATA[distractor information]]></kwd>
<kwd lng="en"><![CDATA[ability evaluation]]></kwd>
<kwd lng="es"><![CDATA[items de elección múltiple]]></kwd>
<kwd lng="es"><![CDATA[modelo logit anidado]]></kwd>
<kwd lng="es"><![CDATA[información distractora]]></kwd>
<kwd lng="es"><![CDATA[capacidad de evaluación]]></kwd>
</kwd-group>
</article-meta>
</front><body><![CDATA[ <p>&nbsp;</p>     <p>&nbsp;</p> <a name="top"></a>    <p><font face="Verdana" size="4"><b>An investigation of enhancement of ability evaluation by using a nested logit model for multiple-choice items</b></font></p>     <p><font face="Verdana" size="4"><b>Una investigación de la mejora de la capacidad de evaluación mediante el uso de un modelo logit anidado para items de elección múltiple</b></font></p>     <p>&nbsp;</p>     <p>&nbsp;</p>     <p><font face="Verdana" size="2"><b>Liu Tour<sup>1</sup>, Wang Mengcheng<sup>2</sup> and Xin Tao<sup>3</sup></b></font></p>     <p><font face="Verdana" size="2"><sup>1</sup> School of Education Science, Tianjin Normal University, Tianjin (China).    <br><sup>2</sup> Center for Psychometric and Latent Variable Modeling, Guangzhou University, Guangzhou (China).    <br><sup>3</sup> Collaborative Innovation Center of Assessment toward Basic Education Quality (CICA-BEQ), Bejing Normal University, Beijing (China).</font></p>     ]]></body>
<body><![CDATA[<p><font face="Verdana" size="2"><a href="#bajo">Correspondence</a></font></p>     <p>&nbsp;</p>     <p>&nbsp;</p> <hr size="1">     <p><font face="Verdana" size="2"><b>ABSTRACT</b></font></p>     <p><font face="Verdana" size="2">Multiple-choice items are wildly used in psychological and educational test. The present study investigated that if a multiple-choice item have an advantage over a dichotomous item on ability or latent trait evaluation. An item response model, 2-parameter logistic nested logit model (2PL-NLM), was used to fit the multiple-choice data. Both simulation study and empirical study indicated that the accuracy and the stability of ability estimation were enhanced by using multiple-choice model rather than dichotomous model, because more information was included in multiple-choice items' distractors. But the accuracy of ability estimation showed little differences in four-choice items, five-choice items and six-choice items. Moreover, 2PL-NLM could extract more information from low-level respondents than from high-level ones, because they had more distractor chosen behaviors. In the empirical study, respondents at different trait levels would be attracted by different distractors from the Chinese Vocabulary Test for Grade 1 by using the changing traces of distractor probabilities calculated from 2PL-NLM. It is suggested that the responses of students at different levels could reflect the students' vocabulary development process.</font></p>     <p><font face="Verdana" size="2"><b>Key words:</b> multiple-choice item; nested logit model; distractor information; ability evaluation.</font></p> <hr size="1">     <p><font face="Verdana" size="2"><b>RESUMEN</b></font></p>     <p><font face="Verdana" size="2">Los items de elección múltiple se han usado ampliamente en tests psicológicos y educativos. Este estudio investiga si los items de elección múltiple tiene ventajas sobre los items dicotómicos o sobre la evaluación de rasgo latente. Un modelo de respuesta al item, con un modelo logit anidado, logístico 2-parámetros (2PL-NKM), fue usado para ajustar los datos de elección múltiple. Los estudios de simulación y empíricos indicaron que la precisión y la estabilidad de la estimación de capacidad mejoró usando el modelo de elección múltiple en contraposición al modelo dicotómico, debido a la mayor información incluida en los items distractores de la elección múltiple. Pero la precisión y la capacidad de estimación mostró pequeñas diferencias en items de cuatro elecciones, cinco y seis elecciones. Además, el modelo 2PL-NLM puede extraer más información respondientes de bajo nivel que de los de alto nivel, debido a que tienen conductas de elección con más distractores. En el estudio empírico, los respondientes en diferentes niveles de rasgo fueron atraídos por diferentes distractrores del Test de Vocabulario chino en el primer grado, usando trazos cambiantes en la probabilidad de distractor a partir de 2PL-NLM. Esto sugiere que las respuestas de los estudiantes a diferentes niveles puede reflejar un proceso evolutivo de vocabulario en los estudiantes.</font></p>     <p><font face="Verdana" size="2"><b>Palabras clave:</b> items de elección múltiple; modelo logit anidado; información distractora; capacidad de evaluación.</font></p> <hr size="1">     <p>&nbsp;</p>     ]]></body>
<body><![CDATA[<p><font face="Verdana" size="2"><b>Introduction</b></font></p>     <p><font face="Verdana" size="2">Multiple-choice items are widely used in cognitive tests, aptitude tests, educational tests and some intelligence tests, since Frederick J. Kelly's introduction in 1914. A multiple-choice item usually includes one correct option and several incorrect options (distractors). The correct option makes a multiple-choice item have an objective scoring standard as well as a dichotomous item, rather than an open-ended item or an essay-type item. While the several incorrect options may distract respondents and decrease the guessing behaviors of respondents. Multiple-choice item may have some advantages over dichotomous item, but writing a multiple-choice item is more complicated than writing a dichotomous item. Also it will increase the cognitive loads of respondents. Consequently, if multiple-choice items can not make estimation of ability (latent trait) more accurate, it will be economical to use dichotomous items. Or if distractor can not provide more information about respondents, it will be better to credit multiple-choice item as binary item.</font></p>     <p><font face="Verdana" size="2"><b>Distractor Information</b></font></p>     <p><font face="Verdana" size="2">Including several distractors is a distinctive feature of a multiple-choice item, compared with a dichotomous item. Various approaches have been proposed by many researchers to extract information from distractors of multiple-choice items. For one thing, some strategies of multiple-choice item construction were taken into consideration (Haladyna &amp; Downing, 1989; Haladyna, Downing, &amp; Rodriguez, 2002; Tamir, 1971, 1989). Briggs, Alonzo, Schwab, and Wilson (2006) suggested constructing multiple-choice items with ordered options which could seek to more diagnostic information. Liu, Lee, and Linn (2011) believed a multiple-choice item needed some explanatory components as a new tier, following a typical multiple-choice item. In addition, the optimal number of item options had been discussed by Haladyna and Dowing (1993). For another thing, technical treatments were good ways to mine the potential information from distractors without constructing a new test. To assign different weights to options was one way (Davis &amp; Fifer, 1959). To create an augmented data matrix transformed from a raw response matrix by using some special scoring rules was another (Luecht, 2007). And also some indices, such as distractor selection ratio and point-biserial correlation, based on certain statistical models were useful (Attali &amp; Fraenkel, 2000; Love, 1997).</font></p>     <p><font face="Verdana" size="2">The information extracted from distractor can be used as auxiliary information in many ways. In person-fit area, Wollack (1997), Drasgow, Levine, and Williams (1985) were used distractor information to detect aberrant response behaviors. Kim (2006) found that linking procedure was improved when distractor information was added in. Roediger and Marsh (2005) thought that distractors could bring in some psychological consequences. Besides, the differential functioning relative to distractors, namely differential distractor functioning (DDF), could also provide some explanation for detecting measurement bias (Green, Crone, &amp; Folk, 1989; Penfield, 2011; Suh &amp; Bolt, 2011; Suh &amp; Talley, 2015).</font></p>     <p><font face="Verdana" size="2"><b>Multiple-Choice Item Modeling</b></font></p>     <p><font face="Verdana" size="2">To achieve accurate ability evaluation is an ultimate goal of all educational and psychological tests. However, test score is not a good indicator to support that distractor information can enhance ability evaluation. Sigel (1963) found no relationship between error patterns and respondents' scores. Jacob and Vandeventer 's (1970) study showed that types of error and total score were related. But they still did not found the distractor information could improve the ability evaluation. Along with the development of item response theory (IRT), information of respondents' score is accurate to item-level in contrast with test-level in classical test theory (CTT). IRT models are good ways to figure out the relationship between the distractor chosen behaviors and respondents' abilities (Levine &amp; Drasgow, 1983; Thissen, 1976).</font></p>     <p><font face="Verdana" size="2">Generally, there are two kinds of IRT models can fit multiple-choice data. The multiple-choice data, in most instance, are transformed into binary data and fitted dichotomous IRT models for the sake of convenience. The distractor information is totally ignored when dichotomous IRT models (e.g., 2-parameter logistic model, 2PLM) are used. The second kind of IRT models were polytomous IRT models. In consideration of order of options in multiple-choice items, there are mainly two ways of modeling when polytomous models were used. One is transforming unordered categories into ordered ones and fitting ordered poly-tomous models, for example GPCM (Muraki, 1992). The other is fitting unordered polytomous models like Bock's (1982) nominal response model (NRM) and Thissen and Steinberg' s (1984) multiple-choice model (MCM) which was general model of NRM. The former way requires item options to be ordered. But multiple-choice items are unordered in many cases. The latter way is more flexible, and both NRM and MCM had been used to analysis empirical multiple-choice tests (Sadler, 1998; Thissen, Steinberg, &amp; Fitzpatrick, 1989). However, a limitation of NRM and MCM is the same level of all options to an item, which means that a respondent is possible to choose any option of an item every time he responses. Practically, high-ability respondents may choose the correct options directly without a glance of the distractors, while low-ability respondents may be attracted by several distractors. Therefore distractors in multiple-choice tests show a collapsibility property. As a result, Suh and Bolt (2010) proposed a framework of nested logit models (NLMs) for multiple-choice items. The 2-parameter logistic version of NLMs (2PL-NLM) can be demonstrated as</font></p>     <p><img src="/img/revistas/ap/v33n3/multidisciplinar1_formula1.jpg"></p>     <p><font face="Verdana" size="2">where <i>&#945;<sub>i</sub></i> and <i>&#946;<sub>i</sub></i> are the slop parameter and the difficulty parameter. Equation (1), the 2PLM term, defines the probability that a respondent of <i>&#952;<sub>j</sub></i> chooses the correct option on item <i>i</i>. A respondent whose ability <i>&#952;</i> exceeds the item difficulty (<i>&#946;)</i> will have a higher probability of correct response (<i>P(u<sub>ij</sub>=1&#1472;&#952;<sub>j</sub>))</i>. Meanwhile, a respondent whose ability can not reach the item difficulty will have a higher probability of incorrect response (<i>P(u<sub>ij</sub>=0&#1472;&#952;<sub>j</sub>)=1-P(u<sub>ij</sub>=1&#1472;&#952;<sub>j</sub>))</i>. The second term of the equation (2) is the NRM, nested in the 2PLM, describes a propensity toward each distractor category <i>v</i> conditional upon an incorrect response. That means the probability of incorrect response can be further separated into three probabilities in a four-choice item. The distractor "difficulties" (&#958;<sub>iv</sub>) determine which distractor will be probably chosen. Naturally, NLMs present a better approximation to individuals' response behaviors on a multiple-choice item.</font></p>     ]]></body>
<body><![CDATA[<p><font face="Verdana" size="2">2PL-NLM can be easily transformed to 3-parameter logistic nested logit model (3PL-NLM) which treats respondents' guessing behaviors by adding a "guessing" parameter. And also NLMs had been generalized to fit multidimensional multiple-choice tests (Bolt, Wollack, &amp; Suh, 2012). In this study, the unidimensional 2PL-NLM was used.</font></p>     <p><font face="Verdana" size="2"><b>Overview</b></font></p>     <p><font face="Verdana" size="2">Multiple-choice items may have more information than dichotomous items owing to distractors. However, in contrast with dichotomous items, they will also increase cognitive loads of respondents and item-writers, and they require a complicated model like 2PL-NLM rather than a simple one like 2PLM, for a 2PL-NLM has four more item parameters (&#955;<sub>1</sub>,&#955;<sub>2</sub>, &#958;<sub>1</sub>, &#958;<sub>1</sub>) than a 2PLM in a four-choice item. If a multiple-choice items can not enhance the ability evaluation, it is doubtful whether a multiple-choice format, rather than a simple dichotomous format, is necessary in psychological and educational tests.</font></p>     <p><font face="Verdana" size="2">In this study, 2PL-NLM is used to assess the ability evaluation based on multiple-choice items, since NLMs are more appropriate for multiple-choice items theoretically and conceptually as described above. Therefore, the purpose of present study is to clarify whether or not multiple-choice items, instead of dichotomous ones, should be used to enhance the ability evaluation. More specifically, whether destructor information from multiple-choice items (1) can improve the ability (person parameter) estimation, and (2) can offer some psychological explanations to respondents' distractor chosen behaviors. The rest of this paper is organized as follows. Study 1, a simulation study focuses on the enhancement of ability estimation, when 2PL-NLM is used in multiple-choice tests under different conditions, by contrast with dichotomous item tests fitted 2PLM. Study 2, a real multiple-choice test is used to assess distractor chosen behaviors of respondents at different ability levels.</font></p>     <p>&nbsp;</p>     <p><font face="Verdana" size="2"><b>Method</b></font></p>     <p><font face="Verdana" size="2"><b>Simulation Study</b></font></p>     <p><font face="Verdana" size="2">The purpose of simulation study is to investigate whether multiple-choice model can extract more distractor information to enhance ability estimation than dichotomous model do. For this purpose simulation study is composed of two parts. Part 1 is to describe whether multiple-choice model should be used or not in a multiple-choice test. And part 2 is to explore the optimal number of distractors should a multiple-choice item have.</font></p>     <p><font face="Verdana" size="2">Multiple-choice model is compared with a dichotomous model in part 1. 2PLM for dichotomous format and 2PL-NLM for multiple-choice format, rather than 3PLM and 3PL-NLM, are introduced for two reasons. First, eliminating the randomly guessing factor may present more pure enhancement by distractor information. Second, the empirical data used in this study showed a better fitness of 2PLM in a pilot analysis. The accuracy of ability estimation for the 2PLM and 2PL-NLM was evaluated for varying sample size (1000, 2000, and 4000 respondents) and test length (5-, 10-, 20-, 30-, 40-, and 50-item tests) conditions. The condition when the number of respondents is less than 1000 was not included. Because 2PL-NLM has much more item parameters than 2PLM. For example, There are 120 ((8-2)*20) item parameters to estimate in a 20-item test. While 2PLM only have 40 (2*20) item parameters. More item parameters needs larger sample size. Embretson and Reise (2000) recommended over 500 respondents when graded response model which has less item parameters than 2PL-NLM dose. 100 replications were executed for each combination of conditions. For each combination of conditions, Respondents' true ability were generated from &#952;&#8764;Normal (0, 1). On account of order of multiple-choice options, GPCM and NRM, presented ordered multiple-choice items and unordered multiple-choice items respectively, were used to generate responses. Item parameters were generated randomly from the following distributions: slope parameter a&#8764;Uniform (0.5, 2) and intercept parameters &#948;<sub>v</sub>&#8764;Uniform (-2, 2) for the GPCM, followed by the imposition of constraints <i>&#948;<sub>1</sub></i> &lt; <i>&#948;<sub>2</sub></i> &lt; <i>&#948;<sub>3</sub></i> (Muraki, 1992), and slop parameter &#955;<sub>v</sub>&#8764;Uniform (-2, 2) and intercept parameter &#955;<sub>v</sub>&#8764;Uniform (-2, 2) for the NRM, followed by the imposition of constraints <img src="/img/revistas/ap/v33n3/multidisciplinar1_formula2.jpg"> (Suh &amp; Bolt, 2010). Besides, all simulated items in this part included four options, namely one correct option and three distractors.</font></p>     <p><font face="Verdana" size="2">In part 2 of simulation study, a fully crossed design was implemented under following conditions: 3 (the number of distractors) &#215; 4 (test length). Specifically, test length was examined at four levels: 5 items, 10 items, 20 items and 30 items. The number of distractors was examined at three levels: 3 distractors (4 options), 4 distractors (5 options), and 5 distractors (6 options). The test-length conditions were designed based on some results of part 1. And the distractor conditions were the common settings in a real test. Response data of 2000 person were generated from 2PL-NLM with slop parameter <i>a<sub>i</sub></i>&#8764;Uniform (0.5, 2) and difficulty parameter <i>&#946;<sub>i</sub></i>&#8764;Uniform (-2, 2). The accuracy of ability estimation was evaluated under each condition.</font></p>     ]]></body>
<body><![CDATA[<p><font face="Verdana" size="2"><b>Data Analysis</b></font></p>     <p><font face="Verdana" size="2">To assess the accuracy of ability estimation, mean absolute bias (<i>M<sub>bias</sub>)</i> and standard deviation of absolute bias (<i>SD<sub>bias</sub></i>) were used. These two indices are commonly used to assess bias of estimators in Statistics. They present the mean and the standard deviation of all differences between the estimated values and the true values (&#124; <i><img src="/img/revistas/ap/v33n3/multidisciplinar1_signo2.jpg"><sub>j</sub></i> - <i>&#952;<sub>j</sub></i> &#124;) (Hofmann, 2007; Walther &amp; Moore, 2005). And they can be demonstrated as</font></p>     <p><img src="/img/revistas/ap/v33n3/multidisciplinar1_formula3.jpg"></p>     <p><font face="Verdana" size="2">where <img src="/img/revistas/ap/v33n3/multidisciplinar1_signo2.jpg"> is estimated from model. Small indices indicate that the estimating bias and the estimating variance are small. Both simulation study and empirical study were administrated in R Project version 3.1.2.</font></p>     <p><font face="Verdana" size="2"><b>Empirical Study</b></font></p>     <p><font face="Verdana" size="2">A 32-item Chinese Vocabulary Test for Grade 1 (CVT-G1) data was analyzed. This test is one of a battery of Chinese Vocabulary Tests which contains twelve tests for twelve grades from Grade 1 of primary school to Grade 12 of high school in China. All these tests were constructed based on 2PLM and composed of five-option multiple-choice items (Cao, 1999). 1035 grade 1 students' responses were used to fit three models: 2PLM, NRM and 2PL-NLM in this study. The average of discrimination parameters from 2PLM is 1.050, with a range of 0.554 to 1.720, and the average of difficulty parameters is -0.571, with a range of -1.631 to 0.875.</font></p>     <p><font face="Verdana" size="2">The advantages of NLMs and an example of distractor analysis procedure will be discussed in this study. First, the performance of NLM on short tests was explored. 32-item CVT-G1 was used to construct another three versions of short tests: 24-item tests, 16-item tests and 8-item tests. Three short tests were constructed by randomly canceling items from the full CVT-G1 test (32 items). The ability parameters obtained from full test were treated as "true values" (<i><img src="/img/revistas/ap/v33n3/multidisciplinar1_signo2.jpg"></i>), because estimating values from long test is supposed to be more accurate than short test theoretically. And the person parameters from short CVT-G1 tests were treated as target estimating values (<i><img src="/img/revistas/ap/v33n3/multidisciplinar1_signo2.jpg"></i>). The randomly canceling procedure had repeated 30 times for each length version. Then the average of <i>M<sub>bia</sub>s</i> and <i>SD<sub>bias</sub>s</i> were calculated. In second part of this study, probabilities of distractor responses were used to extract the potential psychological meaning.</font></p>     <p>&nbsp;</p>     <p><font face="Verdana" size="2"><b>Results</b></font></p>     <p><font face="Verdana" size="2"><b>Enhancement of Ability Estimation</b></font></p>     ]]></body>
<body><![CDATA[<p><font face="Verdana" size="2"><a target="_blank" href="/img/revistas/ap/v33n3/multidisciplinar1_tabla1.jpg">Table 1</a> presents the biases (<i>M<sub>bias</sub>s)</i> of ability estimation under each conditions. Estimating biases in GPCM and NRM are treated as baselines respectively. Theoretically, more items a test has, more accurate ability parameters (&#952;s) will be gained. Results from <a target="_blank" href="/img/revistas/ap/v33n3/multidisciplinar1_tabla1.jpg">Table 1</a> show that the enhancement of ability estimation appears under each condition when 2PL-NLM is used. Moreover when the number of items is less than 30, the performance of 2PL-NLM is close to the basic model NRM and GPCM, while 2PLM presents greater bias estimating especially when GPCM is the basic model. In general, 2PL-NLM shows smaller estimating bias than 2PLM when either distractors are ordered or unordered. And 2PLM may lose more information of ability estimation under ordered condition. In addition, the biases rise when the number of items is over 30. That is because more item parameters require more respondents.</font></p>     <p><font face="Verdana" size="2">The <i>M<sub>bias</sub></i> can provide how accurate the estimation is, and the <i>SD<sub>bias</sub></i> can provide variation information of estimation. <a href="#f1">Figure 1</a> demonstrates the <i>SD<sub>bias</sub></i> under various test length conditions when 2000 persons were generated. The other two sample size conditions are not here since they show similar shapes. The decreasing tendency of all <i>SD<sub>bias</sub></i> curves show that more item information enhances the stability of estimation. <a href="#f1">Figure 1</a> also demonstrates that the distance between NLM curves and baselines (NRM curve, GPCM curve) are much smaller than the distance between 2PLM ones and baselines under 2000-person condition before 30-item condition. In other words, the ability estimation based on NLM is as stable as baseline, more stable than 2PLM. In addition, it is notable that the difference between the NLM and 2PLM decreases. This indicates NLM is losing its ad vantages of distractor information gradually, as the number of items increases.</font></p>     <p>&nbsp;</p>     <p align="center"><a name="f1"></a><img src="/img/revistas/ap/v33n3/multidisciplinar1_figura1.jpg"></p>     <p>&nbsp;</p>     <p><font face="Verdana" size="2">Based on the results above the optimal number of distractors under four different test-length conditions had been explored. Three distractor conditions (three distractors, four distractors and five distractors) had been discussed, because too few (e.g., two) or too many (e.g., more than five) distractors are rarely used in reality. The results of estimating biases under different conditions are shown in <a target="_blank" href="/img/revistas/ap/v33n3/multidisciplinar1_tabla2.jpg">Table 2</a>. More distractors only promote slight enhancement of ability estimation in short tests. Considering that the challenges of more distractors in five- or six- choice items, four-choice items are accurate enough.</font></p>     <p><font face="Verdana" size="2">Next, <img src="/img/revistas/ap/v33n3/multidisciplinar1_signo2.jpg"> s were divided into five levels to look into the details. Date generated from NRM was used because of its generalizability. Four conditions (5-, 10-, 20-, and 30-item) were taken into consideration for reason of results shown above. The differences of <i>M<sub>bias</sub></i> between 2PLM and 2PL-NLM are shown in <a href="#t3">Table 3</a>. Results show that there are very small differences of estimating bias at intermediate levels, larger differences at high levels, and the largest differences at the low levels. It can be inferred that respondents at low levels chose more distractors which could offer more information by using NLM.</font></p>     <p>&nbsp;</p>     <p align="center"><a name="t3"></a><img src="/img/revistas/ap/v33n3/multidisciplinar1_tabla3.jpg"></p>     <p>&nbsp;</p>     ]]></body>
<body><![CDATA[<p><font face="Verdana" size="2"><b>Ability Evaluation Based on Empirical Data</b></font></p>     <p><font face="Verdana" size="2">The simulation results have shown that distractor information can enhance the accuracy of person parameter estimation by using NLM. In other words, NLM will fit the short tests better than 2PLM. The averages of <i>M<sub>bias</sub>s</i> and <i>SD<sub>bias</sub>s</i>, presented the bias between the full-test estimation and short-test estimation, are shown in <a href="#t4">Table 4</a>. The results illustrate that the <i>M<sub>bias</sub></i> based on 2PLM is as good as 2PL-NLM, <i>SD<sub>bias</sub></i> even a bit better by using 24-item test. But when shorter tests are used, both two indices based on 2PL-NLM are smaller than 2PLM. These results also prove that distractors can offer more information for ability estimation in short tests.</font></p>     <p>&nbsp;</p>     <p align="center"><a name="t4"></a><img src="/img/revistas/ap/v33n3/multidisciplinar1_tabla4.jpg"></p>     <p>&nbsp;</p>     <p><font face="Verdana" size="2">Except for the enhancement of estimation, the probabilities of responses on each distractor obtained from NLM are available. Two steps were carried out to explore the students' response behavior on distractors. Step 1, divide 1035 respondents into five levels by their estimating <i>&#952;</i> values as follow: level 1 &#952; &#8712; (-&#8734;,-2), level 2&#952; &#8712; (-2,-1), level 3 &#952; &#8712; (-1,-1), level 4 &#952; &#8712; (1,2), level 5 &#952; &#8712; (2,+&#8734;). Step 2, calculate the mean probabilities of respondents at each level on every distractor. These probabilities can reveal the degree of distractor attractiveness for different levels of respondents. Take item 11 for example (see <a target="_blank" href="/img/revistas/ap/v33n3/multidisciplinar1_tabla5.jpg">Table 5</a>). The stem of item 11 is <img src="/img/revistas/ap/v33n3/multidisciplinar1_signo.jpg"> (hurry up)". Respondents were required to choose the best interpretation to this phrase. If a respondent randomly chooses an option, the probability will be 0.2 for a five-choice item. So the probabilities over 0.2 were picked out. Distractor 1 is the most attractive distractor for level 1 respondents with a probability of 0.329, and distractor 3 is the most attractive distractor for level 3 and level 4 respondents with probabilities of 0.263 and 0.220. For level 2 respondents, the probabilities of distractor 1 and distractor 3 are proximate. It reflects that higher vocabulary ability respondents could be attracted by distractor 3 rather than distractor 1. Obviously, these probabilities form a changing trace of distractor responses. And a developmental psychological explanation can be given by analyzing the distractor contents along this changing trace. Distractor 1 shares the same first Chinese character with item stem. Distractor 3 is something about time as well as stem. This contents analysis reveals that respondents at low levels interpreted the item stem in terms of images, and respondents at higher levels began to understand the abstract meanings of words gradually. The procedure of analyzing response probability changing trace is meaningful, but it is impossible to explain all the changing traces item by item in this paper.</font></p>     <p>&nbsp;</p>     <p><font face="Verdana" size="2"><b>Discussion and Conclusion</b></font></p>     <p><font face="Verdana" size="2"><b>Model Selection</b></font></p>     <p><font face="Verdana" size="2">Model selection for multiple-choice items has been discussed for a long time. Some conclusions were indeed conflicting (Divgi, 1986; Henning, 1989). In practical assessment, a few researchers used various polytomous models for multiple-choice tests. Much more administrators used dichotomous models for them. However, neither polytomous models nor dichotomous models were proposed upon multiple-choice data. NLMs model the response behaviors in a multiple-choice test theoretically (Suh &amp; Bolt, 2010), so it is worth investigating.</font></p>     ]]></body>
<body><![CDATA[<p><font face="Verdana" size="2">The simulation study of this research showed the conditions under which NLMs should be used rather than simple 2PLM in terms of the enhancement of ability estimation. Obviously, distractor information was effective when short tests were used, especially for the test below 30 items. Increasing the number of item would provide more information for both NLM and 2PLM. But when the number of item exceeded 40, a large number of NLM item parameters might bring in a negative effect. Therefore, it is suggested that a shorter multiple-choice test is acceptable by using NLMs. Yet if the multiple-choice test is too long, over 30 items for example, dichotomous models can offer accurate and stable estimation, and NLMs will not be recommended. With respect to the order of item options, NRM and GPCM were used to generate responses. The results showed that estimating biases of 2PL-NLM under GPCM condition were smaller than the ones under NRM condition. On the contrary, estimating biases of 2PLM under GPCM condition were larger than the ones under NRM condition. So if an ordered multiple-choice item, in which options may represent the cognitive level of respondents, is used, recoding multiple-choice data into binary data will lose more useful information.</font></p>     <p><font face="Verdana" size="2">The empirical study results were similar with simulation study results. However, the estimation differences across three versions of tests (24-item test, 16-item test and 8-item test) between two models were smaller in empirical study than they were in simulation study. For one reason, the <i>&#952;s</i> from real full test (32-item test) were not "true". For another reason, the CVT-G1 is an easy test (mean difficulty parameter equal to -0.571), and easy test means fewer distractor chosen behaviors.</font></p>     <p><font face="Verdana" size="2">It is hard to tell if NLMs could offer more accurate estimation than NRM in this study. However, the NRM was proposed to model the nominal response rather than multiple-choice response. There are no correct option and dis-tractors constructionally, so it is troublesome to explain whether an item is discriminating or not. Yet NLM inherited the advantage of 2PLM in this aspect. Two correlation coefficients of discrimination parameters (slope parameters) of NLM, NRM, and 2PLM were calculated. The correlation between NLM and 2PLM is 0.991, instead 0.647 between NRM and 2PLM. In a word, NLM can be used to guide the item construction and revision as conveniently as 2PLM because of the 2PL-term, and also can provide more distractor information to make better ability estimating.</font></p>     <p><font face="Verdana" size="2"><b>Distractor Analysis</b></font></p>     <p><font face="Verdana" size="2">Distractor information can not only enhance the ability estimation, but also provide some psychological explanation. By analyzing the changing traces of distractor response probabilities together with distractor contents, some meaningful psychological inference could be drawn. A good multiple-choice item with good distractors can indicate that which trait level the respondents are, and it also can reflect some cognitive developmental information and thinking strategies.</font></p>     <p><font face="Verdana" size="2">Distractors are a kind of wrong options on earth, since high level respondents choose few. This could be concluded from results of <a href="#t3">Table 3</a>. And it is unnecessary to add more options to a four-option item according to the results in <a target="_blank" href="/img/revistas/ap/v33n3/multidisciplinar1_tabla2.jpg">Table 2</a>. Multiple-choice item with three or four distractors is recommended in item writing. Furthermore, when test length over 30, distractor information can help little.</font></p>     <p><font face="Verdana" size="2"><b>Limitation and Future Research</b></font></p>     <p><font face="Verdana" size="2">The conclusions resulted from this study were based on 2PL-NLM. That is, the guessing behaviors were ignored and the dimension of test was unique. In some cases, tests are multidimensional and respondents may use guessing strategies. Consequently, the performance of multiple-choice item on ability evaluation could be different from this study. And also explanation to the distractor chosen behaviors should be much more complex. So the more generalized multiple-choice model (e.g., 3PL-NLM) have to be discussed in those cases.</font></p>     <p><font face="Verdana" size="2">Less items and more accurate is an ideal aim of psychological assessment. On this point of view, making maximally use of item information to enhancing the estimating accuracy based on multiple-choice items by using NLMs is, to some extent, similar to computerized adaptive testing (CAT). And multiple-choice items are also popular in CAT. However, the conclusion from this study was based on paper-pencil test. So how to applying NLMs to CAT with multiple-choice items still need to be explored.</font></p>     <p><font face="Verdana" size="2">The changing trace of distractor response is a simple way to explain the response behaviors of respondents. But sometimes it is difficult to directly analyze distractors from a long test. Future research will focus on how to establish a more effective distractor analysis procedure to extract the explanatory information for practical application.</font></p>     ]]></body>
<body><![CDATA[<p>&nbsp;</p>     <p><font face="Verdana" size="2"><b>References</b></font></p>     <!-- ref --><p><font face="Verdana" size="2">1. Attali, Y, &amp; Fraenkel, T. (2000). The Point-Biserial as a Discrimination Index for Distractors in Multiple-Choice Items: Deficiencies in Usage and an alternative. Journal of Educational Measurement, 37(1), 77-86. doi: 10.1111/j.1745-3984.2000.tb01077.x.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787652&pid=S0212-9728201700030001000001&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">2. Bock, R. D. (1972). Estimating item parameters and latent ability when responses are scored in two or more nominal categories. Psychometrika, 37, 29-51. doi: 10.1007/BF02291411.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787654&pid=S0212-9728201700030001000002&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">3. Bolt, D. M., Wollack, J. A., &amp; Suh, Y. (2012). Application of a multidimensional nested logit model to multiple-choice test items. Psychometrika, 77, 339-357. doi: 10.1007/S11336-012-9257-5.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787656&pid=S0212-9728201700030001000003&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">4. Briggs, D. C., Alonzo, A. C., Schwab, C., &amp; Wilson, M. (2006). Diagnostic assessment with ordered multiple-choice items. Educational Assessment, 11, 33-63. doi: 10.1207/s15326977ea1101_2.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787658&pid=S0212-9728201700030001000004&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    ]]></body>
<body><![CDATA[<!-- ref --><p><font face="Verdana" size="2">5. Cao, Y W (1999). Construction of vocabulary tests for junior school level. Acta Psychologica Sinica, 31, 460-467.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787660&pid=S0212-9728201700030001000005&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">6. Davis, F. B., &amp; Fifer, G. (1959). The effect on test reliability and validity of scoring aptitude and achievement tests with weights for every choice. Educational and Psychological Measurement, 19, 159-170. doi: 10.1177/001316445901900202.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787662&pid=S0212-9728201700030001000006&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">7. Divgi, D. R. (1986). Does the Rasch model really work for multiple choice items? Not if you look closely. Journal of Educational Measurement, 23, 283-298. doi: 10.1111/j.1745-3984.1986.tb00251.x.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787664&pid=S0212-9728201700030001000007&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">8. Drasgow, F., Levine, M. V., &amp; Williams, E. A. (1985). Appropriateness Measurement with Polychotomous Item Response Models and Standardized Indices. British Journal of Mathematical and Statistical Psychology, 38, 67-86. doi: 10.1111/j.2044-8317.1985.tb00817.x.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787666&pid=S0212-9728201700030001000008&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">9. Embretson, S. E., &amp; Reise, S. P. (2000). Item Response Theory for Psychologists. Mahwah, New Jersey: Lawrence Erlbaum Associates, Inc.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787668&pid=S0212-9728201700030001000009&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    ]]></body>
<body><![CDATA[<!-- ref --><p><font face="Verdana" size="2">10. Green, B. F., Crone, C. R., &amp; Folk, V. G. (1989). A Method for Studying Differential Distractor Functioning. Journal of Educational Measurement, 26, 147-160. doi: 10.1111/j.1745-3984.1989.tb00325.x.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787670&pid=S0212-9728201700030001000010&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">11. Haladyna, T. M., &amp; Downing, S. M. (1989). A Taxonomy of Multiple-Choice Item-Writing Rules. Applied Measurement in Education, 2, 37-50. doi: 10.1207/s15324818ame0201_3.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787672&pid=S0212-9728201700030001000011&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">12. Haladyna, T. M., &amp; Downing, S. M. (1993). How Many Options is Enough for a Multiple-Choice Testing Item. Educational and Psychological Measurement, 53, 999-1010. doi: 10.1177/0013164493053004013.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787674&pid=S0212-9728201700030001000012&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">13. Haladyna, T. M., Downing, S. M., &amp; Rodriguez, M. C. (2002). A Review of Multiple-Choice Item-Writing Guidelines for Classroom Assessment. Applied Measurement in Education, 15, 309-333. doi: 10.1207/S15324818AME1503_5.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787676&pid=S0212-9728201700030001000013&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">14. Henning, G. (1989). Does the Rasch model really work for multiple-choice items? Take another look: a response to Divgi. Journal of Educational Measurement, 26, 91-97. doi: 10.1111/j.1745-3984.1989.tb00321.x.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787678&pid=S0212-9728201700030001000014&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    ]]></body>
<body><![CDATA[<!-- ref --><p><font face="Verdana" size="2">15. Hofmann, K. P. (2007). Psychology of Decision Making in Economics, Business and Finance. Nova Publishers.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787680&pid=S0212-9728201700030001000015&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">16. Jacobs, P. I., &amp; Vandeventer, M. (1970). Information in wrong responses. Psychological Reports, 26, 311-315. doi: 10.2466/pr0.1970.26.1.311.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787682&pid=S0212-9728201700030001000016&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">17. Kim, J. (2006). Using the Distractor Categories of Multiple-Choice Items to Improve IRT Linking. Journal of Educational Measurement, 43, 193-213. doi: 10.1111/j.1745-3984.2006.00013.x.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787684&pid=S0212-9728201700030001000017&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">18. Levine, M. V., &amp; Drasgow, F. (1983). The relation between incorrect option choice and estimated ability. Educational and Psychological Measurement, 43, 675-685. doi: 10.1177/001316448304300301.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787686&pid=S0212-9728201700030001000018&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">19. Liu, O. L., Lee, H., &amp; Linn, M. C. (2011). An investigation of explanation multiple-choice items in science assessment. Educational Assessment, 16, 164-184. doi: 10.1080/10627197.2011.611702.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787688&pid=S0212-9728201700030001000019&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    ]]></body>
<body><![CDATA[<!-- ref --><p><font face="Verdana" size="2">20. Love, T. E. (1997). Distractor selection ratios. Psychometrika, 62, 51-62. doi: 10.1007/BF02294780.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787690&pid=S0212-9728201700030001000020&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">21. Luecht, R. M. (2007). Using information from multiple-choice distractors to enhance cognitive-diagnostic score reporting. In J. P Leighton &amp; M. J. Gierl (Eds.), Cognitive diagnostic assessment for education: Theory and practices (pp. 319-340). Cambridge University Press. doi: 10.1017/CBO9780511611186.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787692&pid=S0212-9728201700030001000021&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">22. Muraki, E. (1992). A Generalized Partial Credit Model: Application of an EM Algorithm. Applied Psychological Measurement, 16, 159-176. doi: 10.1002/j.2333-8504.1992.tb01436.x.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787694&pid=S0212-9728201700030001000022&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">23. Penfield, R. D. (2011). How are the Form and Magnitude of DIF Effects in Multiple-Choice Items Determined by Distractor-Level Invariance Effects?. Educational And Psychological Measurement, 71, 54-67. doi: 10.1177/0013164410387340.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787696&pid=S0212-9728201700030001000023&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">24. Roediger III, H. L., &amp; Marsh, E. J. (2005). The positive and negative consequences of multiple-choice testing. Journal of Experimental Psychology: Learning, Memory, and Cognition, 31, 1155. doi: 10.1037/0278-7393.31.5.1155.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787698&pid=S0212-9728201700030001000024&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    ]]></body>
<body><![CDATA[<!-- ref --><p><font face="Verdana" size="2">25. Sadler, P M. (1998). Psychometric models of student conceptions in science: Reconciling qualitative studies and distractor-driven assessment instruments. Journal of Research in science Teaching, 35, 265-296. doi: 10.1002/(SICI)1098-2736(199803)35:3&lt;265::AID-TEA3&gt;3.0.CO;2-P.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787700&pid=S0212-9728201700030001000025&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">26. Sigel, I. E. (1963). How intelligence tests limit understanding of intelligence. Merrill-Paker Quarterly, 9, 39-56.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787702&pid=S0212-9728201700030001000026&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">27. Suh, Y, &amp; Bolt, D. M. (2010). Nested logit models for multiple-choice item response data. Psychometrika, 75, 454-473. doi: 10.1007/s11336-010-9163-7.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787704&pid=S0212-9728201700030001000027&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">28. Suh, Y, &amp; Bolt, D. M. (2011). A Nested Logit Approach for Investigating Distractors as Cause of Different Item Functioning. Journal of Educational Measurement, 48, 188-205. doi: 10.1111/j.1745-3984.2011.00139.x.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787706&pid=S0212-9728201700030001000028&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">29. Suh, Y., &amp; Talley, A. E. (2015). An Empirical Comparison of DDF Detection Methods for Understanding the Causes of DIF in Multiple-Choice Items. Applied Measurement in Education, 28, 48-67. doi: 10.1080/08957347.2014.973560.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787708&pid=S0212-9728201700030001000029&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    ]]></body>
<body><![CDATA[<!-- ref --><p><font face="Verdana" size="2">30. Tamir, P (1971). An alternative approach to the construction of multiple choice test items. Journal of Biological Education, 5, 305-307. doi: 10.1080/00219266.1971.9653728.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787710&pid=S0212-9728201700030001000030&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">31. Tamir, P (1989). Some issues related to the use of justifications to multiple-choice answers. Journal of Biological Education, 23, 285-292. doi: 10.1080/00219266.1989.9655083.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787712&pid=S0212-9728201700030001000031&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">32. Thissen, D. M. (1976). Information in wrong responses to the Raven Progressive Matrices. Journal of Educational Measurement, 13, 201-214. doi: 10.1111/j.1745-3984.1976.tb00011.x.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787714&pid=S0212-9728201700030001000032&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">33. Thissen, D., &amp; Steinberg, L. (1984). A Response Model for Multiple Choice Items. Psychometrika, 49, 501-519. doi: 10.1007/BF02302588.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787716&pid=S0212-9728201700030001000033&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">34.Thissen, D., Steinberg, L., &amp; Fitzpatrick, A. R. (1989). Multiple-Choice Models: The Distractors Are Also Part of the Item. Journal of Educational Measurement, 26, 161-176. doi: 10.1111 /j.1745-3984.1989.tb00326.x.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787718&pid=S0212-9728201700030001000034&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    ]]></body>
<body><![CDATA[<!-- ref --><p><font face="Verdana" size="2">35. Walther B. A., &amp; Moore J. L. (2005). The concepts of bias, precision and accuracy, and their use in testing the performance of species richness estimators, with a literature review of estimator performance. Ecography, 28, 815-829. doi: 10.1111/j.2005.0906-7590.04112.x.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787720&pid=S0212-9728201700030001000035&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>    <!-- ref --><p><font face="Verdana" size="2">36. Wollack, J. A. (1997). A Nominal Response Model Approach for Detecting Answer Copying. Applied Psychological Measurement, 21, 307-320. doi: 10.1177/01466216970214002.    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;[&#160;<a href="javascript:void(0);" onclick="javascript: window.open('/scielo.php?script=sci_nlinks&ref=787722&pid=S0212-9728201700030001000036&lng=','','width=640,height=500,resizable=yes,scrollbars=1,menubar=yes,');">Links</a>&#160;]<!-- end-ref --></font></p>     <p>&nbsp;</p>     <p>&nbsp;</p>     <p><font face="Verdana" size="2"><a href="#top"><img border="0" src="/img/revistas/ap/v33n3/seta.gif" width="15" height="17"></a><a name="bajo"></a><b>Correspondence:</b>    <br>Xin Tao.    <br>Collaborative Innovation Center of Assessment    <br>toward Basic Education Quality,    ]]></body>
<body><![CDATA[<br>Beijing Normal University,    <br>Beijing, (P. R. China).    <br>E-mail: <a href="mailto:mikebonita@sina.cn">mikebonita@sina.cn</a></font></p>     <p><font face="Verdana" size="2">Article received: 02-10-2015    <br>revised: 09-03-2016    <br>accepted: 16-05-2016</font></p>      ]]></body><back>
<ref-list>
<ref id="B1">
<label>1</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Attali]]></surname>
<given-names><![CDATA[Y]]></given-names>
</name>
<name>
<surname><![CDATA[Fraenkel]]></surname>
<given-names><![CDATA[T.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[The Point-Biserial as a Discrimination Index for Distractors in Multiple-Choice Items: Deficiencies in Usage and an alternative]]></article-title>
<source><![CDATA[Journal of Educational Measurement]]></source>
<year>2000</year>
<volume>37</volume>
<numero>1</numero>
<issue>1</issue>
<page-range>77-86</page-range></nlm-citation>
</ref>
<ref id="B2">
<label>2</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Bock]]></surname>
<given-names><![CDATA[R. D.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Estimating item parameters and latent ability when responses are scored in two or more nominal categories]]></article-title>
<source><![CDATA[Psychometrika]]></source>
<year>1972</year>
<volume>37</volume>
<page-range>29-51</page-range></nlm-citation>
</ref>
<ref id="B3">
<label>3</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Bolt]]></surname>
<given-names><![CDATA[D. M.]]></given-names>
</name>
<name>
<surname><![CDATA[Wollack]]></surname>
<given-names><![CDATA[J. A.]]></given-names>
</name>
<name>
<surname><![CDATA[Suh]]></surname>
<given-names><![CDATA[Y.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Application of a multidimensional nested logit model to multiple-choice test items]]></article-title>
<source><![CDATA[Psychometrika]]></source>
<year>2012</year>
<volume>77</volume>
<page-range>339-357</page-range></nlm-citation>
</ref>
<ref id="B4">
<label>4</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Briggs]]></surname>
<given-names><![CDATA[D. C.]]></given-names>
</name>
<name>
<surname><![CDATA[Alonzo]]></surname>
<given-names><![CDATA[A. C.]]></given-names>
</name>
<name>
<surname><![CDATA[Schwab]]></surname>
<given-names><![CDATA[C.]]></given-names>
</name>
<name>
<surname><![CDATA[Wilson]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Diagnostic assessment with ordered multiple-choice items]]></article-title>
<source><![CDATA[Educational Assessment]]></source>
<year>2006</year>
<volume>11</volume>
<page-range>33-63</page-range></nlm-citation>
</ref>
<ref id="B5">
<label>5</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Cao]]></surname>
<given-names><![CDATA[Y W]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Construction of vocabulary tests for junior school level]]></article-title>
<source><![CDATA[Acta Psychologica Sinica]]></source>
<year>1999</year>
<volume>31</volume>
<page-range>460-467</page-range></nlm-citation>
</ref>
<ref id="B6">
<label>6</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Davis]]></surname>
<given-names><![CDATA[F. B.]]></given-names>
</name>
<name>
<surname><![CDATA[Fifer]]></surname>
<given-names><![CDATA[G.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[The effect on test reliability and validity of scoring aptitude and achievement tests with weights for every choice]]></article-title>
<source><![CDATA[Educational and Psychological Measurement]]></source>
<year>1959</year>
<volume>19</volume>
<page-range>159-170</page-range></nlm-citation>
</ref>
<ref id="B7">
<label>7</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Divgi]]></surname>
<given-names><![CDATA[D. R.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Does the Rasch model really work for multiple choice items?: Not if you look closely]]></article-title>
<source><![CDATA[Journal of Educational Measurement]]></source>
<year>1986</year>
<volume>23</volume>
<page-range>283-298</page-range></nlm-citation>
</ref>
<ref id="B8">
<label>8</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Drasgow]]></surname>
<given-names><![CDATA[F.]]></given-names>
</name>
<name>
<surname><![CDATA[Levine]]></surname>
<given-names><![CDATA[M. V.]]></given-names>
</name>
<name>
<surname><![CDATA[Williams]]></surname>
<given-names><![CDATA[E. A.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Appropriateness Measurement with Polychotomous Item Response Models and Standardized Indices]]></article-title>
<source><![CDATA[British Journal of Mathematical and Statistical Psychology]]></source>
<year>1985</year>
<volume>38</volume>
<page-range>67-86</page-range></nlm-citation>
</ref>
<ref id="B9">
<label>9</label><nlm-citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Embretson]]></surname>
<given-names><![CDATA[S. E.]]></given-names>
</name>
<name>
<surname><![CDATA[Reise]]></surname>
<given-names><![CDATA[S. P.]]></given-names>
</name>
</person-group>
<source><![CDATA[Item Response Theory for Psychologists]]></source>
<year>2000</year>
<publisher-loc><![CDATA[Mahwah^eNew Jersey New Jersey]]></publisher-loc>
<publisher-name><![CDATA[Lawrence Erlbaum Associates]]></publisher-name>
</nlm-citation>
</ref>
<ref id="B10">
<label>10</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Green]]></surname>
<given-names><![CDATA[B. F.]]></given-names>
</name>
<name>
<surname><![CDATA[Crone]]></surname>
<given-names><![CDATA[C. R.]]></given-names>
</name>
<name>
<surname><![CDATA[Folk]]></surname>
<given-names><![CDATA[V. G.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[A Method for Studying Differential Distractor Functioning]]></article-title>
<source><![CDATA[Journal of Educational Measurement]]></source>
<year>1989</year>
<volume>26</volume>
<page-range>147-160</page-range></nlm-citation>
</ref>
<ref id="B11">
<label>11</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Haladyna]]></surname>
<given-names><![CDATA[T. M.]]></given-names>
</name>
<name>
<surname><![CDATA[Downing]]></surname>
<given-names><![CDATA[S. M.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[A Taxonomy of Multiple-Choice Item-Writing Rules]]></article-title>
<source><![CDATA[Applied Measurement in Education]]></source>
<year>1989</year>
<volume>2</volume>
<page-range>37-50</page-range></nlm-citation>
</ref>
<ref id="B12">
<label>12</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Haladyna]]></surname>
<given-names><![CDATA[T. M.]]></given-names>
</name>
<name>
<surname><![CDATA[Downing]]></surname>
<given-names><![CDATA[S. M.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[How Many Options is Enough for a Multiple-Choice Testing Item]]></article-title>
<source><![CDATA[Educational and Psychological Measurement]]></source>
<year>1993</year>
<volume>53</volume>
<page-range>999-1010</page-range></nlm-citation>
</ref>
<ref id="B13">
<label>13</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Haladyna]]></surname>
<given-names><![CDATA[T. M.]]></given-names>
</name>
<name>
<surname><![CDATA[Downing]]></surname>
<given-names><![CDATA[S. M.]]></given-names>
</name>
<name>
<surname><![CDATA[Rodriguez]]></surname>
<given-names><![CDATA[M. C.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[A Review of Multiple-Choice Item-Writing Guidelines for Classroom Assessment]]></article-title>
<source><![CDATA[Applied Measurement in Education]]></source>
<year>2002</year>
<volume>15</volume>
<page-range>309-333</page-range></nlm-citation>
</ref>
<ref id="B14">
<label>14</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Henning]]></surname>
<given-names><![CDATA[G.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Does the Rasch model really work for multiple-choice items? Take another look: a response to Divgi]]></article-title>
<source><![CDATA[Journal of Educational Measurement]]></source>
<year>1989</year>
<volume>26</volume>
<page-range>91-97</page-range></nlm-citation>
</ref>
<ref id="B15">
<label>15</label><nlm-citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Hofmann]]></surname>
<given-names><![CDATA[K. P.]]></given-names>
</name>
</person-group>
<source><![CDATA[Psychology of Decision Making in Economics, Business and Finance]]></source>
<year>2007</year>
<publisher-name><![CDATA[Nova Publishers]]></publisher-name>
</nlm-citation>
</ref>
<ref id="B16">
<label>16</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Jacobs]]></surname>
<given-names><![CDATA[P. I.]]></given-names>
</name>
<name>
<surname><![CDATA[Vandeventer]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Information in wrong responses]]></article-title>
<source><![CDATA[Psychological Reports]]></source>
<year>1970</year>
<volume>26</volume>
<page-range>311-315</page-range></nlm-citation>
</ref>
<ref id="B17">
<label>17</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Kim]]></surname>
<given-names><![CDATA[J.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Using the Distractor Categories of Multiple-Choice Items to Improve IRT Linking]]></article-title>
<source><![CDATA[Journal of Educational Measurement]]></source>
<year>2006</year>
<volume>43</volume>
<page-range>193-213</page-range></nlm-citation>
</ref>
<ref id="B18">
<label>18</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Levine]]></surname>
<given-names><![CDATA[M. V.]]></given-names>
</name>
<name>
<surname><![CDATA[Drasgow]]></surname>
<given-names><![CDATA[F.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[The relation between incorrect option choice and estimated ability]]></article-title>
<source><![CDATA[Educational and Psychological Measurement]]></source>
<year>1983</year>
<volume>43</volume>
<page-range>675-685</page-range></nlm-citation>
</ref>
<ref id="B19">
<label>19</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Liu]]></surname>
<given-names><![CDATA[O. L.]]></given-names>
</name>
<name>
<surname><![CDATA[Lee]]></surname>
<given-names><![CDATA[H.]]></given-names>
</name>
<name>
<surname><![CDATA[Linn]]></surname>
<given-names><![CDATA[M. C.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[An investigation of explanation multiple-choice items in science assessment]]></article-title>
<source><![CDATA[Educational Assessment]]></source>
<year>2011</year>
<volume>16</volume>
<page-range>164-184</page-range></nlm-citation>
</ref>
<ref id="B20">
<label>20</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Love]]></surname>
<given-names><![CDATA[T. E.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Distractor selection ratios]]></article-title>
<source><![CDATA[Psychometrika]]></source>
<year>1997</year>
<volume>62</volume>
<page-range>51-62</page-range></nlm-citation>
</ref>
<ref id="B21">
<label>21</label><nlm-citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Luecht]]></surname>
<given-names><![CDATA[R. M.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Using information from multiple-choice distractors to enhance cognitive-diagnostic score reporting]]></article-title>
<person-group person-group-type="editor">
<name>
<surname><![CDATA[Leighton]]></surname>
<given-names><![CDATA[J. P]]></given-names>
</name>
<name>
<surname><![CDATA[Gierl]]></surname>
<given-names><![CDATA[M. J.]]></given-names>
</name>
</person-group>
<source><![CDATA[Cognitive diagnostic assessment for education: Theory and practices]]></source>
<year>2007</year>
<page-range>319-340</page-range><publisher-name><![CDATA[Cambridge University Press]]></publisher-name>
</nlm-citation>
</ref>
<ref id="B22">
<label>22</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Muraki]]></surname>
<given-names><![CDATA[E.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[A Generalized Partial Credit Model: Application of an EM Algorithm]]></article-title>
<source><![CDATA[Applied Psychological Measurement]]></source>
<year>1992</year>
<volume>16</volume>
<page-range>159-176</page-range></nlm-citation>
</ref>
<ref id="B23">
<label>23</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Penfield]]></surname>
<given-names><![CDATA[R. D.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[How are the Form and Magnitude of DIF Effects in Multiple-Choice Items Determined by Distractor-Level Invariance Effects?]]></article-title>
<source><![CDATA[Educational And Psychological Measurement]]></source>
<year>2011</year>
<volume>71</volume>
<page-range>54-67</page-range></nlm-citation>
</ref>
<ref id="B24">
<label>24</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Roediger III]]></surname>
<given-names><![CDATA[H. L.]]></given-names>
</name>
<name>
<surname><![CDATA[Marsh]]></surname>
<given-names><![CDATA[E. J.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[The positive and negative consequences of multiple-choice testing]]></article-title>
<source><![CDATA[Journal of Experimental Psychology: Learning, Memory, and Cognition]]></source>
<year>2005</year>
<volume>31</volume>
<page-range>1155</page-range></nlm-citation>
</ref>
<ref id="B25">
<label>25</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Sadler]]></surname>
<given-names><![CDATA[P M.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Psychometric models of student conceptions in science: Reconciling qualitative studies and distractor-driven assessment instruments]]></article-title>
<source><![CDATA[Journal of Research in science Teaching]]></source>
<year>1998</year>
<volume>35</volume>
<page-range>265-296</page-range></nlm-citation>
</ref>
<ref id="B26">
<label>26</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Sigel]]></surname>
<given-names><![CDATA[I. E.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[How intelligence tests limit understanding of intelligence]]></article-title>
<source><![CDATA[Merrill-Paker Quarterly]]></source>
<year>1963</year>
<volume>9</volume>
<page-range>39-56</page-range></nlm-citation>
</ref>
<ref id="B27">
<label>27</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Suh]]></surname>
<given-names><![CDATA[Y]]></given-names>
</name>
<name>
<surname><![CDATA[Bolt]]></surname>
<given-names><![CDATA[D. M.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Nested logit models for multiple-choice item response data]]></article-title>
<source><![CDATA[Psychometrika]]></source>
<year>2010</year>
<volume>75</volume>
<page-range>454-473</page-range></nlm-citation>
</ref>
<ref id="B28">
<label>28</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Suh]]></surname>
<given-names><![CDATA[Y]]></given-names>
</name>
<name>
<surname><![CDATA[Bolt]]></surname>
<given-names><![CDATA[D. M.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[A Nested Logit Approach for Investigating Distractors as Cause of Different Item Functioning]]></article-title>
<source><![CDATA[Journal of Educational Measurement]]></source>
<year>2011</year>
<volume>48</volume>
<page-range>188-205</page-range></nlm-citation>
</ref>
<ref id="B29">
<label>29</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Suh]]></surname>
<given-names><![CDATA[Y.]]></given-names>
</name>
<name>
<surname><![CDATA[Talley]]></surname>
<given-names><![CDATA[A. E.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[An Empirical Comparison of DDF Detection Methods for Understanding the Causes of DIF in Multiple-Choice Items]]></article-title>
<source><![CDATA[Applied Measurement in Education]]></source>
<year>2015</year>
<volume>28</volume>
<page-range>48-67</page-range></nlm-citation>
</ref>
<ref id="B30">
<label>30</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Tamir]]></surname>
<given-names><![CDATA[P]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[An alternative approach to the construction of multiple choice test items]]></article-title>
<source><![CDATA[Journal of Biological Education]]></source>
<year>1971</year>
<volume>5</volume>
<page-range>305-307</page-range></nlm-citation>
</ref>
<ref id="B31">
<label>31</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Tamir]]></surname>
<given-names><![CDATA[P]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Some issues related to the use of justifications to multiple-choice answers]]></article-title>
<source><![CDATA[Journal of Biological Education]]></source>
<year>1989</year>
<volume>23</volume>
<page-range>285-292</page-range></nlm-citation>
</ref>
<ref id="B32">
<label>32</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Thissen]]></surname>
<given-names><![CDATA[D. M.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Information in wrong responses to the Raven Progressive Matrices]]></article-title>
<source><![CDATA[Journal of Educational Measurement]]></source>
<year>1976</year>
<volume>13</volume>
<page-range>201-214</page-range></nlm-citation>
</ref>
<ref id="B33">
<label>33</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Thissen]]></surname>
<given-names><![CDATA[D.]]></given-names>
</name>
<name>
<surname><![CDATA[Steinberg]]></surname>
<given-names><![CDATA[L.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[A Response Model for Multiple Choice Items]]></article-title>
<source><![CDATA[Psychometrika]]></source>
<year>1984</year>
<volume>49</volume>
<page-range>501-519</page-range></nlm-citation>
</ref>
<ref id="B34">
<label>34</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Thissen]]></surname>
<given-names><![CDATA[D.]]></given-names>
</name>
<name>
<surname><![CDATA[Steinberg]]></surname>
<given-names><![CDATA[L.]]></given-names>
</name>
<name>
<surname><![CDATA[Fitzpatrick]]></surname>
<given-names><![CDATA[A. R.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[Multiple-Choice Models: The Distractors Are Also Part of the Item]]></article-title>
<source><![CDATA[Journal of Educational Measurement]]></source>
<year>1989</year>
<volume>26</volume>
<page-range>161-176</page-range></nlm-citation>
</ref>
<ref id="B35">
<label>35</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Walther]]></surname>
<given-names><![CDATA[B. A.]]></given-names>
</name>
<name>
<surname><![CDATA[Moore]]></surname>
<given-names><![CDATA[J. L.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[The concepts of bias, precision and accuracy, and their use in testing the performance of species richness estimators, with a literature review of estimator performance]]></article-title>
<source><![CDATA[Ecography]]></source>
<year>2005</year>
<volume>28</volume>
<page-range>815-829</page-range></nlm-citation>
</ref>
<ref id="B36">
<label>36</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Wollack]]></surname>
<given-names><![CDATA[J. A.]]></given-names>
</name>
</person-group>
<article-title xml:lang="en"><![CDATA[A Nominal Response Model Approach for Detecting Answer Copying]]></article-title>
<source><![CDATA[Applied Psychological Measurement]]></source>
<year>1997</year>
<volume>21</volume>
<page-range>307-320</page-range></nlm-citation>
</ref>
</ref-list>
</back>
</article>
