跳到主要內容

Doctoral thesis: Task Difficulty in Semi-direct Speaking Tests

Abstract

This dissertation discusses the effect of intra-task variations in association with performance conditions on task difficulty in a semi-direct speaking test, namely, the GEPT-I. The variables identified for investigation are linguistic demand of task input in relation to code complexity; amount of time allowed for performance in relation to communicative demand; type of pre-task planning in relation to communicative demand; and test-takers’ familiarity with the non-verbal propositional content of the task input in relation to cognitive complexity.

The dissertation takes a multi-dimensional approach in measuring the effects of the variables in a comparison of data collected from the performances in a controlled task and in an experimental task. The dimensions of the comparison are based on three different sources of data, including task scores, test-takers’ responses to the post-task questionnaires eliciting views of difficulty, and interlanguage measures in the areas of accuracy, fluency, complexity, and lexical density. In addition, learners’ proficiency in English is treated as a moderator variable in order to investigate the extent to which learners’ English proficiency interacts with the effects.

A study to establish the parallel nature of the tasks to be used in the study was first carried out to demonstrate quantitatively and qualitatively that the controlled tasks and the experimental tasks were equivalent before the manipulations of the variables were made. 239 Taiwanese learners participated in the main studies and their performance was analysed. The results indicate that the variation in performance conditions altered the degree of difficulty as measured in part or all of the dimensions of the comparisons. The results also confirm that the effect was altered in some cases due to learners’ proficiency in English; in general the more proficient learners were found to react to the variable more strongly than the less proficient learners.

Significant implications for language testing, learning and research in general are suggested. In particular, recommendations in relation to reliability (i.e., for parallel tasks/forms and inter-/intra-raters) and content validity are made to exam boards.

留言

這個網誌中的熱門文章

何謂測驗信度、效度? What is Test Reliability and Validity?

原文刊載於 中華民國 91 年 3 月 23 日 《 中央日報.英語教與學 》 國內近來掀起了一股英語能力檢測的熱潮,各種英語測驗紛紛出籠,如全民英檢、托福、愛普、多益、劍橋認證等。這些測驗各自標榜特色,有的強調測驗簡單易考、有的則是強調含聽、讀、說、寫四項測驗的全方位英語能力評量,讓大家真是眼花撩亂。而對有興趣要報考英語測驗的人來說,更是難以做選擇。其實選擇英語測驗就像選購商品一樣,除了功能、價格等因素外,最重要的就是品質了。賣毛衣的商家為了說明商品的品質良好,會說所賣的毛衣絕對是純羊毛做的,而多數人大概也知道如何判斷羊毛衣的真假。相對於賣毛衣的商家,辦理英語測驗的機構會以測驗具有「信度」與「效度」來說明測驗的品質,專業的測驗機構甚至會提出些數據加以補充說明。但是很多人卻連「信度」與「效度」這兩個名詞的意義是甚麼都弄不清楚,更別提判斷其所言的真假了。 事實上,「信度」與「效度」是測驗理論的術語,一般人較感陌生。作者從事「全民英檢」的研發工作,認為測驗單位應有責任提供測驗的使用者( stakeholders )有關「信度」與「效度」的資訊,協助大家了解,以便在選擇採用英語測驗時做出正確的判斷。 壹、「信度」( reliability ) 信度是指測驗分數可靠的程度,也就是這一測驗受信賴的程度。而一測驗為甚麼會受到信賴,關鍵在於結果的一致。同一位考生在能力沒有變化的前提下,在不同時間或不同的測試狀況下重複受測,其所得的分數應該是一致的,否則就產生測驗的誤差。一個測驗有誤差是難免的,但是當誤差過大時就影響了測驗的公平性了。測驗理論上有一個基本假定:實得分數等於真實分數加上誤差( X=T+E ),但是真實分數是一個未知數。例如甲生考了某英語閱讀測驗的 A 卷得到 80 分,一天後考了 B 卷得了 82 分。雖然有 2 分的差距,但這是可被接受的誤差值,顯示該測驗結果的一致性頗高。又例如乙生考了某寫作測驗,閱卷老師 A 給 60 分,閱卷老師 B 給 62 分,這個結果顯示評分標準相當一致,而測驗的信度自然就高。總之,實得分數與真實分數愈接近即表示誤差愈小,測驗的分數就愈能代表考生的能力,如此,測驗的可信度也就愈高。 貳、「效度」( validity ) 測驗效度即指測驗分數的正確性,簡單的說,就是指一測驗是否評量到它所要評量的...

「全民英檢」高級測驗-口說能力測驗

原文刊載於中華民國 91 年 4 月 《 中央日報.英語教與學 》 繼「全民英檢」高級閱讀能力測驗及寫作能力測驗之後,本文將為讀者介紹口說能力測驗。以下分別就測驗題型、評分重點及應考準備方向做說明。 測驗題型 「全民英檢」研究委員會訂定之高級口說能力測驗目標為「英語流利順暢,僅有少許錯誤,」應用能力擴及學術或專業領域」,依據該目標,高級口說能力測驗設計為三個部分,分別是「暖身面談」( Warm-up Interview )、「訊息交換」( Information Exchange )及「申述」( Presentation )。 不同於目前「全民英檢」其他級數的口說能力測驗,高級口說能力測驗採考生與主考 ( Interlocutor ) 面對面的施測方式( OPI, Oral Proficiency Interview ),每場測驗均有兩至三名考生參加。測驗全程錄音、錄影, Interlocutor 負責向考生提問並依據整體式評分量表給分,試場內的另一主考( Assessor )則負責依據分項式評分量表給分。分項式評分與整體式評分的優缺點在文獻上多有論述( Hughes, 1989; Bachman & Palmer, 1996 )。 高級口說能力測驗採整體式及分項式評分並行制,可達到兩種評分方式相互驗證的目的( Hughes, 1989 )同時提高評分者信 度。 兩項量表計分由低至高分為 l~5 五個級分, 3 級分為通過標準。下文即針對高級口說能力測驗之內容做進一步說明。 * 第一部分「暖身面談」:為考生與主考之間的交談,以問答的方式進行,約五分鐘,主要評量考生自我介紹及回答問題之口語能力。 * 第二部分「訊息交換」:包括考生相互之間訊息交換、討論及回答主考提問等,約七分鐘,主要評量考生口語互動與討論之能 力。 * 第三部分「申述」:考生依主考所提問題思考兩分鐘後發表,另一考生則對該生之意見發表作口頭摘要,歷時約十分鐘。此主要評量考生對特定主題作較深入表述及在短時間內作口頭摘要的能力。 評分重點 高級口說能力測驗之評分重點包含發音、語調( Pronunciation & Intonation )、切題度( Relevance & Adequacy )、語彙使...

「全民英檢」寫作測驗之評分標準與程序 GEPT Writing Tests: Rating Criteria and Process

原文刊載於中華民國 91 年 2 月 10 日 《 中央日報.英語教與學 》 「全民英檢」的初、中、中高級均含寫作測驗,主要目的是評量考生的文字表達能力,也就是語言的使用能力。有別於聽力、閱讀能力測驗之使用客觀題、採電腦閱卷,寫作測驗則是主觀題,需要由專業的評分老師做人工評分。既然是人工評分,則難免因評分老師的個人主觀判斷或個人因素(如疲倦)影響評分。要使寫作評分能正確的反應考生的真實寫作能力,如果排除考生本身的因素,則命題與評分是最關鍵的兩個因素。 為提高評分的一致性( inter-rater consistency ),「全民英檢」的寫作測驗題型不是「自由寫作」( free writing ),而是「引導寫作」( guided writing ),利用圖片、大綱等提示明確的要求考生寫作的內容。這種「引導寫作」的測驗方式有助於降低評分老師的主觀判斷。然而對寫作評分影響最大的還是評分過程。不同的評分老師可能閱了同一篇作文而給了不同的分數,因此如何建立評分者之間的一致性( inter-rater consistency ),也就是評分的信度( reliability )是非常重要的。評分的信度越高,(信度越接近 1 )表示評分者之間的給分標準趨於一致,評分越可靠。一般而言,信度達 0.85 以上時,就表示評分相當可靠。 「全民英檢」一向重視閱卷信度的確保,採取質量並重的控管措施盡量減低評分者的評分誤差。開辦兩年以來,寫作測驗與口說能力測驗的評分信度均保持在 0.86-0.90 之間,達到不錯的水準。這個數值與大家所熟知的「托福」寫作測驗( TWE—Test of Written English )、口說能力測驗( TSE—Test of Spoken English ) 0.87-0.90 的信度相當。我們是怎麼辦到的?本文特別針對「全民英檢」的寫作測驗評分程序提出說明(口說能力測驗的評分程序與寫作測驗類似,故不重複),希望有助於外界對「全民英檢」的認識。 一、「全民英檢」各級寫作測驗均訂有評分指標( 0-5 級分),評分人員在確切掌握評分指標後,依據考生的整體表現評分。每一篇作文皆由兩位評分老師分別獨立評分,若兩者評分差距在 1 級分以內,求其平均值;兩者評分差距大於 l 級分以上,則由第三位(資深)評分老師複閱,並以其評分為最...