Transcription
Thank you, Nancy. And hello, everyone. I'm Evelina and this is Nahal. Hi, everyone. We work in the Research and Thought Leadership Group at Cambridge English. Welcome to this webinar on understanding assessment: what every teacher should know.
What we'll try to do today is help you understand some key ideas about assessment through simple explanations and some examples as well. So, our aim today is, first of all, to help you understand key assessment concepts. Secondly, we hope the webinar will help you to develop high-quality classroom tests which help with learning. And also, we hope that after today's webinar you will be able to evaluate existing tests and decide if they're appropriate for your learners.
But first of all, let's start with an activity for you. We're going to give you two tasks, which are used in speaking exams. And after you look at those two tasks, we will ask you which task, in your opinion is better at testing speaking skills?
So, here's the first example on your screen. It is a reading aloud task from a business exam. Take a look at the task. It involves six sentences which test takers have to read aloud. Okay, this is the first task.
Now, let's take a look at the second task that is now on the screen for you. Now, this task is different. It involves talking about information in pictures. Again, similar to the previous task think about the question: Does this task test speaking?
Now, let's take a vote on your ideas. Please use the voting buttons to tell us which task, in your opinion is better at testing speaking skills? Is it task one, which is the reading aloud task or is it task two which is the task that involves talking about information in pictures? So, please use the voting buttons, the two options, one and two, and not the text chat. And the options are once again, task one, reading aloud and task two, talking about information in pictures. So, let's see what you think. Which of these is better at testing speaking?
So, we have actually quite a strong majority here. What I'm seeing on the poll is that 99% are saying that task B, talking about information in pictures is better testing speaking. One of you has said "reading aloud" and that's interesting actually, we'll come back to the idea of the usefulness of these two tasks. But I agree I would say that task two is probably better at testing speaking skills.
Now, here's another question for you. How do you know what made one of these tasks more suitable for testing speaking than the other task? Now, this time, type your ideas into the text chat now. So tell us, what do you think made one task more suitable than the other task? Type your ideas into the chat box and let's see what you're telling us.
So, what made task two better than task one at testing speaking? So, one person here is saying that it is spontaneous production. This is my tip. She's saying there's spontaneous production. I agree. By this I think you mean that you have to produce some language instead of reading language which is given to you and... another person is saying creativity is involved. I think that, again, gets to the idea that you have to create language versus just reading something on the page. Another person is saying that there is more fluency involved in the second task. So, again, the idea of having to produce some speaking yourself. I see that someone else is also saying that the first task involves reading much more than speaking, so I agree with this.
Just to summarise what you said, I would say that the first task test reading which is one of the points that you made and not really speaking. The first task simply involves also producing words but the second task requires turning ideas into language. So, that was the idea of spontaneous production that some of you mentioned. And also, there is a very narrow idea of speaking in the first task whereas there's a broader definition of speaking the second task. So, you gave us some very, very good ideas that we agree with.
Now, when evaluating the previous two tasks in terms of how appropriate they were for testing speaking you're basically putting assessment literacy into practice. Assessment literacy means understanding assessment and knowing what assessment is doing. It means that if you evaluate a test or you try to develop a test you understand and you can apply basic assessment principles.
So, why is it that assessment literacy is so important? Before we answer that, let's take a look at what students and teachers say about assessment, and then we'll answer that question again.
Students might say something like this, "When you do some exams it is useful because it tells you how well you're doing in English even though I don't actually like the test so much." In other words, learners feel the tests are useful as a tool for helping them understand their progress in English. But at the same time, they also tend to dislike the tests because of pressure anxiety and things like competition.
Now, turning to teachers, a typical comment you might hear often is something like "The best test should mirror what happens in the classroom. But sometimes instead of focusing on improving learning I'm actually focusing on improving test scores." The teacher's comment illustrates this tension between the role of tests on the one hand, and motivating students perhaps to study harder but also highlights the fact that often the test might not actually be connected to what is happening in the classroom. For example, when you use a multiple choice grammar test to assess writing skills. So in a way, students and teachers have some sort of a love-hate relationship with tests. In other words, while they each can see a value in tests they also have criticisms of them.
Now, what does this have to do with assessment literacy? Well, assessment literacy, which is the knowledge about good practice and assessment helps teachers to develop tests which support learning. It also helps them evaluate and select tests which have a positive impact on learning. So, when a teacher has knowledge of assessment this will help their students because high-quality tests actually assist learning. So basically, what we're going to do in this webinar is that we'll help you to develop assessment literacy.
Now let's move on to some key conceptual assessment. A general overarching concept in assessment is Validity. Validity refers to whether the test is measuring what it aims to be measuring. For example, if we think of a driving test a driving test which claims to validity must include a practical driving component and not just theoretical knowledge of the rules of driving. Now, let's think of a language test, a language test for university entry. A test such as that must include several components and one of them should be the ability to write essays because writing essays is a very important component of academic English at university.
Validity is a broad concept which has different elements and we've given you those elements on the right, on the slide and they are: Test Purpose, Test Takers, Test Construct, Test Tasks Test Reliability and Test Impact. What we're going to do now is look at each one of these key concepts of validity in turn, and we're going to do it through a set of questions for you.
The first question relates to why? Why am I testing? And here are some possible reasons that we've given you. One reason is to check learning at the end of a unit at school. Another reason for testing is to diagnose what learners know and what they don't know. Another reason is to place learners into groups based on their ability. Another reason could be to provide test takers with a certificate of language proficiency. Now, please use your buttons to select which of these four purposes for testing apply to you: A, B, C, or D. It is obviously possible that you are familiar with different purposes for testing but use the one that is most familiar to your situation. So, what is the main reason you test? What is the test purpose that applies to you the most? Please tell us by selecting A, B, C or D.
Well, the poll is still happening as you're putting your votes in but what I can see is that the vast majority of you are choosing A or B. A is to check learning at the end of a unit in school and B is to diagnose what learners know and don't know. And this actually makes sense in a classroom setting what teachers do is often test for these two purposes to check learning and to diagnose.
Okay, now each of these reasons for testing represents a different test purpose and test purpose is a fundamental concept in any assessment. It is fundamental because the purpose of the test determines the type of test you're going to produce and that means the kinds of tasks you're going to choose, and the test items the length of the test. Now, imagine, for example that you need to produce a placement test for medical doctors. Such a test might be used to place them into a language course. But in contrast, a certificated proficiency test for doctors may be used to decide if they're able to start practising medicine in an English-speaking country. And the important point here is that the different test purpose would result in a different kind of test.
Now, the second key question that we'd like to focus on is, who am I testing? Is it, for example, primary school children? Is it teenagers? Is it adults? Is it airline pilots or doctors and so on and so forth? So the key term here is therefore Test Takers. The reason why this is so important is because the test has to be appropriate for the test takers it is aimed for, for example if our test takers are primary school children we might want to give them more interactive tasks or games to test their language ability. But we might not necessarily give such tasks to adult learners or we might use role plays with doctors when testing their listening skills. But we might use more lectures and monologues with students at university in order to make the tasks more relevant to our test takers.
Now, let's look at the third question, and that question is: What am I testing? For example, am I testing communicative language ability or am I testing something narrower, such as speaking ability or maybe even narrower such as pronunciation? Or am I testing grammatical knowledge or the use of that grammatical knowledge? The key concept here is Test Construct and this is one of the most fundamental concepts in assessment.
Now, let's take a look at it in a little bit more detail. A construct is an ability or a skill. In technical language we refer to this as a latent trait. "Latent" means something which is not easily observable. So, cognitive ability in your brain, and "trait" means an ability or a skill. Now, some examples of constructs are maths. Knowledge of mathematics is a construct. Intelligence is another construct. Personality is a construct. Anxiety is a construct. English language ability is a construct. Pronunciation is another example of a construct. So, there are different ways that a construct may be measured. For example, if we wanted to test personality we might use a multiple choice questionnaire or we might use observations of somebody. When we want to measure anxiety we might give questionnaires, or we might measure someone's pulse rate. But even if we don't use formal tests for some constructs we still have to understand how we can measure those constructs.
So, constructs as Evelina said, are fundamental to language testing and the key question is, what is our construct? Another key question is, how are we actually going to go about testing that construct? For example, when interested in grammar and vocabulary are we going to use multiple choice items or other types of items and tasks? With reading, are we going to use one text followed by questions or are we going to use several texts? With a listening exam are we going to use a lecture or are we going to use a series of short conversation followed by some comprehension questions? When assessing writing, what exactly are we going to ask our test takers to write? Finally, with assessing and speaking are we going to use read-aloud tasks as the ones that we showed you earlier at the beginning of the webinar or are we going to engage our learners in face-to-face interaction? These are only a few examples of ways we could decide to test a specific skill. The focus here is on test tasks. The test tasks are the way in which we elicit and measure the construct. That is the ability we're interested in testing. The test tasks are like a menu of options which is available for us to choose from, but we must be sure to choose the right task or the right range of tasks for the kind of ability that we'd like to measure.
Moving on to the next question and that question is related to the scoring of the language produced in our test. And the question is: How am I scoring the test? For example are the answers to the task I've developed going to be scored as correct or incorrect? Now, this might be the case for a multiple-choice task test, for example. Or am I going to use trained teachers to score some of the tasks? This may be the case for scoring speaking or listening, for example. Or am I going to use specific criteria to assess what the language produced? For example, grammar, vocabulary, pronunciation essay organisation in writing and so on. These are all questions that you need to address, and they relate to reliability. In other words, they relate to how dependable the scores from the test are. Or to put it differently, how can we make sure that the scores that we give in a test, reflect the learner's actual ability and not whether, for example the examiner happened to be in a bad mood that day and was particularly harsh in giving marks.
Now, the final key question that we'd like to talk about is related to the learning value of the test, that is, how is my test benefiting my learners? Is it supporting learning through, for example the use of authentic tasks, and engaging learners in situations which are similar to the ones that they face outside of the classroom? Or is the test benefiting learners through, for example including all four skills such as reading, writing, listening and speaking Because language proficiency includes all of these different components? Or, for example, is it through providing feedback to the learners based on their performance on the test? Now, all of these different factors would help increase the learning value of the test, and they relate to the idea of test impact. In other words the effect the test has on learning.
Now, we have gone through the six key questions and concepts in language assessment let's quickly go over these with a task for you. We'll give you a few assessment situations and we would like you to tell us which concept they relate to. For example, when you use diagnostic tests we're referring to the concept of Test Purpose. Now, what about when we talk about the role the test has in influencing what tasks teachers will use in the classroom? Which assessment concept does this refer to? Once again, the role the test has in influencing what tasks teachers will use in the classroom. Please vote for the correct answer.
I can see answers coming in. Yes, we're having different responses. Some of you are saying Test Purpose, some of you are saying Test Impact. Actually, majority of you are now saying Test Impact. Very different answers. I'm still waiting for the poll to end. Okay. So, interesting, a lot of you are saying either Test Purpose or you're saying Test Impact. The correct answer is actually Test Impact because we were looking at the effect of the test on the classroom.
What about when we try to make scores more dependable by using assessment scales? What concept does this refer to? Please choose the correct answer. Once again, what about when we try to make scores more dependable by using assessment scales? What concept do you think this refers to? Please choose from the six available options. Okay. Still waiting for the answers to come in. I can see the majority of you are going for option E, which is Test Reliability. Some of you have also chosen Test Purpose and Test Impact. Okay, giving you a few more minutes for everybody's answers to come in. And yes, I can say that the majority of you, 65% have chosen Test Reliability and that would be the correct answer.
Now, as we said earlier, these six questions all relate to the concept of Validity. As a reminder, Validity refers to what the test claims to be measuring. Validity, however, does not exist in a vacuum. We can never really say that a test is valid or not valid. Instead, we can say that a test is valid for a particular purpose or it is not valid for a particular purpose. To go back to the two speaking tasks that we started with at the beginning of the seminar the one which involved reading aloud is quite a limited task as far as speaking is concerned, but in some cases it may be a valid task to use. For example, you may want to test reading fluency or you may want to test pronunciation. We can therefore argue that this task is valid for that particular purpose. But if you were interested in testing interaction skills this task would not actually be fit for purpose and a face-to-face speaking test would be more appropriate. So, we always need to consider the fitness for purpose of a task or a test.
So far, we've looked at the six key questions. Now, in the rest of the webinar, we would like to focus in a bit more detail on four of these concepts. These are: Test Construct, Test Tasks, Test Reliability, and Test Impact.
Now, let's start with the Test Construct. As a reminder, Test Construct refers to the sometimes hidden ability we're trying to measure. Now, a construct has two elements a cognitive element and a task element. In other words, we can make the cognitive ability observable through a task. Now, the point here is that we don't just randomly put any tasks in a test but we put tasks in a test because we want them to activate certain cognitive processes. For example, in a face-to-face speaking test a learner has to both generate ideas and activate their grammatical, lexical knowledge and their pronunciation competence. They also have to pay attention to what the other person is saying and they have to adapt their speech to what the other person has said. These all refer to different cognitive processes in speaking. And so in a test, a key question is, what are the cognitive processes required to complete the task?
Now, let's take a look at an example of cognitive processes in practice. We've chosen an example from a reading task. We're going to give you one sentence to read and we'll ask you three questions about that sentence. Now, the sentence is on your screen. Please now, read the sentence and answer question one. Now, the sentence is not in English but you should be able to answer it by using your knowledge of English. So, please just read it through quickly, and see if you can answer question one. What was the flester doing? Is it A, chandering; B, gollining: C, rangeling? So, what was the flester doing? A, B and C, choose your answers to that question.
Well, the poll is still open but there's actually a very strong majority here and over 90% of you have chosen B, which is "gollining". A few people have chosen the other options but quite a large majority have chosen B, and B is actually the correct answer.
Okay, now we're going to stay with the same sentence but here's another question for you. Where was the flester? Is it A, begrunt the quistly; B, besand chander; C, begrunt the bruck? Again, read the sentence, and choose A, B, or C.
Well, the poll is still open and your answers are coming through so we'll just wait a few more seconds before we see where your answers fall. Quite a lot of you are going for one of those answers, actually. And in fact, the vast majority in this case, 91% have chosen C, begrunt the bruck. And that is actually the correct answer.
Now, let's take a look at one final question. That's question three, about the same sentence. Again, choose the correct answer. And the question is: What is the event described here? So, again, choose A, B or C. Is it A: The flester is participating in a sports competition? Is it B: The flester is cooking a special dinner for friends? Or is it C: The flester is ill in hospital? So, choose A, B or C What is the event described here?
So, your answers are still coming in. Let's see how they shape up. Interesting that a lot of you are actually going for one of those options. Okay, just a few more seconds for the poll to close. And what we're finding is that 80% went for A 11% chose B, and 8% chose C. Now, even though there's quite a strong majority here it's not as strong as with the previous questions. And more of you have gone for any of these options.
Now, let's think about the questions we asked you. What actually makes questions one and two different from question three? What we'd like to ask you is to share your ideas in the chat box. Think about things like the sentence structures between the different questions and how hard it was to pick up the gist of each question. What do you think actually makes question one and two different from question three?
Okay, Mehmet is saying the third one is a guess, yeah. Questions one and two are easier. Some are saying questions one and two are more specific Alex is saying it's about grammar. The third one is about context. Yes, you're right, actually. So, to summarise you're mentioning ideas which are related to, for example finding clues in the sentences or guessing meaning from the Lexis and grammar in the sentences in the first two questions. Whereas, in the third question you need a bit more information than what is provided in the actual sentence. And yes, these are very good suggestions and we would agree with you. To put this in another way questions one and two are tapping into sentence-level knowledge of a language. But question three goes beyond the sentence to inferential skills. The first two questions are asking the learners to read the lines. Whereas the third question asked them to read beyond the lines or between the lines.
Now, let's take a look at this in a bit more detail. Here's an overview of some of the cognitive processes, which reading involves when reading a text carefully. There are generally two levels of comprehending the text. There's a local level, and there's a global level. At the local level we cognitively process the Lexis, grammar and syntactic form and meaning of the sentences. However, at the global level we go beyond the meaning of individual sentences to comprehend main ideas. We make inferences and we try to understand the meaning, which is created by the sentences together in the whole text. So, going back to the three questions, which you had to answer questions one and two tapped into reading at the local level whereas, question three was tapping into reading at the global level. In other words, these questions activated different types of cognitive processes in reading. A key point here is that neither one of these questions is good or bad. What is important is whether the questions in a test trigger appropriate cognitive processes. For example, when developing the test for beginner learners it may be appropriate to have more questions at the local level. This is because beginner learners are still mastering basic linguistic knowledge which is related to vocabulary, grammar, and syntax. However, in a test for intermediate or advanced learners of English you may need to include questions which focus both on the local and a global understanding of the text.
So far, we've talked about the cognitive aspect of a test construct. Now, let's take a look at the task element of the construct. And the key question we have to ask ourselves here is are the tasks appropriate for the test construct? In other words do the selected tasks in our test, activate appropriate cognitive processes? There are many different types of tasks and we've shown you a selection of task types in this word cloud on the screen. For example, we've included task types such as discrete-point tasks. Discrete means they're separate independent questions. Integrated tasks multiple-choice tasks and so on. We'd like you to ask you to quickly type in the task types that you're most familiar with. So just look at all the tasks on the slide and tell us which three you use most often. Which task types do you use most often?
Let's see what you're saying in terms of task types. Okay, still waiting for your ideas to start coming in. So, which task types do you use most often in the classroom? I see that Zainal is saying gut feeling. There's another idea for role-play. Julie is saying gut feeling as well. Multiple-choice is coming up. True-false is coming up as well. Integrated tasks, someone is typing in. Short answer tasks is another option. Gap-filling, again. Cloze tests are coming up. Multiple-choice. Okay, so a lot of ideas. We don't have time to look at all of them but I can see that many of you have used multiple-choice and true-false tasks and quite a few of you are also familiar with integrated tasks as well, and role-play.
Now, we don't have time today to go through all of these task types but what is really important for you to keep in mind is that no task type is naturally good or bad. All tasks have their strengths and they all have their weaknesses. And the important point is that we go back to again and again in this webinar is the idea of fitness for purpose of a test or a task. So, in a test, it's useful to include a range of task types because that way you're minimising the potential problems connected with certain task types.
Now, for example, we can look at the multiple-choice task which is often used in tests and is one that came up over and over again you told us that you use very often in your classrooms. The test taker is usually given three or four options and has to choose the correct one. Now, we've given you an example here. We've chosen it from a listening test where test takers listen to a short conversation between two people and then they have to respond to a multiple-choice comprehension question. So, we have a question with three options at the top and the conversation the test takers listen to underneath.
Now, let's take a look at some of the advantages of this particular task type. One of the advantages is that you can create questions which tap into different levels of cognitive processing, and that enhances the validity of the task because it includes different parts of the construct. Another advantage is that the task is fairly easy to market and the marking can also be successfully done by machines so, it's a very practical option. A potential problem, however with this task type is that there's a chance of getting the correct answer through guessing. A further problem is that answering multiple-choice questions is not necessarily an authentic task, since real life and outside of the classroom we don't actually go around choosing multiple-choice options. Also, research has shown that multiple-choice tasks may favour boys because they're more likely to be risk-takers. So it would be a good idea not to have all items as multiple-choice but rather to balance out some of the weaknesses that we talked about with strength from other task types.
Now, here's one more example, this time from a speaking test. This task type involves a task in which two or three test takers are given some problems and they have to discuss these problems together without the guidance of an examiner or a teacher. The example we've included here asks is it a good idea for students to go on school trips? And it includes possible reasons or ideas that learners can talk about such as getting on with people, help with lessons learning about the world and so on. Now, an advantage of this task type is its high authenticity. In other words that test replicates skills which learners have to use outside of the test or classroom, that is interaction with their peers. Another related advantage of this type of task is that it supports the development of interactional skills. A drawback of the task, however is the fact that performance of the learners could be affected by their partner. For example, the personality of one of the learners in the pair or the group could affect performance on the task. Some might be introverted, some might be extroverted. The age of the learners could also play a role or their gender but also their language ability whether they're a match with somebody from higher ability or lower ability. And these are all factors to consider.
Yes, thank you, Nahal. So, once again a reminder that all types of tasks come with advantages and they come with limitations. It's important, first of all, to be aware of these advantages and limitations. Secondly, it's important to use a range of task types in order to minimise the limitations of tasks and to ensure fitness for purpose.
We've looked at a few examples of tasks. Now we'd like to give you a task to look at. Actually, we're going to give you two tasks taken from a writing test and we'd like you to decide which one of these tasks is better for a writing test. Is it task A or task B? So the two tasks are on your screen. The two examples, this task A, task B. They're writing tasks. Please tell us which of these tasks do you think is better to include in the writing test? Choose A or B in the poll.
So we're still waiting for your answers to come in. As you're deciding which task is better to include in the writing test. Is it A, which is on the left of your screen, or is it B which is on the right of your screen? So, your answers are still coming in. We'll just wait a few seconds before we comment on it. And I can see actually before the polls closed that is actually quite a strong majority. Ninety-three percent of you so over 90% are saying that task B is actually better than task A in a writing test.
Now, let's think about why. Why did you choose task B and not task A this time? Please give us some of your suggestions in the chat box. So, just type a few ideas, a few suggestions why you think that task B is better than task A to include in the writing test. What makes it better?
So, I see that Georgina is saying that task B tries to contextualise the issue. So what Georgina means, I think is that task B creates more of a context around so it gives more information about the situation, and I agree with that. Someone else says task B is more specific. Again, I agree with that. I think... not I think, but it relates to the idea of providing more information so the test takers know what is expected of them. Someone else said that task B is based on communication. I think what they're trying to say here is the idea of authenticity. Communicating in this way is much more knowing exactly what we're doing knowing the purpose of the communication is much more authentic in terms of writing than just write about something, so I agree with that. Debbie says task B gives concrete ideas. Again, I agree that makes the task better because it provides more specific concrete ideas for the test takers.
Okay, let's just summarise quickly what you've told us. And just one more idea here before we move on, the idea of instructions. So task B, you're saying provides clearer instructions than task A and I agree with that. Also to repeat what some of you have said that task B provides more context which tells you who the reader is and that's much more authentic. It relates to what we do in real life. We know the context of the writing we're doing. Task B also, I would add, is fairer because it gives ideas about what to write so that the test taker is not spending time generating the ideas themselves. And this goes back to the point about knowing what we're measuring. In other words, knowing what the construct is. In this case, we're testing writing skills. We're not testing thinking skills or creativity. So that's why giving ideas makes the task fairer because learners focus just on producing language. And also, I would add that task B is more reliable because it asks each test to write about the same information. And so training readers would be more focused because you know what to expect you know what the responses will be. And finally, task A, just to summarise is too open and that makes it a less suitable task for writing.
So far, we've spent some time looking at issues related to the construct of the test and how to test that construct through the most appropriate tasks. Now, let's move on to the idea of reliability, which we looked at briefly earlier. As a reminder, reliability refers to how far we can depend on the scores from the test. In other words, how consistent and accurate are the scores from the test?
Now, I'd like to have another activity for you. On your screens, you will see an apple. Now, what I'd like you to do is to rate this apple on a scale of one to six where one is the lowest and six is the highest score. Again, you have a picture of an apple on your screen. Please give it a score from one to six.
Okay, waiting for your answers to come through. Interesting, I'm getting twos, threes, fours, five,s and some sixes. I haven't got one yet. It is a nice looking apple. It is a nice looking apple, it's true! So, there were quite a few answers you had all chosen all the way from two to six. And because it was a nice apple, we decided that there is no one.
Now, what I'd like you to do is narrow down the task. This time, please rate the apple on a scale of one to six but this time for quality of colour. Again, please rate the apple from one to six, but this time for quality of colour.
Okay, again, very interesting. Again, we're having all sorts of different answers. This time, some of you chose one. Some of you are choosing six. Quite a vast majority are between three and four. So we have about 30% giving a three, and 30% giving a four. Well, there is a bit more agreement amongst you. There is still some disagreement. What we like to argue is that if we give you even more detailed criteria for judging the apple, and we give you some more training you will probably have more agreement amongst you. However, we will guarantee that you will still not all fully agree on a score.
So, what does this tell us about scoring language? If we cannot all agree on something as simple as the colour of an apple how can we agree on something as complex as language? The lesson here is that human raters are bound to disagree on scores but what is important is to aim for an adequate degree of agreement rather than perfect agreement. Humans, after all, are not machines, and it's important to remember that. Indeed, yes.
So now, test reliability is important, as Nahal said. Let's take a look at a few fundamental ways of increasing the reliability of your test. One way is through providing clear item or task instructions. Because if the instructions are clear, the task will produce the kind of language you expect, you need to make it easier to mark. So, the markers will be more consistent in marking because they know what to expect. Now, clear assessment criteria is another way of increasing reliability because if your raters know what they need to focus on when marking then they'll have high levels of agreement. And finally, we've given you one more way of increasing reliability and that's training examiners or teachers. That has a huge impact on reliability because it increases the agreement of the raters. Now, in the case of speaking you have to train your raters or teachers also to deliver the test in a consistent way, and not just to mark it in a consistent way.
Now, we mentioned the need for clear assessment criteria and scales and we've given you an example of rating scales. In this case we've taken it from the Cambridge English First Exam which some of you are probably familiar with. And it shows you a set of analytic scales for speaking. Analytic scales means that the scales are broken down from a holistic... they're not a holistic scale but they're broken down into different criteria. In this case, the criteria are Grammar and Vocabulary, Discourse Management, Pronunciation, and Interactive Communication. And the important point about analytic scales is that learners receive a mark for each one of these criteria.
So far, and to summarise we've talked about concepts such as Validity, Fitness for purpose, and Reliability. We now turn to our last concept which we'd like to cover today. The last concept refers to Test Impact. As we discussed earlier, this relates to the effect of the test on learning. Some of you may be familiar with the term "washback" which is quite closely related to the concept of Test Impact. Washback means the direct effect of tests on classroom practices. The best way to increase test impact is to test those abilities whose development you'd like to encourage and not necessarily what is easiest to test. For example, here are some activities which take place quite regularly in a typical communicative classroom. For example, discussing in pairs or small groups describing photos and visuals, asking and answering questions, information gap activities, reading texts aloud for pronunciation practice, completing dialogues giving presentations and so on and so forth. Now, imagine that a test only includes the following three task types: reading aloud describing pictures and talking about a topic. In such cases, there is a danger that as a consequence of the test the learning which happens in the classroom will be only limited to tasks that resemble those on the test. So, a lot of the time will be spent on production of language and not necessarily on interaction.
Okay, so let's end with a summary of what we've covered today. Good tests are tests which test what they set out to test. In other words, they have construct validity. Good tests produce scores which can be trusted. In other words they have reliability. And also, good tests support learning. In other words, they have positive impact. But also, good tests need to be practical to develop and deliver. Now, imagine, for example, that you've produced a fantastic test for speaking for your school. It has fitness for purpose. It produces reliable marks. It has a positive impact on learning. At the same time, though, your perfect test requires that all the English teachers in your school have to deliver and mark the test over five full days. But they also need to teach their regular classes in that time. Now, that test is clearly not a practical test. It will be very difficult to sustain.
Before we end the webinar, here are some useful tools which will help you with your teaching and with developing tests. We've given you the icons the visual images of the tools and you can find more information on those ideas in the handout which we'll send you after the webinar. The handout is also going to include some additional tasks for you to try. We've really enjoyed sharing all of these ideas with you today and now I'm going to pass you over to Nancy, who will talk a little bit more in some ways the Cambridge English language assessment can help you. Then we'll have a Q&A session and you can ask your questions and we'll try to respond to those questions. So, do type in your questions. Thank you very much.
Thank you very much, Nancy. Now, here are some questions that you typed in during the webinar. Evelina, could you please take the first one? Okay. Some great questions. We have about five minutes to respond to some of them. And let's start with Christine, who says How can we have access to marking scales as a teacher when preparing for exams? There are examples of our marking scales, our rating scales in the handbooks. So, all of the Cambridge exams the handbooks for the Cambridge exams include examples of the speaking scales and also they give you examples of writing and what mark has been given for that writing sample. So I would say go to the Cambridge website, look at the handbooks and that will give examples of the rating scales. But to add to that, Christine, it is not just about the teacher having access to them but your learners. It is really important for your learners to be familiar with those skills as well so that they know what they'll be evaluated on. That'll be a really, really important part of their learning as well what are the criteria that they'll be evaluated against?
Actually, quite related to that quite a few of you have asked about a rating scales and whether rating scales can be more detailed. You need to remember that rating scales, usually, they have quite a few descriptors. So, you have to strike a balance between how much the rater can also cognitively process whether listening to speech and how many words and how long the descriptors are. So, it has to be a balance, I would say, in-between having really lengthy descriptors but also being able to practically use your rating scale especially if you're using it in life context as an examiner. So I would say there really has to be a balance there, if Evelina would agree? I agree, yeah, definitely.
Moving on. So, one of you, I hope I'm pronouncing the name correctly, Minwhei? I hope I haven't butchered that. And a few others, Ren, for example, as well, have asked about Baird speaking tests. And the question basically relates to, well what about the gender of the learners? Doesn't that affect the way they perform? Doesn't that affect the way they perform? What about if one is more dominating? What about the proficiency level? I agree with that. Baird speaking tests do have disadvantages. And one of them is that the background of the learners could affect the way they perform. The key, key, key variable that has the biggest impact is the language proficiency of the learner. So if one is much more advanced than the other, a Baird test doesn't work which is why we don't have Baird testing in all of the Cambridge exams but we will include Baird tests in the ones that focus on a narrow range of English proficiency, for example Cambridge First, test at the B2 level so any differences in proficiency will not play such a big role. But a test like IELTS, for example cannot use a Baird test because then differences in proficiency might be much bigger. Going back to other variables, such as, for example gender or how dominating you are, they might play a role but that makes it more authentic as well so you have to learn to deal with learners of different personalities. The key point is, though, that a Baird test includes only one task which is Baird. There are other task types in that test. So in a Baird test at Cambridge, we have a question and answer task. We have a describer picture task. We have a Baird task. We also have a discussion task with the examiner involved. So we have a range of tasks. And the point is to make the test as authentic as possible but to limit any possible problems with any of these task types by having a different range of tasks that we use.
Thank you. I have another question, very interesting from Samarrah, who is asking: In a modern test is there any room for grammar and vocabulary exercises? I would say that especially in the classroom grammar and vocabulary, you can see them as your building blocks and it's by putting them into use and in communication, that you can then start speaking, and you can start using the language. So, yes, they're important. However, we're not also advocating that you only have a test of grammar or a test of vocabulary, but rather integrate it within a communicative language test and in the same way that you might use them in your classrooms. So, is that something that you-- I agree, yeah definitely.
We have time for one more question. We have about 30 seconds. And I'm going to end with Annalisa. Annalisa says: Is it better to have a continuous assessment or just a final assessment? My very strong, definitive answer is, yes it is better to have continuous assessment to be testing throughout the period of study. Because all the learning checks and all the feedback that you give, that's crucial don't just assess, but give feedback, then helps learners to improve. But then the summative assessment should work in a complementary fashion with the formative one, so that the two are actually leading to learning and you're integrating assessment and learning continuously throughout your teaching.
Okay, we have to end. It's 11 o'clock here in the UK and this is all we have time for. Thank you very, very much for attending today. We've had about, I think, over 800 participants and it's been a true pleasure and a privilege to be able to share our ideas with so many of you. We hope you found it useful and we look forward to seeing you at our future webinars. The next webinar which is coming up is going to be introducing IELTS Life Skills. Thank you very much. Bye-bye. Thank you. Bye.