Is the purpose of the assignment to give a review and critique of the course, saying what you liked and what you didn't like? In particular, did they ask you to give feedback about their teaching methods? Honestly, it reads like an opinion piece we might read in a news paper and would benefit from spending more time on objectively talking about peer evaluations. Also, given what you said the point of the assignment was in the OP, it doesn't seem that you actually did what was asked, i.e. give a grade to each member of the group. My apologies in advance for tearing what you wrote apart...
Here I go...
You say a few times that the class was successful in providing info to you... as a psychologist, you'd have to ask yourself, "how was it successful?"..."according to what specific criteria was it successful?"..."how was this success measured exactly?" ... "what evidence is there to support this claim of success?"
XUL said: Here is an example of one category from our grading instructions: “Contribution – Did they contribute productively to team discussions and work?”
How is "contribute productively" being measured exactly?
1) what are the criteria of 'contributing productively'? 2) why are these criteria important to 'contributing productively', and thus why are they to be used to understand productive contributions? 3) how would these criteria like to have been measured? 4) what specific data/observations were actually available to make the measurements? (and what data wasn't available and why?) 5) were all the available data for the criteria of equal validity/reliability or where some more more valid/reliable than others? (and thus, was some of the data weighted differently because it was less/more valid/reliable?)
XUL said: The way we grade each other seems largely based on empirical data,
You haven't actually graded anyone yet, nor do you seem to grade anyone in this paper, so how can you say that it seems "largely based on empirical data". With your argument that it is impossible to grade anyone in the group, there is no empirical data used in that argument at all.
XUL said: such as inferring the amount of our group members’ correct answers on the i-rat, by comparing their answer to the correct answer on the t-rat – or perhaps by evaluating the information they put forth for the t-rat. As well as measuring each other on test performance, we are left to judge each other based on e-mails, amount of phone calls, tasks done, points made – all empirical data.
Here you seemed to outline some criteria of 'contributing productively'. However, it could read more clearly insofar as being specific about what the criteria are that make up your evaluation of 'contributing productively'... perhaps use dot points or number them. And then, as the reader, I'm wondering why you chose those criteria and how they were to be measured exactly. I'm also wondering why you might choose not to use any of this data (might there be something wrong with it?). I'm curious what data do you actually have for each of these criteria.
It kind of seemed like you were talking about making an evaluation only hypothetically and not actually doing it..
Quote: XUL said: Even though empirical data is important in grading one another, I believe that personality has a lot to do with how we determine our grades. That is, evaluating the subjectivity of a group member by making personality attributions and inferences. If a group member is particularly sprightly and in turn brings about a higher level of group cohesiveness by facilitating discussion, then they therefore have improved group performance and should possibly be awarded better grade points. It seems that there is indeed a subjective side to peer evaluations, which is a haven for all kinds of different interpretations.
Are you saying that there is no way to get an empirical measure of personality? What about the MBTI, NEO, MMPI, and others..?
Here it seems you've added another criterion (or a couple of criteria) for adding to your "peer grading system". But the questions are, should and can 'personality factor(s)' (i.e. "sprightliness") be added into the grading system in their own right? OR are you saying the 'results of personality factors' (i.e. activation of group cohesiveness behaviours) should be all that contributes to your grading system? AND are these both (or individually) examples of 'contribution' as you want to define it earlier or should they exist under a different criteria of grading?
I think it would be easier to avoid talking about personality and just go to the specific factors to be measured, such as "members who contributed to group cohesiveness = a group cohesiveness score" and explain how this score was given - how a member could contribute to group cohesion - such as "the amount of talk that created discussion", and explain how this was measured exactly (by meeting minutes or from memory of group member(s)), and give each member a score according to this measure. Calculating a group cohesiveness score might be more complicated than, for example, just counting how many emails were sent or how many meetings were attended, but empirical data can still be used to calculate something such as 'group cohesiveness', you just need to clearly define what the phenomenon of group cohesion is and what empirical data are used to give it a score.
Quote: XUL said: The peer grading system seems is based on the fact that we cannot grade each other evenly and so, if appropriate, we must either distribute grades in any way we please or ‘regress to the mean’ (if we believe our group members are truly equal). That is, due to the inability to properly grade accurately based on the given instructions we must give equal scores. In a group of six members, If the grading evaluator attempts to regress to the mean by assigning each other grades that are as even as possible then we are left with one low score, three equal scores, and one high score (e.g. 13,14,14,14,15 = 70). Therefore we are forced to determine one “slacker” student, three average students, and one student “leader.” But there is often an outlier, which seems to reveal a flaw. The grading system for peer evaluations does not leave room for an outlier – the group that performs evenly – or believes that it does. The flaw exists in that we cannot, without penalty, grade each other evenly when indeed there is always an off chance that a group could function evenly, or believe that it functions evenly.
This is really confusing. It kind of seems like you don't understand what "regress to the mean" actually means. Regressing to the mean has nothing to do with whether you believe group members are truly equal or not. But you also seem to be going off on a tangent talking about being unable to properly grade members.. why not? Why can't you use measures of the criteria you listed earlier to give a grade?
But on the topic of outliers... Although in a data set there is often an outlier(s), does the absence of an outlier really mean that there is something wrong and must there always be an outlier? Or can a data set be completely fine even though there is no outlier (i.e. can there be valid reasons why there there is no outlier)?? The answer is, "Yes, there can be valid reasons why there might not be any outlier in a data set". So using 'the absence of an outlier' as reason for why a grading system is not valid, is not a good reason.
It might help to stop talking hypothetically. Do you honestly believe that everyone in your group really functioned evenly so that they each really deserve exactly the same grade? If so, how are you specifically making this evaluation? What are the criteria, measures, and how are you empirically substantiating the equal grades you give?
It doesn't make sense to on one hand outline possible measures to base a grading system and then on the other hand, for seemingly no reason at all, ignore those possible measures and say that peer grading can't be done and therefore the grading system is flawed.
BUT, it might make more sense to say it is impossible to come up with a grade IFF you went through the list of possible measures you already provided one by one and explain what exactly you want to measure for each, explain how you are going to measure it, and then explain why you can't do it.
For example, "Meeting attendance is a criterion to be measured and 'the proportion of the total number of meetings each member attended' was to be measured, and meeting logs (documentation) was to be used as the data set, but there is no data available because no logs/documentation were made and everyone in the group has since died or otherwise uncontactable meaning there is no way to obtain the data, thus it is impossible to use attendance as a way to determine each group member's grade" AND/OR "Electronic communication in the form of email is a criterion to be measured and 'the number of emails sent by "reply all"' was to be measured, and this data was collected by doing a search of email history of John Smith (a member of group) and counting the number of emails sent and how much each person contributed to give an 'email score', but the email histories of every group member has been erased by a super virus and as such there is no data set available to measure and thus this measure is not able to be used in the calculation of 'contribution'". Or something like that... but I'd be really surprised if there was not one single criterion for which you could use to measure a grade.. and I think your teachers would agree, hence, an argument that it is impossible to come up with a grade is so far pretty much a waste of your time trying to argue..
My advice would be to abandon the whole "it's all subjective and thus is impossible to measure" argument, because that argument is flawed, particularly because you haven't explained why it is impossible to measure subjective phenomena and also because psychology has proven again and again that it is possible to measure subjective phenomena (even if people can argue over the validity and reliability of the measures).
I think you'd be better off if you stop being hypothetical and simply just: 1) list a bunch of criteria that could make up the grade, then 2) explain why you chose those criteria to evaluate the grade, then 3) explain how you want to measure those criteria, then 4) explain where you're getting the measurements from and what data sets you're using, then 5) present the data for each criterion, then 6) explain how the data of each criterion contributed to an overall score, then 7) give each person's overall score, then 8) discuss the scores, discuss the strengths and weaknesses of the process (i.e. strengths of some data sets were more valid/verifiable for contributing to an overall grade and were less subject to subjective interpretation, while a weakness/limitation might have been no way to exclude researcher bias and that some data was sketchy or unreliable or unavailable or lacking statistical power because of small sample size, etc), discuss how the scoring process could contribute the concept of giving a grade, and discuss what criteria should be used again or not again, what criteria should be used that wasn't used, and what could make such scoring more accurate and reliable in the future.
In effect, what I've listed is using a similar pattern to how scientists outline their research in journals.. "Introduction", "Method", "Results", "Discussion". You could limit the amount of work you have to do by simply limiting the criteria you choose to the "best 5" and persuasively explain why they are the "best five" criteria (out of all the possible criteria you could choose from) and why these 5 criteria are sufficient for calculating an accurate enough and reliable enough grade. If you don't have data for a criterion, such as leadership, then you might want to work out a way to get the data, such as attempting to gather it from somewhere and/or create a questionnaire and ask other group members to give their ratings, then calculate means and standard deviations/confidence intervals of the results (I think data collection from other group members might give you extra points with your teacher).
...all just a suggestion of course. Feel free to ignore. Good luck with it.
--------------------
|