Apr 022014
 

Last year I was watching the Wimbledon final with my elder son. A few days earlier – me being me – I had explained to him the meaning of the word ‘probability’: if you are completely sure that something will happen, its probability is 100%; if you are sure there is no way it will happen, the probability is 0%; if you have no idea whether it will happen or not, it is 50%. In defiance of the pervasive national craze, for some reason (contrarian gene?) he wanted Djokovic to win. So I asked him: What do you think is the probability that he will win? 70% – he said, after a few moments’ pondering. And what about Murray? – I asked. 30% – he said straightaway. He was six years old.

In Kahneman and Tversky (1973) Engineers vs. Lawyers experiment, participants were assumed to be at least as smart as a six years old. So after asking them the probability that the described person was an engineer, K&T did not bother to ask them the probability that he was a lawyer (they actually asked the low priors group the probability of engineer and the high priors group the probability of lawyer, noted that the two added to about 100% and proceeded to express the pooled data in terms of probability of engineer).

Baratgin and Noveck (2000) were not so convinced. So, after running their Math vs. Literature teachers experiment and reproducing results similar to K&T, they ran two additional experiments. In the first, they asked 40 participants, again equally divided into a low priors and a high priors group, the probability that the described person was a math teacher and the probability that he was a literature teacher. In the second, they asked another set of 40 participants the same two questions, but explicitly pointed out to them that the two numbers had to add up to 100%.

If K&T’s assumption is valid – participants are sufficiently switched on as to figure out by themselves that exclusive and exhaustive probabilities must add up to 100% – results in the two new experiments should not differ significantly from the results of the first experiment. But this is not what B&N found.

Average estimates for the new experiments are reported in the following table (from their Table 1), along with the standard format results presented in the earlier post.

Table 1

In the following graph, the probability pairs are presented as yellow dots for the standard format, as red dots for the Suggested Complementarity (SC) format and as purple dots for the Induced Complementarity (IC) format.

Figure 1

Interestingly, there is a clear move towards the Bayesian curve, suggesting that the standard format estimates might have been influenced by an undetected failure to adhere to the complementary constraint. Incredible but true, more than a quarter of the participants in the SC experiment (11 out of 40) failed to comply, thus suggesting that, if detected, failures in the standard experiment might have been even more numerous.

The SC format is the same later adopted in the Gloms vs. Fizos experiment in Villejoubert and Mandel (2002), where I commented that the failure to comply with complementarity (also called the additivity principle) was, in my opinion, a sign that the participants did not really understand the questions. David Mandel disagrees. But it does seem to me that, rather than suggesting additivity, the SC format may actually contribute to generate confusion, leading the less alert participants to conclude that, since both numbers are requested, perhaps they are not meant to add up to 100%.

The IC format, on the other hand, entirely avoids all potential confusion, by focusing participants on the true meaning of the required estimate. The IC format is nothing but the standard format plus a proviso such as “sorry to be obvious, but it goes without saying that, if your estimate is x%, the probability of the alternative is 1-x%”. The proviso proved to be far from redundant. IC estimates – together with very similar estimates from the 29 participants in the SC experiment who complied with additivity – were significantly different from the estimates in the standard format experiment. Differences between low and high priors estimates were much more pronounced for each of the five descriptions, averaging 18% versus 7% in the standard format. In particular, the difference in Paul, the neutral description, went from 6% to 19%, and the apparent bias noticed in the standard experiment disappeared, with low and high priors estimates nicely centred at around 50%, and halfway from the 30/70 Bayesian estimates.

Notice, by the way, that the second assumption in the experiments – that evidence was perceived as symmetric – could not have influenced the results. Treating the problem in odds form – thus allowing for asymmetry: TPR+FPR≠1 – does not change the conclusions. The low priors–high priors locus in the figure above remains the same if drawn for different levels of the Likelihood Ratio LR rather than for different levels of TPR. For each level of LR under asymmetry there is an equivalent level of TPR under symmetry: TPR=LR/(1+LR). And, under prior indifference, PP=TPR is equivalent to PO=LR, i.e. PP=LR/(1+LR).

B&N’s experiments underscore the fact that prior indifference should be seen as a tendency, not as invariant mechanism. By centring Paul at 41% and 59%, Figure 1 becomes:

Figure 2

So, while in K&T, and in B&N standard experiment, prior indifference explained most of the departure from the Bayesian curve – the best fit Base Rate was 47% – in the IC format prior indifference accounted for about one half of the move. Not as strong, but still a potent effect. B&N’s results might have been affected by other issues – use of averages instead of medians is one that I would like to know about – but it clearly indicates that, for a correct measurement, the standard question should be accompanied by an apparently redundant but, in point of fact, necessary proviso.

Another suggestion is the introduction of a third group, which, as in the Gloms vs. Fizos experiment, is given 50/50 priors. By providing direct measures of accuracy, unencumbered by gaps between actual and perceived Base Rates, this control group may give useful indications as to the overall consistency and transparency of the experiment.

  8 Responses to “Wimbledon’s winner”

  1. “The SC format is the same later adopted in the Gloms vs. Fizos experiment in Villejoubert and Mandel (2002), where I commented that the failure to comply with complementarity (also called the additivity principle) was, in my opinion, a sign that the participants did not really understand the questions. David Mandel disagrees. But it does seem to be that, rather than suggesting additivity, the SC format may actually contribute to generate confusion, leading the less alert participants to conclude that, since both numbers are requested, perhaps they are not meant to add up to 100%.”

    We have to be clear on what complementarity refers to in these discussions. You originally suggested that perhaps we (Villejoubert & Mandel, 2002) should have used an induced complementarity format to convey the diagnostic probability information. In response, I had said that doing so was not obligatory since truncation is normative in language, at least from the stance of pragmatics a la Grice. But in your current post you are talking about the complementarity of the posterior probabilities that subjects in our experiment judged. As you said, we used a suggested complementarity format, which is true. But while the first issue — representing the diagnostic information — has to do with information presentation, the latter — eliciting the posterior probabilities — has to do with judgment. Both may affect the logical coherence of judgments, but they are distinguishable issues.

  2. Sorry David – I created a bit of confusion when I said that my preferred format for the probability question was what Baratgin and Noveck call Suggested Complementarity. What I meant was Induced Complementarity (I corrected the mistake). Whereas I was not advocating any change in the diagnostic information format.

    So to recap: your diagnostic information – e.g. 98% of Gloms and 58% of Fizos play the harmonica – is in the standard format, and I think that is fine. But your probability question is in the SC format and, for the reasons I explain in this post, I think it should be in the IC format.

  3. Glad we’ve cleared up one issue, but there’s still one small point I’d like to make regarding your last comment:

    “But your probability question is in the SC format and, for the reasons I explain in this post, I think it should be in the IC format.” (italics added)

    I’m not sure why you are claiming it should be in IC format. I would tend to agree if you had said “it would be of interest to vary the IC/SC formats to see what effect it might have in this paradigm.” I think we both expect that it would have a comparable effect. This would triangulate with previous demonstrations of the same effect, but it would not be front-(journal)page news. But your choice of the term “should” suggests you believe that it is wrong in some sense to use the IC format. If that is in fact what you meant, then I would disagree with you. In the real world, SC represents a kind of intellectual spoon feeding that is exceedingly rare. The fact that people miss the connection between complementary events even when they judge them side by side (in the IC format)–itself a rarity–is, to my mind, the interesting aspect for it shows how easy it is for people to lose track of relevant principles that ensure coherence in judgment. That is, if they lose track. Presumably, some proportion may not even understand the principle and thus have “nothing to lose.”

  4. This is funny – there must a devil inside this SC/IC business: even you messed them up! (read the second half of your comment).

    I do think it is wrong to use the SC format – not in all cases, of course (see Support theory, extensionality etc.), but certainly in cases like your Gloms vs. Fizos, or the K&T and B&N experiments, where failure to comply with additivity can only mean one thing: lack of understanding of a super elementary question: is it a Glom or a Fizo? This is not the same as, for example, Fischhoff et al. (1978) car mechanic. It is a six-year old-proof question: it is either a Glom or a Fizo, period. And if someone tells me, like your average respondent to questions 7 and 8 (see Table 3 in my post), that he reckons there is a 42% probability that it is a Glom and a 10% probability that it is a Fizo, I don’t find that answer “interesting”, but a sure sign that he did not get the question right. Therefore, his answer cannot help me to understand what I am investigating and is just a source of confusing and harmful noise.

    And I would indeed expect different results in your experiment if you use the IC format – in fact it is what you pretty much get after you normalize the estimates (Table 4 in my post), which is what I would expect after imposing prior indifference.

  5. Yes, you’re right, I mixed up SC and IC in my last comment. Clearly, we’re both better at detecting each other’s mixups than our own.

    I don’t think the issue is quite as straightforward as you’re suggesting. People violate additivity even with binary complements for different reasons. One is that they lose track of the principle which involves having to track the consistency of related responses. Is losing track of a consistency principle the same thing as not understanding individual questions? I don’t think so. In fact, I suspect that even if you screened subjects for correct question interpretation, you would still observe violations of additivity in the remainder’s responses. Part of that may be due to losing track of the principle. Another possibility is that a subject doesn’t understand the additivity principle, even though they understand each individual question. I described these alternative sources of nonadditivity in:
    Mandel, D. R. (2005). Are risk assessments of a terrorist attack coherent? Journal of Experimental Psychology: Applied, 11(4), 277-288.
    I also disagree that the nonadditivty is just “confusing and harmful noise” because as my colleagues and I have shown, it could be leveraged to improve to improve judgment. In fact, as we’ve shown it is sometimes better to let incoherence express itself and use it than to take steps to minimize it. If you’re interested, see here:
    Karvetski, C. W., Olson, K. C., Mandel, D. R., & Twardy, C. R. (2013). Probabilistic coherence weighting for optimizing expert forecasts. Decision Analysis, 10(4), 305-326. http://dx.doi.org/10.1287/deca.2013.0279

    And also: http://couchpsychologist.wordpress.com/2014/03/30/harnessing-individual-differences-in-incoherence-to-improve-forecasting-accuracy/

  6. David – there is no question that people do violate additivity. There is plenty of evidence about it, including the experiments described in your interesting papers. However, what we are talking about here is binary complementarity. As you know, that is an axiom in Support theory. Now, as Macchi et al. (1999) have shown, additivity can fail even for binary partitions. But in their experiments, just as in all other experiments, including yours, each participant was asked only one side of the question, not both. That is not what happened in Gloms vs. Fizos. Macchi et al. did show additivity violations for binary partitions, but their “findings are consistent with the intuition that additivity will result from having the same person assess the probabilities of all components of a given partition” (p. 213).

  7. Actually, that’s not so. In Mandel (2005) which I’ve referenced earlier I report both between- and within-subjects tests of additivity and there is additivity violation within-subjects in their estimates of the probabilities of future terrorist attacks occurring or not occurring, more so when the complements are spaced apart by an intervening task.

    In Williams and Mandel (2007), we show how additivity violations for binary complements of conditional probabilities vary as a function of whether they were elicited using valuation frames (as described in the Tversky & Koehler 1994 Support Theory paper) or what we call economy frames which truncates the reference to the complement in the query. Again in that research we report superadditivity for binary complements in both elicitation conditions although it was attenuated in the evaluation frame condition (in line with your argument that moving from the standard format to SC will improve additivity).

    The Decision Analysis paper I referenced the other day (Karvetski et al., 2013) which shows how incoherence due to additivity violation can be leveraged is also within-subjects.

    And of course we showed it in Villejoubert & Mandel (2002). These are only my studies. I would be surprised if I were the only one to have shown additivity violations of binary complements.

    Finally I do realize that additivity of binary complements is axiomatic in support theory. The theory is therefore wrong in that respect since there is ample evidence from both within and between subject designs that people do violate the additivity property even for binary complements. Not always, of course, but they do, and more often than you seem willing to believe. If a theory does not accord with nature, we do not revise nature. We revise the theory, no matter how elegant it may be.

  8. You don’t have to worry about that – in fact, how evidence changes beliefs is the unifying theme in my blog, and everything else I do.

    On this issue, however, I guess we have to agree to disagree. You are right, of course, in saying that everything that happens needs to be explained, not ignored or assumed away. But on binary complementarity, when someone says P(H)=x and, at the same time, P(not H)≠1-x, I think there is only one explanation: he doesn’t understand what he is talking about. A probability is a number such that P(H)+P(not H)=1. Anything else is unacceptable, just like P(H)=300%, or -23. It is like asking people what kind of apples they like and someone answers: conference pears. It is not an error of judgement – it is error of grammar. Remember K&T (1982, On the study of statistical intuitions): “The student of judgment should avoid overly strict interpretations, which treat reasonable answers as errors, as well as overly charitable interpretations, which attempt to rationalize every response”.

Any comments?

This site uses Akismet to reduce spam. Learn how your comment data is processed.