Following suggestions from readers of my paper on the Inverse Fallacy, I have been thinking about possible ways to test my idea that the fallacy is due to prior indifference.
Prior indifference is closely related to Base Rate neglect: being indifferent about whether a hypothesis is true or false implies neglecting its Base Rate. The relationship between the Inverse Fallacy and Base Rate neglect is well established, at least since Kahneman and Tversky (1973), and has been subsequently validated in many other studies (e.g. Bar-Hillel (1980)). As typified in the cab problem, evidence about a hypothesis can cause people to disregard its prior probability. There is plenty of support about this and, I think, no real need for further empirical tests.
My point, however, is that people do not ignore Base Rate information per se. They are actually well aware of it and, in the absence of other evidence, would naturally use it as their best estimate of the probability of the hypothesis. But evidence can blind them to Base Rates, causing a distortion of Bayesian updating. While a correct update would start from a Base Rate and increase it or decrease it according to the Likelihood Ratio of new evidence, evidence itself can trigger an inadvertent shift of the Base Rate to 50% before the update takes place. As a result, the update builds on Knightian uncertainty and perfect ignorance, rather than on prior beliefs. Therefore, the best way to think about the distortion is to call it what it ultimately is: a Prior Indifference Fallacy.
Under prior indifference, PP=TPR/(TPR+FPR). Hence PP is not exactly equal to TPR, except when evidence is symmetric. In that case, since FPR=FNR, we have TPR+FPR=1 and therefore PP=TPR. The extent to which PP differs from TPR depends on the level of asymmetry ASY=FPR-FNR:

Therefore, if the Inverse Fallacy is due to prior indifference, we should expect people to estimate PP with TPR when they are presented with symmetric evidence (as in the cab problem), and to do so irrespective of the Base Rate. But if the presented evidence is asymmetrical, the expected estimate should be TPR/(1+ASY), again irrespective of BR.
Villejoubert and Mandel (2002) describe the results of an experiment where participants – 45 undergraduates from the University of Hertfordshire – were presented with the following problem:
Imagine you visit a planet inhabited by two million invisible creatures: one million Gloms and one million Fizos. You meet twelve of them and, through an interpreter, ask each of them a question, to which they answer Yes or No. Your task is to estimate the probability that each creature is a Glom, based on the following evidence:
Table 1

So, for example, you are told that 98% of Gloms and 58% of Fizos play the harmonica. You ask the first creature if he plays the harmonica, and he says Yes. Next, you are told that 2% of Gloms and 58% of Fizos exhale fire. You ask the second creature if he exhales fire, and he says No. And so on.
In our framework, assuming the hypothesis is “The creature is a Glom”, the Gloms column is the True Positive Rate: the probability that the creature shares the feature, given that he is a Glom. And the Fizos column is the False Positive Rate – the probability that the creature shares the feature, given that he is not a Glom but a Fizo. Here is the full picture:
Table 2

The first four columns report TPR and FPR from Table 1, as well as their complements to 1, FNR and TNR. As shown in the next column, all pieces of evidence are highly asymmetrical. The next column reports the level of accuracy A=(TPR+TNR)/2. Remember A goes from 1 (perfect accuracy) to 0 (perfect contrary accuracy), going through 0.5 (perfect inaccuracy). So the first piece of evidence is a mildly accurate indication that the creature is a Glom, and the second piece of evidence is an equally mildly accurate indication that he is a Fizo. The same is true for the other answers. Next, PP is the probability that the creature is a Glom if the answer is Yes, and NP=FNR/(FNR+TNR) is the probability that he is a Glom if the answer is No. Since odd answers are Yes and even answers are No, the last column reports the relevant estimate, according to Bayes’ Theorem. Notice that the first two answers give the same 63% probability, the second two the same 64% probability, and so on. So they can be grouped, as in the first column of the following table:
Table 3

According to Bayes’ Theorem, the first couple of answers imply a 63% probability that the creature is a Glom and therefore, as reported in the second column, a 37% probability that he is a Fizo. For the second couple the probabilities are 64% and 36%, and so on.
In order to test whether the participants were subject to the Inverse Fallacy, the authors asked them to give, for each of the twelve encounters, the probability that the creature was a Glom, and the probability that he was a Fizo. The average answers are reported in the third and fourth column (derived from Table 2 in the paper). So, for example, the average estimated probability that the creature was a Glom, based on answers 1 and 2, was 77%. And, incredibly, the average estimated probability that he was a Fizo was 46%, for a total of 123% (fifth column). For answers 7 and 8, the probabilities were 42% and 10%, for a total of 52%! Clearly, students were either particularly dumb or, more likely, very confused.
One wonders what their estimates would have been if they had been required to respect the obvious constraint that the two numbers needed to add up to 100% (the additivity principle). After all, in being presented with Table 1, students were left to figure out by themselves that if, for example, 98% of Gloms play the harmonica, that means that 2% don’t. I suspect that deviations from Bayes would have been significantly smaller. To get a rough idea, if we normalise the observed estimates by dividing them by the observed sum, results are as follows:
Table 4

The normalised estimates are much closer to the Bayesian benchmark. This is to be expected if the Inverse Fallacy is due to prior indifference. Since in the experiment the Base Rate is assumed to be 50% – there are one million Gloms and one million Fizos on the planet – prior indifference is justified. And since the evidence is highly asymmetrical, it is wrong to expect PP=TPR and NP=FNR – as the authors do.
It would be interesting to repeat the experiment by:
- Asking only for an estimate of the probability that the creature is a Glom.
- Run parallel experiments with different Base Rates, e.g. a low BR=25% and a high BR=75%.
If the Inverse Fallacy is a Prior Indifference Fallacy, we should observe that the estimates under BR=50% does not differ significantly from the Bayesian benchmark, as well as from the estimates under a low and a high BR.
In an additional experiment, participants could be asked to imagine that, rather than meeting twelve different creatures, they meet just one creature and ask him all the twelve questions. It would interesting to check the extent to which new evidence builds on Bayesian updates or is each time rebased to the Base Rate.
It’s been over a decade since I coauthored that study, but I’m glad to see it still generating interest. I wanted to note that subjects in our study did not receive the information as presented in Table 1 but rather were presented with that information one problem at a time. They also received base rate information about Gloms and Fizos at the start, since they were told that the two types of creature were in equal numbers on the planet.
I was somewhat perplexed by your comment that the participants were either incredibly dumb or very confused because the sum of their posterior probabilities, which should be additive were systematically non-additive. Additivity violations, even for binary complements, are not uncommon. For instance, I reported more severe violations of additivity in students’ estimates of terrorist attacks. In that case, the estimates were superadditive (see Mandel 2005). Likewise, I showed predictable additivity violations in binomial probability assessments in Mandel (2008, Exp. 6). It may be that participants are confused by probability assessment tasks, but how they resolve that confusion is not by random guesswork. As we showed in our Glom/Fizo experiment, violations of additivity are predictable and systematic. The representational account I developed in the 2008 paper goes further in terms of explaining why such violations occur.
Finally, I was puzzled by your statement: “One wonders what their estimates would have been if they had been required to respect the obvious constraint that the two numbers needed to add up to 100% (the additivity principle). After all, in being presented with Table 1, students were left to figure out by themselves that if, for example, 98% of Gloms play the harmonica, that means that 2% don’t. I suspect that deviations from Bayes would have been significantly smaller.”
If it is surprising to you that subjects violate the additivity property, then why would it be important to provide redundant information to them? Presumably if the subjects were non-additive because they were dumb, then wouldn’t a replication with brighter students eliminate the biases shown? I don’t think you’ll find that since the inverse fallacy has been shown in expert studies where presumably the subjects aren’t dumb.
In daily life, redundant information is usually truncated. We don’t usually say “we’re 70% sure it will rain and also 30% sure it won’t.” We just say one or the other. Of course, the more you do to rule out possible errors and biases, the less likely subjects are to deviate from Bayesian probabilities. The restructuring of probability problems into natural frequency trees is a good example of taking such steps. Although scholars still disagree on why it works, most agree that it does work. Likewise, your proposal to take the information framed in terms of attribute possession and flesh it out in terms of attribute non-possession might clarify the subsets of cases pertinent to Bayesian inference.
A final point about incoherence: sometimes it can be exploited to improve judgment. That might seem counter-intuitive, but my colleagues and I (Karvetski et al. 2013) have shown how that can be done in forecasting environments, where multiple experts’ assessments are pooled. Knowing their relative incoherence can provide a valuable cue for weighting their contribution to the pooled estimate.
Thanks for bringing this to my attention and my apologies for the delay in replying to your email. I seldom check my York account.
David
References
Karvetski, C. W., Olson, K. C., Mandel, D. R., & Twardy, C. R. (2013). Probabilistic coherence weighting for optimizing expert forecasts. Decision Analysis, 10(4), 305-326.
Mandel, D. R. (2005). Are risk assessments of a terrorist attack coherent? Journal of Experimental Psychology: Applied, 11(4), 277-288.
Mandel, D. R. (2008). Violations of coherence in subjective probability: A representational and assessment processes account. Cognition, 106(1), 130-156.
Thank you so much for your comments and for pointing out your other papers, which will be very interesting to read. To your points:
1) It is indeed clear that participants received the information in Table 1 (which reproduces your Table 1) one piece at the time, and that they were given the 50/50 base rate at the start. As I say in the last point in my post, it would actually be interesting to see what happens if participants are asked to imagine that, rather than meeting twelve different creatures, they meet just one creature and ask him all the twelve questions.
2) I take your point that additivity violations are common, and worth investigating. I will read your papers. But it seems to me at the moment that, if participants do not get that exclusive and exhaustive posterior probabilities must add up to one, it is not because they are dumb (of course they aren’t) but because they are not really understanding the question.
3) I am not proposing to change the format of the information in Table 1. On the contrary, I am saying that the question about the posterior probability should have the same format, i.e. “What is the probability that the creature is a Glom?”, with an added proviso like “Sorry to insult your intelligence, but let me point out that, as obvious as I am sure it is to you, if you say that the probability of Glom is x% you are implying that the probability of Fizo is 1-x%”. It is what Baratgin and Novek (2000) call the suggested-complementarity format.
4) I am not saying that, by imposing additivity, you will ensure correct Bayesian probabilities. I am saying that it is what would happen in your Glom/Fizo experiment, because you gave participants a 50/50 base rate. But I would expect that if you vary the base rate and repeat the experiment with, say, 25/75 and 75/25 base rates, answers would not be significantly different, because of base rate neglect or, more precisely, of what I call prior indifference.
Hi Massimo — on point 3, I’m not sure why you think there’s a need to do what you’re recommending. If the purpose is to see whether such communications could improve reasoning, then I get it. But, in general, there seems to be no good case for such communication being obligatory. That is to say, we were not playing any sort of “trick” on subjects by not giving that information (we just weren’t jumping through flaming hoops either). I do share your intuition that reminding people of the complement can serve a debiasing function. Joseph Williams and I have shown that in a study we reported in 2007 in the context of conditional probability judgement.
I agree that varying base rates in similar experiments could be informative. We intentionally kept base rates constant since we wanted to examine predictions about the inverse fallacy in an experimental context where base rate neglect could be definitively ruled out as a source of judgment bias.
By the way, I reblogged your post with some commentary on my spanking new blog:
http://couchpsychologist.wordpress.com/2014/03/26/gloms-and-fizos-reloaded/
(This being the first substantive blog entry!)
Reference
Williams, J. J., & Mandel, D. R. (2007). Do evaluation frames improve the quality of conditional probability judgment? In D. S. McNamara & J. G. Trafton (Eds.), Proceedings of the 29th Annual Meeting of the Cognitive Science Society (pp. 1653-1658), Mahwah, NJ: Erlbaum. Available: https://sites.google.com/site/themandelian/home/examples/files/CCS2007.pdf