Evidence is a collection of Likelihood Ratios LR=TPR/FPR. The True Positive Rate TPR and the False Positive Rate FPR define the evidence accuracy: A=1/2+(TPR-FPR)/2. Perfect accuracy (A=1) has TPR=1 and FPR=0. Coin-toss accuracy, i.e. perfect inaccuracy, has TPR=FPR (or LR=1), hence A=50%. Accuracy establishes a trade-off between TPR and FPR: given A, a higher TPR implies a higher FPR.
Evidence is hard if accuracy can be properly measured; it is soft if it can’t. Hard evidence is (or should be) inescapable. But soft evidence is in the eye of the beholder: its perceived accuracy coincides with the observer’s confidence and is therefore determined by trust. Unscrupulous experts manipulate evidence by boosting TPR and hiding the consequent increase in FPR.
Albert Edwards, Chief Global Strategist for Société Générale, has been calling stock market crashes for longer than I care to remember. His unrelenting efforts have gained him some notable Hits in the last two rollercoasting decades, as well as an at least equivalent number of False Alarms. He is like a coin thrower always calling Heads: if properly measured, a generous estimate of his accuracy would be 50%. But it is not what exudes from his bubbly confidence and the consideration of his audience. How does he do it? He has developed brushing Tails under the rug into a fine art:
From yesterday’s Financial Times:
“In the sense that we are so far through the equity bear market, I’m relatively more bullish. I expect the S&P to go below 666. I expect there to be total carnage. But I’m more bullish than I was.”
“Regular readers will know that in the main, my market timing is unerringly inaccurate, normally months if not years too early.”
Yep.
We should perhaps cut Edwards some slack – after all, he’s hardly unique (didn’t Samuelson say something about “Wall St has predicted nine out of the last five recessions” ?). I’ve certainly come across equally egregious abuse of the concept of “accuracy” in fields ranging from meteorology and seismology to pharma. It’s been particularly abused in the debate over the use of animal models. What little quantitative research that has been done into the predictive value of such models has focused on the true positive rate, while ignoring the false positive rate – thus leaving the LR of animal models undefined. When one does the digging to find both rates, it turns out that animal toxicity models tend to have a good true positive rate (that is, if the compound harms the animal, humans are likely to be harmed too) but a poor true negative rate (if the animal shows no harmful effects, that’s weak evidence that humans will be ok).
So I like the idea of coming up with an accuracy parameter which compels the use of both TPR and FPR. A couple of observations, though. There are two types of LR (positive and negative), and one needs to be precise about which is implied (the former, in your formula I think).
More to the point, shouldn’t perfect inaccuracy correspond to A = 0, rather than to the A = 0.5 value given by a coin-toss ?? For me, this highlights a bit of a paradox about predictive accuracy. Specifically, perfect predictive INaccuracy is inferentially identical to perfect predictive accuracy – you just flip the prediction around. Thus, for example, if Punxsutawney Phil were 100 per cent useless at predicting an early spring on Groundhog Day (Feb 2, incidentally), he’d also be 100 per cent perfect at it – as you can just flip whatever he indicates around. And this highlights the fact that, as you indicate, a coin-toss is the only truly useless prediction, and thus should perhaps get a rating of A = 0 rather than A = 0.
So how about this formula for predictive accuracy: A = TPR – FPR
This leads to A = 1 when TPR = 1 and FPR = 0, to A = -1 when these are reversed, and to A = 0 when TPR = FPR (ie a coin-toss). What do you think ?
Thanks for a(nother) thought-provoking post !
Hi Robert. A is just the average of the True Positive Rate and the True Negative Rate: A=(TPR+TNR)/2, so it is a natural measure of overall accuracy – which TPR-FPR=TPR+TNR-1 wouldn’t be (e.g. TPR=TNR=80% would imply A=60%). Coin-toss accuracy (A=50%) is perfect inaccuracy in the sense that it is perfectly useless. As you say, A<50% would be useful as a contrarian indicator – I know quite a few “market experts” who fit the bill. In the extreme, A=0 means that the expert gets exactly everything wrong – which would make him a priceless adviser, if you just do the opposite! The insidious quality of useless experts is precisely this: they are right 50% of the time. The same is true, in general, of useless evidence: LR=1.
In your animal toxicity example, if AT=Animal Toxicity, HT=Human Toxicity, then TPR=Probability of AT given HT, and FPR=Probability of AT given not-HT. In your table in the Appendix, TPR=60% and FPR=33%, hence LR=1.8. Since the Base Rate of HT is 5/23=22%, the Base Rate Odds are 0.3, hence the Posterior Odds are 0.5, giving a posterior probability of 33%. So the AT evidence increases the probability of HT from 22% to 33% – which, as you rightly say, is too low.
Question: ho do you calculate the 95% confidence level for LR?