Pages

Showing posts with label Jordan Ellenberg. Show all posts
Showing posts with label Jordan Ellenberg. Show all posts

Tuesday, October 14, 2014

Book Review: How Not to Be Wrong by Jordan Ellenberg





A while back (11/2/11) I reviewed the book Stats.con by James Penston. That book discussed how the statistics used in randomized clinical trials can be highly deceptive. How Not to be Wrong also covers some aspects of statistical misuse, in more detail, and certainly in a much more entertaining way. Some of his comments are funny as hell. 


Jordan Ellenberg

Consider the widespread use of a statistic call the p value, which estimates the probability that the result of a study could have just been a chance coincidence rather than an actual meaningful finding. A study is generally considered positive if the p value is 5% or less.

5% is of course not 0%. There is a one in twenty probability that the study results that are deemed positive were in fact negative. But what happens if journals only publish the positive studies and not the negative ones, when there might be a large number of negative studies, and when the positive study results are not reproduced (replicated) in a second study? Well, people start believing things that are not true, that's what.

The reason for this is because, as the author points out, improbable things actually happen quite frequently. Especially if you do lots and lots of things – like experiments. 

Another issue he mentions is that, if your sample size is too small, the chances increase dramatically that one of your subjects will be an outlier that dramatically but artificially changes the average for whatever characteristic you are measuring. With a small sample, you are more likely to get a few extra prodigies or slackers in a study of people's ability to perform certain tasks. A famous example: if Bill Gates walks into a bar with a few other people, the average guy in the room is a billionaire.  

Here’s how the author starts out a discussion of the p value problem (pages 145-46) :

"Imagine yourself a haruspex; that is, your profession is to make predictions about future events by sacrificing sheep and then examining the features of their entrails...You do not, of course, consider your predictions to be reliable merely because you follow the practices commanded by the Etruscan deities. That would be ridiculous. You require evidence. And so you and your colleagues submit all your work to the peer-reviewed International Journal of Haruspicy, which demands without exception that all published results clear the bar of statistical significance.    
      
Haruspicy, especially rigorous evidence-based haruspicy, is not an easy gig. For one thing, you spend a lot of your time spattered with blood and bile. For another, a lot of your experiments don't work. You try to use sheep guts to predict the price of Apple stock, and you fail; you try to model Democratic vote share among Hispanics, and you fail…The gods are very picky and it's not always clear precisely which arrangement of the internal organs and which precise incantations will reliably unlock the future. Sometimes different haruspices run the same experiment and it works for one but not the other — who knows why? It's frustrating…
      
But it's all worth it for those moments of discovery, where everything works, and you find that the texture and protrusions of the liver really do predict the severity of the following year's flu season, and, with a silent thank-you to the gods, you publish
      
You might find this happens about one time in twenty.
      
That's what I'd expect, anyway. Because I, unlike you, don't believe in haruspicy. I think the sheep's guts don't know anything about the flu data, and when they match up it's just luck. In other words, in every matter concerning divination from entrails, I'm a proponent of the null hypothesis [that there is no connection between the sheep entrails and the future]. So in my world, it's pretty unlikely that any given haruspectic experiment will succeed.
      
How unlikely? The standard threshold for statistical significance, and thus for publication in IJoH, is fixed by convention to be a p-value of .05, or 1 in 20... If the null hypothesis is always true — that is, if haruspicy is undiluted hocus-pocus —then only one in twenty experiments will be publishable.
      
And yet there are hundreds of haruspices, and thousands of ripped-open sheep, and even one in twenty divinations provides plenty of material to fill each issue of the journal with novel results, demonstrating the efficacy of the methods and the wisdom of the gods. A protocol that worked in one case and gets published usually fails when another harupex tries it, but experiments without statistically significant results do not get published, so no one ever finds out about the failure to replicate. And even if word starts getting around, there are always small differences the experts can point to that explain why the follow-up study didn't succeed."

The book covers many subjects about which the non-mathematically-inclined can learn to think in a mathematical way in order to avoid coming to certain wrong conclusions and to zero in on correct ones. Many of these, however, are irrelevant to this blog – the chapters on lotteries come to mind. I of course found those parts a bit less interesting. But the chapters relevant to medical studies are so right on.

Another important topic the author covers is known mathematically as regression to the mean. This phenomenon can lead, as examples, to overestimates about the genetic component of human traits and explains why fad diets always seem to work at first but then later on everyone seems to forget about them. As mentioned, when you average any measurement applied to human beings, the averages can be deceptive.  

In addition to the sample size considerations described above, you can get into trouble if you start with a sample that contains people who are higher or larger on average on the relevant variable that the average person in the general population.

If two tall people marry, their progeny will usually be, on average, tall compared to others in the general population. However, they are not all that likely to be taller than their parents. As Ellenberg states, “…the children of a great composer, or scientist, or political leader, often excel in the same field, but seldom so much as their illustrious parents” (p. 301). Their heredity mingles with chance environmental considerations, and pushes them back toward the population average. That is the meaning of regression to the mean.

To understand this, think about those who embark on weight loss diets. One needs to consider the fact that most people’s weight tends to fluctuate a few pounds either way depending on a lot of chance factors, such as their happening by an ice cream truck. And when are people most likely to start a diet? When their weight is at the top of their range! So by the law of averages, they are probably in many instances going to lose weight whether they diet or not. But when they do diet, guess what happens? They attribute the loss to the fantastic new diet!

I can not say for certain, but I wonder if studies on borderline personality disorders (BPD) yield misleading results because of regression to the mean. Long term follow-up studies on patients with the disorder seem to indicate that it seems to go away after a few years in a significant percentage of subjects. This finding is misleading, however, when you look closer. 

To make the BPD diagnosis, the subject needs to exhibit 5 of the 9 possible criteria. Many of the "improved" subjects merely went from 5 criteria down to 4 of them, and were therefore not diagnosed with BPD any longer. Actually, they became just what we call "subthreshold" for the disorder. Their problematic relationships, however, were still pretty much the same.

These results could mean that subjects with BPD may naturally vacillate between meeting criteria for the disorder and being subthreshold, or between exhibiting a high number of the criteria and a lower one. Which would mean that if they qualified for the diagnosis at the beginning of the long term follow-up study, a significant proportion of the long-term study subjects were at their worst. If so, the study results may indicate regression to the mean, and therefore say nothing else significant about the long term prognosis for the disorder.

Other important statistical issues the author discusses clearly and brilliantly include assumptions that two variables are related in a linear fashion when the are not (non-linearity - cause and effect relationships that are not based purely on an increase in one variable always leading to either an increase or decrease in another); torturing the data until it confesses (running multiple tests on your study data, controlling for different things, until something significant seems to pop up); and the following problem inherent in studies designed to see if two things like being married and smoking are correlated: 

"Surely the chance is very small that the proportion of married people is exactly the same as the proportion of smokers in the whole population. So, absent a crazy coincidence, marriage and smoking will be correlated, either positively or negatively."

Any one who is serious about critically evaluating the medical literature owes it to themselves to read this book.

Tuesday, September 23, 2014

Hidden Assumptions in Conclusions about Research Data in Psychology





In evaluating the conclusions of the authors from the results of any “empirical” study, two important questions one should ask oneself are: What assumptions are the authors making, and are those assumptions justified?

In today’s world, particularly in studies of the psychology of human beings, study authors often make assumptions which they do not bother to spell out in their reports, so their conclusions may seem logical. However, if they were to spell out those assumptions, everyone would immediately recognize them as completely and obviously preposterous.

In his book How Not to Be Wrong, Jordan Ellenberg mentions an illustrative anecdote about the importance of hidden assumptions that involved a group of government scientists from World War II. Their task was to determine where on warplanes to best place armor, since too much armor weighed the planes down and decreased their maneuverability. The scientists closely examined the airplanes that were returning home safely. 

At first, they inspected the planes in order to determine where the bullet holes mostly were. They figured that the parts of the plane that were hit the most often should be where the most armor should be placed, since (as the thinking went) those places must be where being hit was the most likely. Strangely, the engine seemed to be the part of the planes most frequently spared from bullet holes.

Wrong strategy. They should have been looking at where the bullet holes mostly were not. The planes hit in those places were the ones that were not making it home safely! If the engine got hit, the plane crashed. If a plane had been hit in the places they were looking at, it was apparently much less likely to crash, since it made it home. The armor should therefore be put around the engine. But only one scientist in the group made this seemingly obvious point before everyone else saw how it obvious it was! And these were some of the best minds in the field.

So let me take a study that I recently found during my weekly literature search on borderline personality disorder (BPD) on the medical database Ovid.  I’m just going to discuss the abstract, since that is all most doctors are ever going to read, if they read anything at all. (The authors did not spell out their assumptions any better in the body of the paper, but the odds are no one is going to actually read that anyway).  

Here is the abstract:

Authors:  Nicol K.  Pope M.  Sprengelmeyer R.  Young AW.  Hall J.
Title: Social judgment in borderline personality disorder.
Source: PLoS ONE [Electronic Resource].  8(11):e73440, 2013.
Abstract:
  BACKGROUND
:  Those with a diagnosis of BPD often display difficulties with  social  interaction and struggle to form and maintain interpersonal relationships.  Here we investigated the ability of participants with BPD to make social inferences from faces.
  METHOD: 20 participants with BPD and 21 healthy controls were shown a series of faces and asked to judge these according to one of six characteristics (age, distinctiveness, attractiveness, intelligence, approachability, trustworthiness). The number and direction of errors made (compared to population norms) were recorded for analysis.
  RESULTS: Participants with a diagnosis of BPD displayed significant impairments in making judgments from faces. In particular, the BPD Group judged faces as less approachable and less trustworthy than controls. Furthermore, within the BPD Group there was a correlation between scores on the Childhood Trauma Questionnaire (CTQ) and bias towards judging faces as unapproachable.
  CONCLUSION: Individuals with a diagnosis of BPD have difficulty making
  appropriate social judgments about others from their faces. Judging more faces as unapproachable and untrustworthy indicates that this group may have a heightened sensitivity to perceiving potential threat, and this should be considered in clinical management and treatment.

Now, many other studies have shown that patients with BPD are actually better at reading faces than controls, so in trying to draw any conclusions of course we have to figure out why different studies get different results. But ignoring that for the time being, let us just look at this one study abstract in isolation.

The conclusions was that the subjects with BPD had “significant impairments” and "difficulties" in making judgment. To be fair, the authors also used the words "heightened sensitivity to perceiving potential threat," which is actually a far more accurate description of their findings. But it is the words "impairments'" and "difficulties" that will be the ones that will jump out at most readers. And in the body of the paper, those terms are if fact more in line with the conclusions discussed by the authors than the phrase "heightened sensitivity."

In using these nouns, the authors are making some rather strange assumptions. A clue that they are doing that is also in the abstract: It mentions that the patients with BPD were far more traumatized as children than the controls.

That being the case, it is highly likely that the people in the social environment of the BPD subjects were far more likely to have hostile intentions than those of the controls. In such an environment, you’d have to be an idiot not to generally have a high index of suspicion when evaluating the faces of people. 

The assumption the authors seem to be making is that somehow the BPD subjects were just naturally worse at reading faces, rather than they were justifiably more suspicious of other people - the latter conclusion being one that would be predicted by error management theory.

So the assumptions they seem to making that need to be questioned are:

1        1. We can just ignore the social context of research subjects in making these sorts of judgments about people’s abilities.

          2. It is true that people rarely if ever use their brains to develop strategies for dealing with other people that have little to do with their innate abilities.

Clearly, those are really stupid assumptions.