Thursday, May 22, 2008
Announcement: Symposium on Causality, in Germany
Dear colleagues,
We would like to kindly invite you to the Symposium on Causality 2008, scheduled for July 17th to 19th in Dornburg (near Jena), Germany. The symposium brings together different traditions of analysis of causal effects (regression-based analyses, analyses based on propensity scores, analyses with instrumental variables) to discuss the state-of-the-art in the analysis of causal effects, with a special focus on non-standard designs and problems (missing data, non-compliance, multilevel designs, regression discontinuity designs).
The symposium will be structured along seven focus presentations by leading proponents in different fields of causality research. Each focus presentation will be discussed and supplemented by two invited discussants, followed by an open discussion among all participants. Focus presentations will be given by Donald B. Rubin, Thomas D. Cook, William R. Shadish, Rolf Steyer, Steven G. West, Christopher Winship and Michael E. Sobel.
There will also be ample room for participants to present and discuss their research during the symposium. Participants who want to present their research findings are asked to register for the symposium no later than June 15 and submit a title and an abstract for their presentation together with their registration. The mode of presentation (oral presentation or poster) will be determined by the organization committee depending on the total number and quality of the submissions. Other participants should register no later than June 29.
The registration fee for the symposium is 80 Euros including a daily bus transfer from Jena to Dornburg and refreshments during the conference. You can also register for the conference dinner for additional 30 Euros. To register, please visit our webpage [English, German], where you can also find additional information about the contents and structure of the symposium. If you have any questions do not hesitate to contact us.
We hope to see you soon in Jena!
Rolf Steyer and Benjamin Nagengast
Tuesday, May 6, 2008
Special Series of Articles in Developmental Psychology
The March 2008 issue of Developmental Psychology contains a special series of around 15 methodolocially and statistically oriented articles (Table of Contents). Three of the articles explicitly refer in their titles to causal inference, and others of the articles may have relevant ideas, as well. The three titles mentioning causation are as follows:
From statistical associations to causation: What developmentalists can learn from instrumental variables techniques coupled with experimental data (Gennetian, Magnuson, & Morris)
Using full matching to estimate causal effects in nonexperimental studies: Examining the relationship between adolescent marijuana use and adult outcomes (Stuart & Green)
Combining group-based trajectory modeling and propensity score matching for causal inferences in nonexperimental longitudinal data (Haviland, Nagin, Rosenbaum, & Tremblay)
At this stage, I have only skimmed through these (and other) articles in the issue. The techniques of "instrumental variables" and "matching" have, of course, been around for many years. I will be interested to see in greater depth what new contributions these articles make with such established techniques. Only within the past six months did I first hear the term "propensity score;" in skimming the many articles in the issue that use propensity scores, however, I've learned that this approach, too, has been around for decades!
Causal inference from nonexperimental data clearly is a complex, tricky endeavor. Perhaps it is for this reason that the kinds of techniques discussed in the special series have needed a quarter-century or longer to be absorbed, tested in different contexts, and diffused across disciplines.
Monday, April 7, 2008
Note from Les Hayduk
I had a look at the blog Alan provided (see below) and found this easily readable, traditional, and in some ways extremely UN-helpful. I will pick up on two of the things that seem standard, but that slant people's thinking in ways that are unhelpful, and hence where I see improvements are possible. The two matters I will take on are: experiments as the supposed benchmark/gold-standard against which SEM is to be evaluated (I doubt this), and the conditions for causality (2 of the 3 traditional conditions are wrong, the third is imprecise).
There is some substantial artificiality of comparing single experiments and single SEM studies, but I skip this for the moment, though I suspect it may eventually become an important matter.
Some failings of experiments: 1) random assignment of cases (say people) to groups should result in the groups being SIGNIFICANTLY different on 5 out of every 100 characteristics, in the long run. (SEM analysis of the experimental data can include potentially problematic variables, to see if they happen to be among the 5%, if the experimenters are not too proud to combine experiments with SEM.)
2) Experiments minimize, but do NOT statistically control for any remaining measurement error. Such control can and should be done by SEM statistics. This is NOT a feature of the experiment, but involves the statistics that could be connected to the experiment. Notice that comparing experiments and SEM is implicitly comparing two different things: the methods, and the statistics that usually go along with the methods.
SEM should be used IN CONJUNCTION WITH experimentation, so I see Alan as (possibly unknowingly) working against the helpful combining of SEM with experimentation.
Within a single experiment the mechanisms of action WITHIN the study are usually NOT well-investigated with experiments, but can be much better done in a single SEM (via inclusion of indicators of the appropriate/anticipated intervening causal structures).
Model testing is LESS well done in experiments than in SEM. SEM has an overall model test, and experiments do not usually have a comparable test (if ANOVA, or regression, or mean-differences are used as the statistical procedures). These procedures provide parameter tests parallel to those in SEM, but they have no parallel to the OVERALL MODEL TEST in SEM. Often experimenters are not even aware that they do not have an overall model test comparable to SEM's overall model test.
Enough on this for now, so I will move to the criterion for causality. Here is a quote from Alan's blog:
…contemporary SEM practitioners would probably be more comfortable with suggestions of causation if the data were collected longitudinally (more specifically, with a panel design, in which the same respondents are tracked over time). Of the three major criteria for demonstrating causality, longitudinal studies are clearly capable of demonstrating correlation and time-ordering; provided that the most plausible “third variable” candidates are measured and controlled for, the approximation to causality should be good…
Time sequence is NOT required for causation. Causes can go “both ways simultaneously” – there are such things as reciprocal causes (e.g. Rigdon, 1995, Multivariate Behavioral Research 30(3): 359-383) and variables can even cause themselves (for example, see my 1996 book chapter 3, or Hayduk, 1994, Journal of Nonverbal Behavior 18:245-260).
Correlation is NOT required. Suppressor effects can result in a variable causing another variable, and yet other parts of the causal system can counteract the causal covariance contribution, so the covariance between the variables is zero. (See Duncan, 1975, Introduction to Structural Equation Models, page 29 [equation for Greek-ro-subcript23] and realize that one term in the equation can be positive and the other of equal-magnitude yet negative.)
[Moderator’s note: See also Dean Keith Simonton’s discussion of suppressor variables and causal inference, in the posting immediately below.]
“Third variable” control should refer to MANY variable control – there can be many common causes, and many correlated causes that influence the two variables, and even reciprocal effects where the jargon of “third variables” is not quite correct. The full causal structure should be attended to, including misplacement of causally downstream variables to upstream locations. The issue here is the full proper causal specification of the model, not something connected to just third variables.
I notice you mentioned Judea Pearl's work. [Interested readers are encouraged to] have a look at the SEMNET archive for the comments Judea Pearl provided to SEMNET some years back [and] discussion of some of Pearl's work in SEM 2003, 10(2):289-311, which was designed to help SEM people understand one part of Pearl's book that directly connects to SEM.
Sunday, April 6, 2008
Note from Dean Keith Simonton
I was reading your posting when I came across the following page, where you state that there are three criteria for the inference of causality, the first being correlation. Correlation is specified as a necessary but not sufficient standard for causal inference.
This I believe is incorrect. When I teach causal modeling I emphasize a paradoxical version of the commonplace statement that "correlation does not prove causation," namely that "no correlation does not prove no causation." Both are equally true.
The problem is this: Not only can third variables generate spuriously non-zero correlations but they can also generate spuriously zero correlations. Only after we control for these attenuating effects will we discover that the (partial) correlation (or regression coefficient) is actually non-zero. Not only can this happen, but we even have a name for this consequence: suppression. Third variables that enlarge rather than reduce the association between two variables are suppressor variables.
Admittedly, suppression is often seen as something to be avoided. This is especially true when suppression yields standardized partial regression coefficients that are greater than one (or less than minus one) or when the relationship between the two variables changes sign (e.g., from a significantly positive correlation to a significantly negative beta). Often such effects can be seen as artifacts of poor measurement or design (e.g., excessive collinearity among the independent variables measured by tests with numerous shared items).
Yet it can also happen that suppression leads to a superior understanding of the underlying causal process. Sometimes the true model operates in such a fashion that it produces a zero bivariate correlation between two variables that are actually causally related. In such instances, no correlation does not prove no causation.
The example I use in class is the equation I've been developing over the years to predict the greatness assessments of US presidents.* It turns out that one of the best predictors in a 6-variable multiple regression equation is whether or not a president was assassinated while in office. Yet assassination does not have a significant zero-order correlation. How can this be? Well, another major predictor of leader performance is duration of tenure in office, and this variable quite understandably has a negative correlation with assassination. On the average, assassinated presidents have shorter tenures. So the positive association between tenure duration and the global leadership assessment masks the positive impact of assassination. Only when both are put into the same equation will the causal impact of assassination emerge. In addition, the predictive power of tenure duration is increased because its true effect size is no longer obscured by assassination. In the literature, this is sometimes called "cooperative suppression" (a term that seems inappropriate in the current example!).
I could give other empirical illustrations, but the foregoing should suffice. Two variables can have a causal relation even in the absence of a non-zero correlation. Zero-order correlations can be spuriously small as well as spuriously large. This outcome is especially likely in the complex causal networks that likely underlie real-world phenomena. Hence, the three conditions for causal inference from correlational data are misspecified. They probably reduce to two: temporal priority and a non-zero correlation after controlling for all reasonable third variables.
*The original 6-variable equation was published in Simonton, D.K. (1986). Presidential personality: Biographical use of the Gough Adjective Check List. Journal of Personality and Social Psychology, 51, 149-160. An update of the entire research program will appear in Simonton, D.K. (in press). Presidential greatness and its socio-psychological significance: Individual or situation? Performance or attribution? In C. Hoyt, G. R. Goethals, & D. Forsyth (Eds.), Leadership at the crossroads: Psychology and leadership (Vol. 1). Westport, CT: Praeger.
Monday, March 31, 2008
Modeling Causation with Non-Experimental Data (Part II)
Following up on Part I, we now look at a recent empirical study whose discussion provides what I consider a particularly thoughtful exposition on causal inference with longitudinal survey data.
As someone who studies adolescent drinking and how it may be affected by parental socialization, I took notice of an article by Stephanie Madon and colleagues entitled “Self-Fulfilling Prophecy Effects of Mothers’ Beliefs on Children’s Alcohol Use…” (Journal of Personality and Social Psychology, 2006, vol. 90, pp. 911-926).
The substantive findings of longitudinal relations between mothers’ beliefs at earlier waves and children’s drinking at later waves were interesting. However, what really stood out to me was the thoughtful discussion near the end of the article about which criteria for causality the longitudinal correlational design could and could not satisfy, and what other lines of argument could be marshaled on behalf of causal inferences. Within the larger Discussion section, Madon and colleagues’ consideration of causal inference appears under the heading “Interpreting Results From Naturalistic Studies.”
Madon et al.’s discussion notes the major thing their design does: “…rule out the possibility that the dependent variable exerted a causal influence on the predictor variable because measurement of the predictor variable is temporally antecedent to changes in the dependent variable.” It also acknowledges the major limitation of the design, possible unmeasured third variables: “…the potential omission of a valid predictor raises the possibility that mothers based their beliefs on a valid predictor of adolescent alcohol use that was not included in the model. If this were to occur,… [mothers’] self-fulfilling effects would be smaller than reported.”
The authors then proffer “several reasons why we believe that the self-fulfilling prophecy interpretation is more compelling” than a third-variable account.
First, Madon and colleagues claim, their preferred interpretation is consistent with established experimental findings of self-fulfilling prophecies. “Although this convergence does not prove that our results reflect self-fulfilling prophecies, confidence in the validity of a general conclusion increases when naturalistic and experimental studies yield parallel findings.”
Second, the authors cite the breadth of “theoretically and empirically supported predictors” they used as control variables. This step, they claim, reduces the chance of spurious, third-variable causation.
Third, the authors note that maternal beliefs exerted greater predictive power over time, whereas the non-maternal predictors explained less variance over time. The extension of this argument, as I understand it, would then be to ask, How likely is it that any as-yet-unmeasured, potential third variable tracks perfectly with maternal beliefs and not at all with the predictors that served as control variables?
Both by their empirical findings and their citation of the broader adolescent drinking (in terms of their control variables) and self-fulfilling prophecy (for its experimental foundation) literatures, Madon and colleagues provide a great deal of evidence that is consistent with a parenting effect.
Depending on one’s perspective, to use a football analogy, one might conclude that the Madon et al. study has advanced the ball to the opponents’ 40-yard line, or 20-yard line, or even 1-yard line. But most would probably agree that it takes a true experiment to get the ball into the end zone.
Wednesday, March 26, 2008
Modeling Causation with Non-Experimental Data (Part I)
For more than a decade, I've taught a course on structural equation modeling (SEM). The technique, nearly always used with survey data (although in theory also applicable to experimental data), involves drawing shapes to represent variables and arrows between the shapes to represent the researcher's proposed flow of causation between the variables. Regression-type path coefficients are then generated to assess the strength of the hypothesized relations. Though the unidirectional arrows in the diagrams and the tone often used in reports of SEM analyses (e.g., "affected," "influenced," "led to") imply causality, such an inference cannot, of course, be supported with non-experimental data. This is especially true of cross-sectional data (i.e., where all variables are measured concurrently). As will be discussed later, causal inferences can be supported to a greater degree -- though not completely -- with longitudinal data.
It is in this context that I read the chapter on SEM in Judea Pearl's (2000) book Causality: Models, Reasoning, and Inference. Pearl, a UCLA professor with expertise in artificial intelligence, logic, and statistics, writes at a level that is, quite frankly, well over my head. I did find his SEM chapter relatively accessible, however, so that I will discuss.
Pearl's apparent thesis in this chapter is that contempory SEM practitioners are too quick to dismiss the possibility of being able to draw causality from the technique (much like I did in my opening paragraph). Early on, in fact, Pearl writes that, "This chapter is written with the ambitious goal of reinstating the causal interpretation of SEM" (p. 133).
Pearl reviews a number of writings on SEM that he feels, "...bespeak an alarming tendency among economists and social scientists to view a structural equation as an algebraic object that carries functional and statistical assumptions but is void of causal content" (p. 137). He further notes:
The founders of SEM had an entirely different conception of structures and models. Wright (1923, p. 240) declared that "prior knowledge of the causal relations is assumed as prerequisite" in the theory of path coefficients, and Haavelmo (1943) explicitly interpreted each structural equation as a statement about a hypothetical controlled experiment. Likewise, Marschak (1950), Koopmans (1953), and Simon (1953) stated that the purpose of postulating a structure behind the probability distribution is to cope with the hypothetical changes that can be brought about by policy.
An interpretation of the above paragraph that would make the causal interpretation of SEM defensible, in my view, would be as follows: If the directional relations one is modeling with non-experimental (survey) data have previously been demonstrated through experimentation, then SEM can be a useful tool for estimating quantitatively how much of an effect a policy change could have on some outcome criterion. It's when we start talking about "assumed" or "hypothetical" experimental support that things get dicey.
As I noted above, however, contemporary SEM practitioners would probably be more comfortable with suggestions of causation if the data were collected longitudinally (more specifically, with a panel design, in which the same respondents are tracked over time). Of the three major criteria for demonstrating causality, longitudinal studies are clearly capable of demonstrating correlation and time-ordering; provided that the most plausible "third variable" candidates are measured and controlled for, the approximation to causality should be good (these latter two linked documents are from my research methods lecture notes).
MacCallum and colleagues (1993, Psychological Bulletin, 114, 185-199), while acknowledging some limitations to longitudinal studies, succinctly summarize why they're useful: "When variables are measured at different points in time, equivalent models in which effects move backward in time are often not meaningful" (p. 197).
In the upcoming Part II, we highlight a recent empirical study whose discussion provides a particularly thoughtful exposition on causal inference with longitudinal survey data.
Friday, March 7, 2008
Taking Advantage of Random Processes in the Real World
As is taught in beginning research methods courses, the only technique that allows for causal inference is the true experiment and the linchpin of experimentation that permits such inference is random assignment. In the social and behavioral sciences, experiments typically are conducted in university laboratories, with established subject pools to ensure the availability of participants.
Scientists lacking such resources may thus have a hard time conducting experiments, even if they wanted to. Others may simply develop a preference for survey research or other non-experimental methods such as archival research and content analysis; such techniques generally cannot be used to assess causality, but offer potential advantages in terms of mapping onto more natural, realistic situations encountered in daily life. I myself, for whatever reason, gravitated to survey research over the years, even though laboratory experimentation was a major part of my graduate-school experience.
Now, however, even scholars lacking an affinity for the lab may have opportunities to address causation in their research. In what appears to be a growing trend, clever researchers are noticing random assignments in real-world settings and seizing upon them to conduct causal studies from afar.
As one example, University of Durham anthropologists Russell Hill and Robert Barton realized that in Olympic "combat" sports such as boxing and wrestling, competitors are assigned at random to wear either red or blue outfits. The finding that red-clad participants won more often than did their blue counterparts can thus be interpreted causally (a showing that outfit color has some causal effect does not necessarily mean that it's a large effect).
As another example, readers of the 2005 book Freakonomics (by Steven Levitt and Stephen Dubner) may remember Levitt and colleagues' drawing upon the Chicago Public Schools' use of random assignment in the district's school-choice program when more students wanted to go to a particular school than could be accommodated. As Levitt and Dubner wrote (p. 158):
In the interest of fairness [to applicants of the most competitive schools], the CPS resorted to a lottery. For a researcher, this is a remarkable boon. A behavioral scientist could hardly design a better experiment in his laboratory...
The random aspect thus allowed the researchers to make the causal conclusion that:
...the students who won the lottery and went to a "better" school did no better than equivalent students who lost the lottery and were left behind.
In yet another example, Levitt's fellow University of Chicago faculty member, law professor Cass Sunstein, seized upon the fact that appellate cases within the federal circuit courts are heard by random three-judge panels from the larger pool of judges within each geographic circuit. Does a judge appointed by a Democratic (or Republican) president show variation in how often he or she votes on the bench in a liberal or conservative direction, depending on whether he or she is joined in a case by judges appointed by presidents of the same or opposing party? That is the type of question Sunstein and his colleagues can answer (here and here).
Yale law professor (and econometrician) Ian Ayres, whose 2007 book Supercrunchers I recently finished reading, refers to this phenomenon as "piggyback[ing] on pre-existing randomization" (p. 72). The book contains additional examples of its use.