Monday, August 13, 2012

Causal Mediation Conference in Belgium

The Center for Statistics at Belgium's Ghent University will present the symposium "Causal Mediation Analysis" on January 28-29, 2013, at Het Pand, Gent, Belgium. Further information is available here. According to a notice on the event, "This meeting aims to bridge the gap between traditional mediation analysis building on the famous work of Baron and Kenny and state-of-the-art causal mediation analysis. The idea is to discuss recent developments between methodological researchers and to bring them to the wider research community in a non-technical way."

Sunday, March 25, 2012

Michael Nielsen Offers (Relatively) Accessible Explanation of Pearl's "Causal Calculus"

by Alan

Judea Pearl, a UCLA professor of computer science, is one of the world's leading thinkers -- if not the leading thinker -- on conceptual approaches to causal inference. He is author of the book Causality and of numerous articles and presentations. He also operates the UCLA Causality Blog, a link to which appears in the left-hand column of the present page. On top of all this, Pearl recently garnered the Association for Computing Machinery (ACM) Turing Award for his contributions to artificial intelligence.

"Accessible" is not a word I would use to describe Pearl's writings, however. I have previously described the level of Pearl's writing as "quite frankly, well over my head." Heavy with logic symbols, Pearl's texts would, I suspect, challenge even many well-educated students of causality.

Fortunately for those of us seeking greater understanding of Pearl's ideas, Michael Nielsen has written an article trying to explain Pearl's "causal calculus" to a wider audience. I couldn't understand everything Nielsen wrote, but in relative terms, I found his exposition easier to grasp than Pearl's.

Fairly early on, Nielsen introduces the familiar example of smoking and lung cancer to discuss what conclusions can be drawn from correlational (observation) vs. randomized-controlled research designs (he seems to use the word "experimental" generically for any empirical investigation, specifying with terms such as "intervention" or "randomized controlled" when he means that participants are randomly assigned to conditions). Noting that human participants cannot ethically be randomly assigned to smoke cigarettes, Nielsen tantalizes the reader as follows:

We’ll see that even without doing a randomized controlled experiment it’s possible (with the aid of some reasonable assumptions) to infer what the outcome of a randomized controlled experiment would have been, using only relatively easily accessible experimental data, data that doesn’t require experimental intervention to force people to smoke or not, but which can be obtained from purely observational studies.

The main points I gleaned from Nielsen's piece were that (a) we can learn more than I previously thought simply from diagramming hypothetical causal relations between variables as in structural equation modeling or path analysis; and (b) one's conceptual model can be translated into conditional probability statements (i.e., given x, what is the probability of y) that potentially can be manipulated to answer causal questions without a randomized experiment. As Nielsen explains:

...Pearl had what turns out to be a very clever idea: to imagine a hypothetical world in which it really is possible to force someone to (for example) smoke, or not smoke. In particular, he introduced a conditional causal probability p(cancer|do(smoking)), which is the conditional probability of cancer in this hypothetical world. This should be read as the (causal conditional) probability of cancer given that we “do” smoking, i.e., someone has been forced to smoke in a (hypothetical) randomized experiment.

Now, at first sight this appears a rather useless thing to do. But what makes it a clever imaginative leap is that although it may be impossible or impractical to do a controlled experiment to determine p(cancer|do(smoking)), Pearl was able to establish a set of rules – a causal calculus – that such causal conditional probabilities should obey. And, by making use of this causal calculus, it turns out to sometimes be possible to infer the value of probabilities such as p(cancer|do(smoking)), even when a controlled, randomized experiment is impossible.


Returning to the lung-cancer example, it is theoretically possible that smoking leads directly to lung cancer or that an unobserved third variable causes both smoking and lung cancer (also, lung cancer may cause people to begin smoking, but that seems implausible). As Nielsen discusses, we can insert a fourth variable, namely particulate lung residue ("tar"), between smoking and lung cancer in the proposed causal sequence. This inclusion helps us partially break the connection between the hidden third variable and the other variables. Argues Nielsen: "But if the hidden causal factor is genetic, as the tobacco companies argued was the case, then it seems highly unlikely that the genetic factor caused tar in the lungs, except by the indirect route of causing those people to smoke."

Through manipulations such as the above: "the causal calculus lets us do something that seems almost miraculous: we can figure out the probability that someone would get cancer given that they are in the smoking group in a randomized controlled experiment, without needing to do the randomized controlled experiment. And this is true even though there may be a hidden causal factor underlying both smoking and cancer."

Ultimately, the manipulation of equations can lead to a formula to estimate the conditional probability of developing cancer given random assignment to a smoking condition, p(cancer|do(smoking)), as a function of "quantities which may be observed directly from experimental data, and which don’t require intervention to do a randomized, controlled experiment" (see Equation 5 in Nielsen's article). For any given problem, such non-intervention-based probabilities to plug into the equation may or may not be available.

Nielsen concludes the article by exploring possible future directions in the study of causality. For those interested in causal inference without randomized-controlled studies, Nielsen's article is a must-read.

Tuesday, August 30, 2011

Judea Pearl and colleagues have launched the Journal of Causal Inference, to be published by the Berkeley Electronic Press.

Sunday, August 7, 2011

Causal-Inference References from SEMNET

by Alan

Over at the Structural Equation Modeling discussion listserve (SEMNET), participants lately have recommended several recent articles and resources on causal inference (and related topics) with non-experimental (correlational) research designs. For benefit of the larger research community, I have listed these materials below.

I've looked over some of these articles and they seem to vary in the prior training assumed. Some would seem accessible for social scientists without elaborate mathematical training, whereas others refer extensively to more sophisticated math (e.g., matrix algebra). The Antonakis et al. piece, in particular, appears to provide a (mostly) non-technical overview.

Three specific topics are covered in many of the articles:

  • Omitted-variable bias (or specification error, more generally), which is all-important to causal inference, due to the "third-variable" issue.


  • Propensity-score modeling, which already has been discussed extensively on this blog (e.g., here, here, and here).


  • The use of instrumental variables.


  • I am particularly impressed by the range of academic discliplines from which these articles arise. Thanks to those who contributed these items to SEMNET!

    ----------------------------------------------------------------------------------

    Antonakis, J., Bendahan, S., Jacquart, P., & Lalive, R. (2010). On making causal claims: A review and recommendations. The Leadership Quarterly, 21, 1086-1120.

    Austin, P.C. (2011). A tutorial and case study in propensity score analysis: An application to estimating the effect of in-hospital smoking cessation counseling on mortality. Multivariate Behavioral Research, 46, 119-151.

    Beguin, J., Pothier, D., & Côté, S.D. (2011). Deer browsing and soil disturbance induce cascading effects on plant communities: a multilevel path analysis. Ecological Applications, 21, 439–451.

    Bollen, K.A., Kirby, J.B., Curran, P.J., Paxton, P.M., & Chen, F. (2007). Latent variable models under misspecification: Two-Stage Least Squares (2SLS) and Maximum Likelihood (ML) estimators. Sociological Methods & Research, 36, 48-86.

    Bollen, K.A., & Bauer, D.J. (2004). Automating the selection of model-implied instrumental variables. Sociological Methods & Research, 32, 425-452.

    Bollen, K.A., & Maydeu-Olivares, A. (2007). A polychoric instrumental variable (PIV) estimator for structural equation models with categorical variables. Psychometrika, 72, 309-326.

    Clarke, K. (2005). The Phantom Menace: Omitted variable bias in econometric research. Conflict Management and Peace Science, 22, 341-352.

    Clarke, K. (2009). Return of the Phantom Menace: Omitted variable bias in econometric research. Conflict Management and Peace Science, 26, 46-66.

    Coffman, D.L. (2011). Estimating causal effects in mediation analysis using propensity scores. Structural Equation Modeling, 18, 357-369.

    Freedman, D.A., Collier, D., Sekhon, J.S., & Stark, P.B. (Eds.). (2009). Statistical models and causal inference: A dialogue with the social sciences. Cambridge University Press. ISBN: 978-0521123909.

    Frosch, C.A., & Johnson-Laird, P.N. (2011). Is everyday causation deterministic or probabilistic? Acta Psychologica, 137, 280-291.

    Hancock, G. R., & Harring, J. R. (2011, May). Using phantom variables in structural equation modeling to assess model sensitivity to external misspecification. Paper presented at the Modern Modeling Methods conference, Storrs, CT. (Hancock webpage to request copy.)

    Hoshino, T. (2008). A Bayesian propensity score adjustment for latent variable modeling and MCMC algorithm. Computational Statistics & Data Analysis, 52, 1413-1429.

    Kirby, J.B., & Bollen, K.A. (2009). Using instrumental variable (IV) tests to evaluate model specification in latent variable structural equation models. Sociological Methodology, 39, 327–355. (Public copy)

    Mahoney, J. (2008). Toward a unified theory of causality. Comparative Political Studies, 41, 412-436.

    Markus, K.A. (2011). Mulaik on atomism, contraposition and causation. Quality and Quantity. Online First (subscription needed), http://www.springerlink.com/content/r754405614228w0v/

    Markus, K.A. (2011). Real causes and ideal manipulations: Pearl's theory of causal inference from the point of view of psychological resarch methods. In P. McKay Illari, F. Russo & J. Williamson (Eds.),Causality in the sciences (pp. 240-269). Oxford, UK: Oxford University Press. (Errata)

    Pearl, J. (2010). On a class of bias-amplifying variables that endanger effect estimates. Technical Report R-356. In P. Grunwald & P. Spirtes (Eds.), Proceedings of UAI, 417-424. Corvallis, OR: AUAI.

    Pearl, J. (2011, August). The causal foundations of structural equation modeling. UCLA Cognitive Systems Laboratory, Technical Report (R-370), http://ftp.cs.ucla.edu/pub/stat_ser/r370.pdf. Chapter for R. H. Hoyle (Ed.), Handbook of structural equation modeling. New York: Guilford Press.

    Shadish, W.R., & Steiner, P.M. (2010). A primer on propensity score analysis. Newborn & Infant Nursing Review, 10, 19-26.

    Shipley, B. (2009). Confirmatory path analysis in a generalized multilevel context. Ecology, 90, 363-368.

    Shipley, B. The Causal Toolbox: A collection of programs for testing or exploring causal relationships [website]. http://pages.usherbrooke.ca/jshipley/recherche/book.htm

    Spector, P.E., & Brannick, M.T. (2011). Methodological urban legends: The misuse of statistical control variables. Organizational Research Methods, 14, 287-305.

    Steiner, P.M., Cook, T.D., Shadish, W.R., & Clark, M.H. (2010). The importance of covariate selection in controlling for selection bias in observational studies. Psychological Methods, 15, 250-267.

    Thoemmes, F.J., & Kim, E.S. (2011). A systematic review of propensity score methods in the social sciences. Multivariate Behavioral Research, 46, 90-118.

    Tuesday, February 1, 2011

    Sunday, December 19, 2010

    Correlation, Causality, and Parenting Studies

    Alan and Bo are both quoted in this new article from Brain, Child magazine. In addition to discussing substantive issues of causal inference, author Katy Read also probes the chain of diffusion of social-science research.

    The chain begins, of course, with the investigators who conducted the research. Those who conduct correlational studies typically include a statement of limitations at the end of their articles, noting that the findings are open to alternative causal interpretations. In less-guarded moments, however, even research scientists will use phraseology that implies a preferred causal direction.

    Universities, research institutes, and/or professional organizations may then issue press releases on a particular study. Ultimately, a research finding may make it into the media. At each reporting step removed from the (methodologically trained) scientific investigators, therefore, statements of caution regarding causality are less likely to appear.

    Saturday, November 27, 2010

    Studying Personality as a Causal Agent

    by Alan

    Angela Lee Duckworth, Eli Tsukayama, and Henry May have published an article entitled, "Establishing Causality Using Longitudinal Hierarchical Linear Modeling: An Illustration Predicting Achievement From Self-Control" (abstract) in the October 2010 issue of the new journal Social Psychological and Personality Science.

    The authors are interested in the personality trait of self-control (which encompasses such abilities as persistence and delay of gratification) and whether it can actually be shown to cause academic achievement (GPA) during the years from fifth to eighth grade. Duckworth and colleagues acknowledge the difficulty early on:

    "Manipulating personality in a random-assignment experiment could, in theory, establish its causal role for later outcomes but, alas, personality is not easily manipulated. To our knowledge, no empirical investigation to date has successfully manipulated trait-level self-control and measured subsequent effects on life outcomes" (p. 311).

    The authors also carefully review the arguments for why longitudinal predictive models using path-analysis and structural equation modeling (controlling for prior levels of the dependent variable), though an improvement over cross-sectional correlations, are still vulnerable to unmeasured third variables. "Longitudinal growth-curve modeling using hierarchical linear models (HLM)" is then presented as offering "a partial solution to the third variable problem" (p. 312).

    The core of the argument appears to be this: "By treating a predictor as a time-varying covariate in the prediction of trajectories, one can rule out the possibility of all time invariant confounds (e.g., relatively stable variables such as socioeconomic status). Specifically, if short-term changes in a predictor predict subsequent short-term changes in achievement, a confounding variable z would have to predict these changes and also be tightly yoked to changes in the predictor over time (i.e., the confound and predictor would have to go up and down together in synchrony over time)" (p. 312).

    The authors indeed found self-control to predict GPA longitudinally (and not the reverse), concluding as follows:

    "This longitudinal HLM study illustrated an innovative analytic strategy that effectively controlled for all time-invariant third-variable confounds... What our analyses did not rule out, however, is the possibility of an unmeasured time-varying third variable that changes in sync with self-control and causally determines subsequent academic performance" (p. 316; my emphasis added).

    I would recommend this article for its excellent exposition on causality with non-experimental designs and for the approach it demonstrates. There are a lot of technical details in the article, as well, and even readers experienced with some of the general techniques used may find themselves having to pause periodically to wrap their minds around each specific analytic decision the authors describe.