Jeromy Anglim's Blog: Psychology and Statistics


Saturday, September 19, 2009

Introduction to Journal Article Deconstruction

One of the most powerful strategies that I use to learn how to write journal articles is to consciously study the writing conventions of good journal articles. I often want to communicate this strategy to other researchers who are battling the process of writing research. This is particularly the case with results sections. Thus, I plan to post various case studies using this approach to demonstrate how it works. Perhaps the principles generated from the deconstruction will also be relevant to others. I'll post all such instances with the label Article Deconstruction.

Friday, September 18, 2009

Variable Importance and Multiple Regression

Many researchers are interested in questions related to the relative importance of a set of predictors in multiple regression. This is important to both consultants and academics. I assume the motivation derives from the assumption (typically wrong at least to some extent) that the predictors flagged as important will have  larger causal effects and are therefore better targets for manipulation in an  intervention. Some examples include: a) a set of risk factors on clinical symptoms in psychology; b) a set of personality measures on performance; c) a set of beliefs on overall attitude.

R Community in Australia

One of the nice aspects of R is the community of users that has built up around it. The open-source model seems to create an orientation of sharing and contribution. Users benefit from R and then they give back in the form of new packages, free documentation, blogs, presentations, and so on.

Thursday, September 17, 2009

Tuesday, September 15, 2009

Setting up a Blog on Blogger

Given the number of academics in the world, there are surprisingly few blogs on psychology and research methods. There are many possible reasons. Two barriers are: 1) lack of knowledge of the ease of creating a blog; and 2) lack of knowledge of the benefits of having a blog.

The details below set out my setup for my blog account and my blogging statistics. When I set it up originally, I did look into the various options in terms of blogging providers and so on. I make no claim to my choices being optimal for me or other people. But I have found them more than adequate for my purposes. In particular, usage statistics (and comments) are a great form of feedback that is not necessarily available in other forms of academic communication. For further discussion of the benefits of blogging and related technologies in academic, Gideon Burton provides a great exposition.

Confidence Intervals and Correlations

In a previous post, I discussed the various scenarios for running significance tests on correlations.

A researcher recently asked me how to calculate confidence intervals for two correlations that share a common variable (i.e., dependent correlations).

Thursday, September 10, 2009

Pen and Paper

For a long while I did almost all my thinking and writing either in my head or on the computer. More and more lately I find myself returning to pen and paper.
Examples:

Wednesday, September 9, 2009

Experiments with a mixture of repeated measures and between subjects factors

I often speak to researchers who need to analyse an experiment with a combination of between subjects and repeated measures factors.
Some Examples:
  • 5 x 3 Design: 5 levels of task type (repeated measures); and 3 levels of group (between subjects)
  • 2 x 2 x 2 Design: 2 levels of order (between subjects); by 2 levels of instructions (between subjects); by 2 levels of task feature (repeated measures)
This post directs such researchers to some resources on the web.

Resources:
  • UCLA has several examples of how examine such designs using SPSS Repeated Measures ANOVA; It also talks about how to test contrasts and run follow up test more generally.
  • Andy Field provides a gentle introduction to repeated measures ANOVA using SPSS.

Tuesday, September 8, 2009

Cluster analysis and single dominant factors

I often chat with researchers wanting to use cluster analysis to group cases. I just wanted to point out a common scenario where cluster analysis may not be a good way of proceeding.

Monday, September 7, 2009

Logistic Regression Resources in SPSS or R

Question: A researcher asked me: "What resources are available for running and interpreting a logistic regression?"

Significance Tests on Correlations

OVERVIEW: I often speak to researchers wanting to compare the significance of two correlations. The two scenarios most commonly encountered are: 1) comparing dependent correlations; and 2) comparing independent correlations.

Wednesday, September 2, 2009

Repeated Measures Experiments with Many trials in SPSS (PASW)

I was recently talking to a researcher who was in the process of analysing experimental data based on a 2 x 2 x 2 x 4 repeated measures design. Each combination of the levels of the repeated measures factors involved five trials. The dependent variable was reaction time. The researcher had their data laid out in a one person per participant format (wide format). I had the following advice for the researcher:

Data Format:
Create a long format data file called “trials” where each row is the combination of one participant and one trial. And have a separate data file called “subjects” that contains one row per participant and includes data on participants that is constant throughout the experiment (e.g., gender, age, personality measures, etc.).

Tuesday, August 18, 2009

Social Network Analysis Resources for R

Social Network Analysis is an increasingly popular tool for modelling dependence structures between social actors. In my department researchers are developing new models for representing such dependence structures (MELNET). In 2007 I gave a talk on my consulting experience using social network analysis to provide insights on team dynamics. Since then I have switched to mainly using R for analysing social network datasets.

Monday, August 17, 2009

Selecting University Students: Perspectives from Selection and Recruitment

The discussion paper mentions other means for selecting students in addition to their performance in the final year of high school (ENTER).
I see the issue as involving two elements:

  1. What criterion of an effective selection system does the university want to use?
  2. What measurement system can be put in place to maximise this criterion?
The first issue is a matter of values and policy. The second issue is empirical.
From The Age article I gather that there is concern for: representation from disadvantaged student groups and academic potential.
Many other values could be mentioned, including transparency, procedural justice, cost of administration, and so on. No doubt any decision would be a synthesis of such concerns. However, if the choice of criterion could be resolved, the question of a measurement system is largely an empirical question.

Wednesday, August 5, 2009

My Procedure for Upgrading R: Windows XP with StatET

R releases upgrades around every 6 months. I run R on Windows XP using the StatET plug-in and Eclipse. This is my current upgrade procedure. I've posted this mainly for my own future reference. But who knows? it might be useful for someone else:

Friday, June 12, 2009

Redmond Barry Building Views: 4 Seasons in One Day

A series of pictures from my office window.

Saturday, June 6, 2009

Upcoming Conferences

I'll be attending a couple of conferences in Europe in early to mid July, 2009. So, if anyone who reads this blog will be attending, feel free to say hi.
  1. 11th European Congress on Psychology in Oslo: I'll be presenting a talk: "The effect of warnings on personality test faking in employee selection"
  2. Directions in Statistical Computing in Copenhagen

Normality and Transformations: A few thoughts

A lot of researchers ask me about normality and transformations. This post sets out some my advice on assessing normality and dealing with violations of the assumption.

Learning R for Researchers in Psychology

R is a powerful open source environment for statistical computing. This post provides a selective list of resources for getting started with R including thoughts on books, online manuals, blogs, videos, user interfaces, and more. At the end of the post are some R resources specific to researchers in psychology. (UPDATED 4th May 2011)

Friday, May 29, 2009

Pronunciation Guides for Mathematical Notation, Expressions, and Greek Letters

When doing research in psychology you are sometimes required to study new statistical or mathematical techniques on your own. However, mathematical books rarely tell you how to pronounce the mathematical symbols. And even if you know how to pronounce the symbols in isolation, this does not guaranty that you can pronounce a mathematical expression made up of multiple symbols. Being able to read mathematical notation is a basic first step in aiding memory and conceptual understanding. The following links provide resources on reading mathematical symbols and provide a good reference if you encounter symbols and expressions with which you are not familiar.

Mathematics Pronunciation Guides

  • VÄliaho's guide to Pronunciation of Mathematical Expressions: This is the place to start. It covers many important rules in a 3 page document
  • Handbook for Spoken Mathematics: If VÄliaho's guide did not meet your requirements, check out this extensive resource. It covers many major branches of mathematics such as logic and set theory, geometry, statistics, calculus, and linear algebra. It is the most comprehensive guide that I have found with around 100 pages and around 500 symbols with pronunciation. I'd recommend studying all the symbols if mathematical pronunciation is an issue for you. The symbols are distributed over many pages making it a little difficult to look up a single symbol of interest. Also, when a choice exists, the guide often chooses a more verbose and less ambiguous form of pronunciation. For example, it suggests for "x_i", "x sub i" instead of "x i". This emphasis on unambiguous verbal communication is sometimes more than required when verbalising the symbols in your head or when verbalising symbols in a context where the actual symbolic math is also displayed.
  • RPI's Saying Mathematics Guide
  • Oanca et al's Reading Mathematical Expressions
  • Wikipedia guide to mathematical symbols: meaning of common mathematical symbols with links to their meaning.
  • Greek letters: Lower and upper case Greek letters with pronunciation
  • Tips on displaying formulas can even be useful for some obscure mathematical symbols
The other option is to actually listen to mathematics lecturers and assume that the pronunciation will rub off: see my earlier post on free online mathematics video courses.

Books on mathematical pronunciation

  • Lawrence Change (1983). Handbook for Spoken Mathematics: (Larry's Speakeasy).

Related Posts

Thursday, May 28, 2009

Introduction to Statistical Modelling in Psychology: NSS Presentation

Today I gave an introductory talk for the Neuropsychological Students’ Society at the University of Melbourne on the topic of Statistical Modelling in Psychology.

The slides with notes from the talk are available for download at the following link: Introduction to Statistical Modelling in Psychology.

Friday, May 22, 2009

Bootstrapping and the boot package in R

I was recently asked about options for bootstrapping. The following post sets out some applications of bootstrapping and strategies for implementing it in R. I've found bootstrapping useful in several settings:

  • where the statistic I'm interested in is a little unusual: the average R-square across five separate regressions; the difference in the average correlation of a set of variables between two groups
  • non parametric statistics, such as the median
  • when assumptions such as normality of homoscedasticity are not satisfied

Thursday, May 21, 2009

Self-Archiving of journal articles in academia

I'm a big fan of academics who post copies of their journal articles on their own website (e.g., here and here and here). Even if my university has online subscriptions to most journals, it is just more convenient to access all of a person's work in one place. Such articles are also typically then available through Google Scholar. It also opens the work up to others outside of academia who do not have access to university journal subscriptions. Given that academic work is typically directly or indirectly funded by public money, this seems particularly important.

I have found Sherpa Romeo to be a useful site for looking up the standard publisher copyright transfer agreement for different journals. This can be useful in selecting a journal outlet that allows for publishing on your own website. It also sets out the conditions for putting journal articles on your own website.

See here for a discussion of author rights.

Wednesday, May 20, 2009

Endnote Collaboration

I'm currently exploring options for using Endnote for collaboration:
UQ has a nice discussion of strategies for collaborating with Endnote.
I'm also considering using the Endnote web option.

Friday, May 15, 2009

Statistics for a Psychology Thesis

In 2007 I presented a talk to postgraduate psychology students at The University of Melbourne. As part of the talk I produced a handout which summarised many of the key points that I felt were relevant for such an audience who needed to complete a thesis involving quantitative analysis. Reading over it two years later, I still agree with the ideas, even if my understanding may be a little more nuanced. For example, I'd now see meta-analytic thinking as a simple version of Bayesian statistics. Anyway, I thought I'd post it on the blog.
The audio (17MB) for the talk is available online, as are the Slides, and a PDF version of the content below.

Wednesday, May 13, 2009

Online Mathematics Video Courses For Self Study in Psychology

This post provides links to free online video courses on mathematics and statistics.

Verifying How Composite Scores Have Been Computed

This post sets out a procedure for checking how a composite was computed from items. This is particularly useful when working with self-report psychological scales and composite ability tests. Computing total scores for a series of subscales and total scores on personality tests and other self-report inventories can be fiddly business, and errors often arise.

Wednesday, March 25, 2009

Calculating Composite Scores of Ability and Other Tests in SPSS

Researchers in psychology often have a large number of variables. Science aims for parsimonious explanations of the world. Thus, the challenge is to develop a principled approach to dealing with the multiplicities that arise in psychological research. One common approach is to combine tests that measure similar things into composites. This post looks at how to form composites. The emphasis is on settings where you have multiple ability tests and you want to create a composite ability factor. Also, particular emphasis is given to how to do it in SPSS.

Monday, March 23, 2009

Inquisit: Simple Reaction Time, Four Choice Reaction Time, and Typing Test

In 2007 I was looking for a tool for running psychological experiments online. It was important that the online tool could record reaction time. After a little exploration I started using Inquisit.

I’ve attached a script that I wrote which provides a basic measure of simple reaction time, four-choice reaction time, and a typing test. I thought I’d make the script available for others to use and modify. The program was written for Inquisit 3.0. The Inquisit website has a pile of additional scripts. It also has a fully functional trial-version of the software.

Click on the link below to download the script. Change the file extension to ".exp" to associate with Inquisit.

Saturday, March 14, 2009

Photos from My Office - The University of Melbourne




For all its quirks, the Redmond Barry Building provides some very nice views.

User Interface for R: StatET and Eclipse

R is tremendously flexible. It’s flexible in how commands can be written, and it’s flexible in the user interface that can be used to run it. I’ve been playing around with different user interfaces on Windows for a while now.

Thursday, March 12, 2009

Saving Citation Information from Google Scholar

I have often wanted to be able to export citations from Google Scholar into Endnote. While Google Scholar makes it easy to do this for single articles, it does not at present permit you to export an entire search.
I’ve found a work around. Anne-Wil Harzing’s Publish or Perish Software (free) searches Google Scholar and allows you to export a file that contains the ciation information of the entire search. This can then be imported into Endnote. It also exports in Bibtex and CSV format. 
It’s also a great way to assess the impact and citation rates of individual scholars, journals, and particular articles.

Friday, February 27, 2009

Books on Writing, Grammar, and Style in Psychology and Beyond

This post lists some books on academic writing, which I've found useful.

Thursday, February 19, 2009

Assorted issues with Multiple Regression and Moderation

1. Do variables in moderation have to be normally distributed before centering? 
No, normality is not a requirement. However, you could apply a transformation if you felt it was justified, such as when it better reflects the meaning of the variable.
See the following for a discussion of issues in transformations: 

Assumptions in moderation are the same as in normal multiple regression. They mainly concern the pattern of the residuals. For a discussion see: http://www.statisticshell.com/multireg.pdf

2. Do I use the R squared or adjusted R squared when reporting a multiple regression?
You can use either or both. The larger your sample, the smaller the difference is between adjusted R square and R-squared. R-squared is a biased estimate.

3. p=0.52, is that still significant?
Jacob Cohen once said something along the lines that surely God must love .06 almost as much as .04. 
You can write p = .05
Whether this is statistically significant depends on your alpha. If your alpha is .05, then your obtained p value is not less than .05, and thus, it is not statistically significant. Reflecting the seemingly arbitrary cut-off consequences off null hypothesis significance testing, researchers often write that there was a trend toward significance when p <.10 and >=.05.

4. Would a significance level of 0.06 be slightly significant or not at all?
Trend toward significance is a common label, see above.

5. Do you have any examples of how to write up a moderator regression? 
Frazier, P., Tix, A., & Barron, K. (2004). Testing moderator and mediator effects in counseling psychology research. Journal of Counseling Psychology, 51, 115-134.

McAlister, A. R., Pachana, N., & Jackson, C. J. (2005). Predictors of young dating adults' inclination to engage in extradyadic sexual activities: A multi-perspective study. BRITISH JOURNAL OF PSYCHOLOGY, 96(3), 331.
This article is online at:

Single Group Correlational Study: Basic Analyses

This post sets a basic procedure for analysing a single group observational study in psychology. It aims to provide a starting point, particularly for researchers who do analyse data infrequently.

Formatting Correlation Matrices in Psychology

Researchers in psychology often want to present a correlation matrix of the main variables in a study. This post sets out one way of producing a formatted correlation matrix that conforms to APA style.

Friday, January 16, 2009

Avoiding typing references into Endnote using Google Scholar

This post provides a tip for using Google Scholar to make importing references more efficient.

Thursday, January 15, 2009

Saving Tab Delimited Text Files in Excel without Warnings

Since using R, I have started generally storing my data in tab delimited text files. I tend to edit this data in Excel. Excel is easier to use than the current editor built into R.

Wednesday, December 10, 2008

Basic Analyses of a Dyadic Data Set

This post follows on data manipulation for dyadic data. Once the pragmatic issue of merging data files was resolved, the researcher was interested in considering statistical analysis options for data involving measures of the same variables on a female and her partner.

Merging Data Files for Dyadic Data Analysis

The following post discusses data management for dyadic data.