Sunday, November 10, 2013

The Evolutionary Puzzle of Extreme Rituals

These are my notes for the introduction of the recent Science Express Event on Extreme Rituals at Te Papa, sponsored by the Royal Society of New Zealand Wellington Branch and supported by the Centre for Applied Cross-Cultural Research, Wellington.  Many thanks for all those who came and engaged in a really interesting and stimulating discussion. There will be a podcast available soon, till then here are my typed-up notes for the introductory presentation.

Collective Rituals

I am interested in collective rituals. Many of us here have been to support the All Blacks, Hurricanes or the Wellington Phoenix at Westpac Stadium;  have been to a concert, watched a theater play or have been dancing in one of the bars. Collective rituals are common. 
Why do humans do this? There are no obvious evolutionary functions to these gatherings. They do not help us to get food, ward off predators or to create more offspring (although there may be a bit of  'that' going on after a good night out or when the All Blacks smash the Wallabies ;) 
We could say these rituals just serve pure entertainment purposes and are a byproduct of our big social brains. Obviously it does not hurt to sit and watch grown-up guys chase after an odd-shaped ball or sit through a 1 hour lecture at Te Papa. 

Yet, there are collective rituals in many societies and cultures that are painful, uncomfortable or may even injure or harm participants. It is difficult to convey the suffering that some people inflict on themselves during some of these events  in a talk like this. People chose to walk over glowing hot coals of 600-700 degrees that burns paper within seconds. Others walk over burning hot asphalt for hours in the midday sun without drinking or seeking shade while balancing a hot pot of milk on your head; they hit their face or back with swords, chains or sharp objects till they bleed profusely. Can you imagine hobbling on knees large distance and then circle churches or temples on your knees, over a period that takes a few hours. Most of us would feel uncomfortable sitting on our knees after a few minutes. There are collective rituals where people drag heavy wooden carts hooked into their skin over a distance of 6 km, a feat that takes 5 or more hours. Would you want to pierce your tongue or face with a number of random objects, needles, skewers, metal rods, swords, guns, tree branches or giant beach umbrellas?

Why do people do these things? The British anthropologist Harvey Whitehouse has argued that these more extreme rituals are the historically more ancient types of rituals. But what is the evolutionary purpose and why have such rituals survived over the millennia till today?

These are difficult questions to answer. There is much speculation and many attempts to explain such events by observers and social scientist.  But it is important to test these ideas more rigorously. I have joined forces with a dream-team of scientists to tackle these big questions in the field. 

Separating Collective from Individual Benefits


These rituals may have benefits for both individuals and communities and it is important to separate these effects.
There have been speculations going back more than one century all the way to Darwin and Durkheim that extreme ritual may act as a social bonding devise at the community level. Extreme rituals are thought to increase prosociality and cooperation in the community. They signal commitment to the group:  behaviour speaks louder than words. Groups with strong rituals are more likely to survive. Colleagues and friends such as Joseph Bulbulia, Bill Irons, Richard Sosis and Joe Henrich have written extensively on these theories. How can we test these mechanisms in modern societies?


Extreme Ritual: Studying Social Bonding Effects in Mauritius 

In one study with Dimitris Xygalatas, Joseph Bulbulia and other colleagues, we decided to study an extreme ritual in the small island of Mauritius. Our aim was to study the social aspects of extreme ritual by studying groups of individuals that differ in their involvement in the ritual (see this blog by Matthew Rossano on the larger context). We measured their behavioural responses – how much money are they giving to the community depending on the ritual they participate in. We used a considerable sum of money to ensure that we can tap into real prosocial motivations, where helping the in-group really hurts your own pocket. We studied individuals who participated in a low ordeal ritual, people who watched the high ordeal ritual and those individuals who engaged in what objectively appeared to be rather painful acts (such as piercing themselves with needles, skewers and rods, carrying wooden shrines on their head or pulling them through the streets by attaching hooks to their skin). Supporting the prosocial theories of extreme rituals, those in the high ordeal ritual gave more money to their temple. More importantly, the perception of pain helped us to explain the pattern  – the more pain you perceive, the more prosocial and charitable you become! It is the amount of perceived pain that binds the group together!!!!!


Extreme Ritual: What might be the motivation for the individuals? 


But what is in it for the individuals that perform these rather extreme actions? Why should I care that my group is tighter and more cooperative if it means that I have to hurt myself?
Let's examine one potential mechanism: By committing acts in a ritual that is important and central to the community, I can increase my status and prestige in the community. It signals my commitment to the group and therefore increases my social standing within the group. I become a fully fledged member and can benefit from the support and help of other group members in the future.
To test this social status mechanism I went to Thailand to study another extreme ritual. I again compared groups of individuals – high ordeal devotees who pierced themselves with all sorts of things, assistants who supported the high-ordeal performers and spectators trying to get the blessing of the high ordeal devotees. In order to measure the motivations of individuals, I measured their values (as motivational goals) in life. High ordeal devotees were much more oriented towards getting ahead in life – they are motivated to perform these actions to achieve some level of status and power in the group. In short, for individuals performing these actions is one possible mechanism to move up the social ladder. Often the more disadvantaged and marginalized group members engage in the more audacious acts. Bringing it back to our society, you may just  want to look at who are the boldest and best sports players in a lot of contexts. Often it is minorities that use sports (soccer, rugby, athletics, etc.) as a mechanism to achieve status in a society that values these sports.
Although these rituals may seem bizarre or extreme from an outsiders perspective, they involve universal human characteristics. It is a human universal: we want to be included in our groups. At the same time, we needed strong groups to survive in history. Bringing these threads together, it becomes clear why extreme rituals are appealing to individuals and why these rituals have survived for so long in our history. After all, these extreme rituals are just another expression of our shared humanity.

I hope I have convinced you that it is possible to scientifically study extreme rituals to address interesting questions about us humans. It is amazing to me that despite the popularity and widespread participation in quite extreme collective rituals the world over, we know very little about what is happening to individuals and communities involved in these rituals. I look forward to hearing your comments and reactions and an interesting discussion.

******end of my typed out introductory notes*******

Closing thoughts: Damn, it was a really interesting and fascinating discussion that we had at Te Papa. Many thanks again to everyone for asking questions and sharing their thoughts and experiences. Let's continue the discussion and even more importantly, let's do empirical work on this fascinating aspect of human experience. 

Friday, October 25, 2013

Extreme rituals in transition: Some reflections

I just returned from a turbulent five days in Phuket, Thailand. Although many people associate Phuket with beach, sand and holidays, it is also one of the biggest centres for a Buddhist religious festival, which in all likelihood is one of the most extreme rituals still existing in modern societies. The rituals have a long 5,000 year history, going back to pre-Chinese rituals (see work by Margaret Chan). The ceremonies arrived in Thailand with Chinese groups that worked in mines during the 19th century. I will share some brief impressions of my short visit. I could talk about what I saw for ages, so this is just a quick rambling of thoughts and observations. The purpose was to collect some data on the values and personality of participants in this fascinating ritual and to examine health and well-being effects of a ritual like this. Watch this space for the results coming soon...

The Processions

The festival lasts about 10 days and involves thousands of worshipers from Thailand and the region. It starts with the raising of a lantern pole which invites some Emperor Gods to descend from heaven.

In the following days, believers who act as mediums go into trance and impersonate various gods and warriors. These mediums will walk over heaps of glowing coal and climb bladed ladders (both events have specific religious significance and are performed on days before important events of the larger ritual). The most impressive show of these mediums are the processions though. Starting in the early hours of day, believers will go into trance at the temple which will sponsor and organize the procession of the day. Supporters will dress the mediums (once they are in trance) with the characteristic Chinese style garments and hand them artefacts that show their heavenly power (typically some flag with Chinese inscriptions and a whip with a snake-shaped grip).

Depending on the spirits that possess them, a good number of the believers will engage in piercings and acts of self-mortification. The piercings in some cases can be hard to stomach for outsiders. It ranged from needles and skewers to guns, beach umbrellas, metal saws, swords, tree branches, basket ball hops attached to a car tire, beads and any unimaginable object. I asked how people decide what objects to drive through their cheeks. The answer was that they dream it (the spirit telling them what to do). Apparently, the greater the pain, the more blessing for the individual and the community because this wards off evil spirits (many of the mediums are supposedly fierce warriors that battle evil).

There was blood, sometimes lots of it. Some were cutting their tongue on swords or hit their back with sharp objects. There were moments when I thought it went a bit too far (I remember one moment when I was wondering whether the weight of the object inserted into a person's cheek would rip off the skin completely from his face - well, it did not and I was glad about that. I spare you the photos...). The mediums parade through town in a procession that lasts several hours (covering up to 20 km in distance). They walk barefoot over hot asphalt in the tropical heat.

Onlookers invite them to bless them, their family and their house. In order to increase the torment, firecrackers are thrown on these mediums, with the belief that the more noise being made, the more fortune for the family. Despite all the goriness of the piercings, the throwing of the firecrackers leading to war-like scenes left probably the most impressions on me. I felt transported into a war zone between aliens from a different planet.

The smoke from the firecrackers made breathing difficult. Visibility was reduced to a few meters, only the light of the firecrackers able to break through the fog. The sound of the firecrackers echoing back and forth between the buildings in the narrow streets. Transported into an apocalyptic war zone, the best protection of my face was to hold my camera and shoot back (photos against firecrackers). It was hard to imagine how people were able to endure this for hours. Trying to shoot photos, I got a few hits from firecrackers. The physical sensation of a firecracker exploding on your back, in front of your face or your bare feet (call me well-prepared for venturing out in sandals) is not to be underestimated.



Once back at the temple, the piercings were removed by priests and other participants in the ritual and the spirits were asked to leave (ie. exorcised) in a small ceremony in front of the shrines inside the temple. Soon after the procession, there would be more rituals... a non-stop chain of religious rituals and a kaleidoscope of passion, devotion, shared food, blood and ritualistic suffering.


The other extreme: Silence

In addition to these spectacular extreme rituals, there are other ceremonies that oscillate to the other extreme of tedium and stillness. One of the rituals called Propitiation of Seven Stars involves the warrior mediums in trance guarding a platform structure where more mediums assembled. The temple community is crowded around in silent prayer. One individual beats a drum monotonously for about half an hour. There is not movement and no sound except the tiny dingdingdingding of this drum for the whole period. Unfortunately, I could not really got more information on what is happening during this ritual beyond some general vague stories. But the contrast between extremes could not have been stronger.


Transitions: Between Ancient Traditions and Modernity

Transitions and change were notable. For one, tablets, cell phones and cameras were everywhere to document the extremities. Often it appeared that friends and family of the medium were documenting every single step of the medium, as if to document the suffering to document for posterity. What are these people doing with these documents? Many of the mediums would stop and pose for photos - what is the purpose of this exhibitionism? 

There is a line of research that argues that costly rituals bring communities together and increase the status of individuals who engage in the most extreme forms (see for example work by my colleagues Joseph Bulbulia and Markus Frean). Are these documents used later to increase the status and prestige of the individuals who engage in these activities? Is the posing for pictures a mechanism to increase one's self-esteem? Is the greater shock value of extreme acts translated into social capital within the community?


The extreme nature of some of these piercings was striking. I had seen some postings of such piercings before on the internet, but seeing it in person is different from seeing it from the comfortable safety of your computer. There seems to have been a shift towards more extremes. We did some work on Thaipusam last year in Mauritius (a major field site for studying rituals in the field, directed and organized by Dimitris Xygalatas). In conversations with Temple leaders, they commented that piercings and size of the kavadees that believers carry has increased over the years. An interview with a medium from Phuket posted on the internet came to similar observations. The choice of piercings is supposedly dictated by spirits during dreams.... To me it nearly seemed like a ratcheting of extreme actions (think of the 'keeping up with Joneses' effect). To what extent is the posting of extreme piercings, sensational violence, and all sorts of other evidence of self-inflicted harm on internet and global media leading to a 'spiritual arms race'? 


Yet, at the same time, other changes led to changes in the rituals. Blade-walking and bladed ladder climbing used to be a significant part of the ritual. Yet, these practices in the times of HIV are dangerous and can infect a large number of participants. For this reason, a number of temples have abandoned this practice. 

Some changes were also quite cute. It was fun to see how kids and youngsters went through some of the more boring parts of the ritual by secretly using their cellphones and tablets for entertainment. Some of these folks had devised ingenious techniques for entertaining themselves while pretending to be good believers. 


In many ways, these simple observations of human reactions during a period that felt so extreme punctuated the humanness of it all. The pained expressions of bystanders, the open eyes in awe and disbelief of some of the suffering that was displayed, the innocent attempts to ward off both shock and boredom revealed the human aspect. It is these small observations that brought it home to me that through the extremity it actually shows the universality of what it means to be human. There will be people who go to extremes for all sorts of reasons in all cultures. There will be pain and boredom in all places. And it is the reaction to these universal elements that reveals our shared humanity.  





I am glad I went. Many thanks to Janpaphat Kruekaew for introducing and opening this fascinating world for me. Quite often when I was too tired to continue, she tirelessly continued asking people and getting interviews and responses from participants. It was a humbling experience working with her. I am looking forward to going back and getting a better understanding of the minds of some of these people, those who go to extremes and those who just stand by and watch the procession unfold in front of them.

Tuesday, October 1, 2013

Sexy, hot and easy! Should we trust student evaluations of university courses and teachers?

How should we rate the effectiveness of university teachers and university education in general? This is a million dollar question that is hardly ever been questioned, yet determines the course of countless lives on both sides of the divide (student and teacher). 
Teaching evaluations are treated with suspicion by profs and teachers, but are loved by university bureaucrats and administrators. Students may often not realize, but these evaluations, specifically the mean numbers that come out after some basic number crunching often have a huge impact on the careers and success of academics. A rather small difference between a 1.9 and 2.5 can determine whether somebody gets promoted, is hired or gets a bonus. In extreme cases, a teacher may lose a contract. So how much evidence is there that these numbers provide good evidence of teacher effectiveness? And are teachers who do get higher ratings actually those teachers that help students succeed in other courses?

A number of recent studies cast some big doubt on the usefulness of these criteria. Let's start with a fun example. 

Quality of teaching or just easy and 'hot'? 

James Felton, a Professor of Finance at Central Michigan University and colleagues examined the evaluations submitted to http://www.ratemyprofessors.com/. There are a couple of different criteria that students can rate their professors on. The two core areas are 'helpfulness' (how helpful and approachable a teacher is) and 'clarity' (how organized, clear and effective is a teacher). These two are averaged to get a rating of overall teaching quality. There are two more evaluations though. The first one is 'easiness', meaning how easy or difficult the classes are and how much work is needed to get an A. The second is 'hotness', a simple rating of whether a student thinks that a teacher is hot or not. Obviously, we would want to have teachers that are effective and helpful, but these perceptions should not be driven by how easy a course is or how attractive a teacher is. When looking at the data from ratings for 6,852 profs from 369 institutions... the answer is that the hotter you are and the easier your course is, the better are your evaluations. The correlation between easiness and quality is a whopping .62, whereas hotness and quality correlate .64. 

They offer this explanation: 
We see Quality as a function of Easiness, but it could be argued that Easiness is a function of Quality, where professors who are skilled in the classroom take difficult material and make it seem easy. We wish that were the case, but we see Quality as a function of Easiness the majority of the time for two reasons. First, as stated previously, Ratemyprofessors.com (2004) defines Easiness as the ability to get a high grade without having to work very hard. Second, professors with high Easiness scores usually have student comments regarding a light work load and high grades. Similarly, we see Quality as a function of Hotness, but it could be argued that Hotness is a function of Quality, where a brilliant professor, regardless of physical appearance, is considered sexy by his or her students. Again we wish that were the case, but most student comments point toward Quality as a function of Hotness when they focus on physical characteristics of their professors that could be captured in photographs.

The lesson that might be learned from this correlational study is that it does not hurt to dumb down your lecture content and hit the gym (well, the latter would be good regardless). 

Lecture fluency or welcome back, Dr Fox...

Now, let's enter study number 2. Shana Carpenter and colleagues from Iowa State University in a study recently published showed students a short video of the same teacher presenting the same material. The major difference was that in one video the prof acted in what was called a fluent way: upright, confident, with eye contact and speaking fluently without notes. In the other condition, the prof acted disfluent: slumped, looking away, speaking haltingly and relying on notes. In two experiments, students were tested on how much they actually learned and ratings of the prof were also obtained. The results very clearly showed that the fluent prof was rated much better (surprise surprise), but also that students thought that they had learned more and would remember more from the fluent prof compared to the disfluent prof. However, when later tested, there were no differences between the two groups. This means instructor fluency increases perceptions of learning but not actual learning! There were also some curious smaller findings. For example, for the disfluent group -when given the opportunity to read the transcript of the lecture, students who spent more time rehearsing had higher test scores. This was not the case for the fluent group. It is an ambiguous finding, but could indicate that fluent lectures may decrease the attention paid to study material when preparing for an exam. Not sure whether this is desirable.

This really sounds like the famous Dr Fox effect.  Talk nonsense as long as you are dynamic, engage the audience and are make jokes...

Some disturbing findings when using random assignment of students to profs

The most concerning study though used a controlled random assignment of students to courses that overcomes a lot of the shortcomings of previous studies (including self-selection of students to courses and professors). Scott Carrell and James West studied student achievement and course feedback as students moved through mandatory classes in maths, science and engineering. The unique aspect of their study is that professors rotated in sections of the course, assessment was not done by the professors themselves and students were randomly allocated to professors (but all studied the same content). A first finding that is of practical importance is that less academically qualified instructors got students more (erroneously?) interested in the topics which resulted in better immediate student performance, but then led to lower scores in follow-on related courses. More experienced and qualified professors in contrast had students that did not well in the introductory classes, but those students than excelled later on. Those students were able to build on what they had learned during the initial courses. 

What is even more important is that professors who were rated positively by students did better in the initial courses. However, the rating of the effectiveness of the professor did not predict later performance! In fact, in a number of cases the correlation flipped - students studying with the more highly rated professors did worse in half the courses than those who studied with a prof who was not rated as highly (note: only one of these correlations was significant - the point remains the same though: ratings of teacher effectiveness does not predict long-term student achievement). As Carreel and West argue:
'Since many U.S. colleges and universities use student evaluations as a measurement of teaching quality for academic promotion and tenure decisions, this finding draws into question the value and accuracy of this practice.' 

Are there alternatives? Yes! 

The reliance on student evaluations for courses and teachers is problematic, if these evaluations are not considered in a larger context of what is achieved in a course. In the business world, this has been long realized. Teaching is effectively training. In the organizational world, Donald Kirkpatrick developed a famous four stage model of training evaluation. The four criteria for the evaluation of the effectiveness of training are:

  1. Student reactions - this is essentially equivalent of student evaluations, assessments of students thought they had learned and how they felt about the course/the teaching
  2. Learning - this is measured by the increase in knowledge or capability after the course, we could consider the test performance in a test or exam as a good measure of this (of course only if the assessment is independent of the teacher - see above the problem with the easiness of a course)
  3. Behaviour change - this refers to the changes in the behaviour outside the teaching environment that are a result of the teaching, including applications of what has been learned to new situations outside the teaching/training context
  4. Results - this is the effect of the teaching on the business or the larger environment that results from the performance and the behaviour changes induced by the teaching/training
The 3rd and 4th points are what universities (and society) should be concerned about. There has been a lot of questioning of the value of tertiary education recently (see for example here, here and here). These criteria can help in re-adjusting both the focus of universities as well as criteria that are used to evaluate professors. 

Students and society deserve better, not just the profs ;)

Comments are welcome as usual :)

Tuesday, May 14, 2013

Unpackaging culture & cultural differences

One of the most fascinating questions arises when we observe that individuals in a different cultural system behave or act in a different way. Why do they do that? What is the explanation or reason for showing these particular behaviours or responses? For example, we may have stepped on an exotic island and observe that the inhabitants eat way more chocolate ice cream that we do. Or they may tell you that there are lots of little ghosts out there taking care of them, many more than you ever thought would be possible to inhabit a small island like this. Or they may simple say in some interviews or surveys that they do not like to work as hard as you normally would expect in adult samples. How could we explain any of these differences?




Given the perpetual problem of potential bias in comparative research, we can never really rule out that our observations were simply erroneous - we might have had the wrong instruments, there may have been language problems in interactions (remember Lost in Translation?), we may have mis-interpreted the data or it may have simply been a chance difference. 

One persuasive idea that has been around for quite a while in the social sciences is the idea of unpackaging. The terms goes back to a classic study conducted by Whiting and Whiting (1975). They orchestrated a large ethnographic study of child development among six communities: a New England Baptist community; a Philippine barrio; an Okinawan village; an Indian village in Mexico; a northern Indian caste group; and a rural tribal group in Kenya. They reported differences in a number of psychological processes, socialization and child-rearing patterns. Going beyond just noting these differences, they reasoned that there must be specific contextual variables that could explain the differences found, linking ecological constraints faced by these communities to psychological processes via adaptive socialization practices. For instance, they compared the activities of children from the same families, some of whom were living in cities and others in villages. They also compared families in which young boys helped with baby-tending with those in which girls did the helping. Therefore, these social conditions were linked to observed behavioural differences, leading to one plausible explanation of why they may have occurred in the first place. This is essentially what psychologists study as mediation:  processes and variables that explain the relationship between an independent or predictor variable and the dependent or criterion variable. It is about the causal theoretical processes, the how and why of the observed effects. We often think of mediators as internalized psychological processes of external conditions that lead to other outcomes down the causal chain. In cross-cultural work, it does not necessarily always be an internal psychological variable - it could also be living conditions or social constraints and norms that can act as mediators. 

Put differently, unpackaging studies are extensions of basic cross-cultural comparisons in which the active ingredient presumed to cause the observed differences in psychological processes is directly measured and explicitly tested for its role in explaining the outcome. Have a look at the graphical representation of mediation. Unpackaging culture is one term often found in the literature, others include ‘linkage studies’ (Matsumoto & Yoo, 2006), ‘mediation studies’ (Kirkman, Lowe & Gibson, 2006) or ‘covariate studies/strategies’ (Leung & van de Vijver, 2008).




 For example, Tinsley (2001) found that differences in the conflict management strategies of German, Japanese and US managers were completely mediated by the values held by members of these cultural groups, and Felfe, Yan and Six (2008) reported that individuals’ scores on a ‘collectivism’ scale mediated differences in organizational commitment across samples of Romanian, German and Chinese employees.  

In an ideal test of mediation, the researcher tests whether other relevant variables that are not related to the hypothesis also yield mediation effects. This provides greater certainty in establishing exactly what the causes the results that are obtained. For instance, Y. Chen, Brockner, and Katz (1998) showed that a measure of individual-collective primacy mediated the intergroup effects that they had predicted and found. They then tested whether six other measures derived from the concept of individualism-collectivism also mediated these effects, and found that they did not. Studies of this kind help to clarify the loose and varied ways in which the psychological aspects of individualism and collectivism have been employed by different authors. 

What are some general concerns?
In the above examples, the mediator and dependent or criterion variable were measured using the same or similar methods. If there is some third unmeasured variable that is related or unrelated to the independent variable, we may end up with a situation where it appears that there is mediation, whereas in reality, there is none. Having multiple mediators measured with the same method as the DV may lead to some reassurance about the findings, but ultimately, the best test would be a test using independent methods

Experiments have been much in vogue recently to study cultural differences. One of the major concerns is whether the manipulation was effective or not. This is again the problem of potential bias in comparative studies. If we have a psychological mediator in our experiment that highlights how the manipulation is affecting the DV, then we are much safer grounds in terms of explaining the psychological processes. 

In summary, unpackaging has two important and inter-related features: identification of the theoretical factors or processes that may cause cultural differences in psychological outcomes of interest, and an explicit empirical test of the proposed processes leading to these outcomes. Therefore, it is as much about theory as it is about methods and stats. Having unpackaged an observed difference in behaviour, attitudes or beliefs and having ruled out alternative theoretical explanations (other potential mediators), we can also place much more confidence in our results. I leave it up to you to come up with potentially meaningful variables that we could use in those three semi-silly examples (ice cream, ghosts and motivation). Once you have some mechanism, the next phase would be to test whether the mediator(s) actually do the job. Ideally, this is one of the best ways to rule out measurement bias - explain where the difference came from and that the difference is not driven by artefacts. 


Some more technical explanations are available in Fischer (2009); Leung & Van de Vijver (2008) and Poortinga & Van de Vijver (1987, this is an excellent discussion early discussion with some great multi-method examples). Excellent resources and tutorials for running state-of-the-art mediation analyses are available from on Kristopher Preacher's and Andrew Hayes' websites. Use them!!!!!

Overall, I think this is the most exciting part of cross-cultural research - put on your detective hat and find out where any difference that you perceive in the world ultimately stems from. 


Thursday, February 7, 2013

My 7.5 General Guidelines for Reviewing Journal Articles

I am involved in a few editorial boards and as part of these duties, I get the occasional question about what it means to be a reviewer and what is required if doing a review. There are some great resources on the web, look for example here, here, and here. These links have some excellent suggestions for evaluating the suitability and quality of a manuscript.
Hence, my goal is to just simply add a more personal view on reviewing and some guidelines.
Science in its current form rests on a peer review, the scientific research process relies on improving ideas, methods, designs and theories through discussion with peers. The review process is just one aspect of this.
So here are some rather random guidelines:

1. Be fair 

This may be the most important guideline. Research should be objective and dispassionate, but this is often hard to maintain. Researchers spent most of their time working on a project and become highly identified with their scientific 'babies'. It is easy to become opinionated about your and others research. The theoretical lenses through which we conduct our research will lead to biases and preferences. Do not let these professional blindspots guide your reviews. Evaluate the submitted manuscript on its merits and what you consider to be weaknesses. State your own biases or assumptions, if this helps to clarify why you argue a particular point. If you know the person (experienced researcher will easily identify the author of an article), do not be tempted to engage in personal feuds (e.g., "I will get back to you about that nasty comment you made at my last presentation in XYZ" ;). Stay professional and evaluate the manuscript on its scientific merit.

2. Be supportive

Help authors to find the parts in their manuscript that are less clear. Researchers are passionate about their research, acknowledge the strengths of the study. It often helps to quickly summarize what you consider to be the key points. This will show that you have understood the manuscript and also may highlight some additional points that the authors may not have thought about yet.

3. Be critical

Evaluate the whole manuscript in its style and content. What areas need improvement? Are there alternative interpretations of the data or the results? What are flaws or problems in the argument or theorizing that you can see? Where are ambiguities in style or expression?

4. Offer constructive suggestions

Help authors to improve their work. This is the reason why we have a peer-review process! Provide references to additional literature. Suggest theories or interpretations that help to shed light on the research.

5. Be open

Research is about charting unknown territory. There may be ideas or approaches that may appear strange or  unconventional. Don't judge ideas prematurely.

6. Consider the bigger picture

Research often tends to focus on very narrow aspects or specific questions. It can be helpful to consider the wider picture again, especially if there are implications for real world problems, people or communities. If you see some implications or applications, highlight them. I believe it is important to consider how research can contribute to society. As a reviewer, you can help authors in this respect.

7. Do it!

One of the most frustrating experiences as an editor is finding reviewers. We often spend hours on the internet and going through papers, journals or books in order to identify some suitable reviewer. To get declined review requests can be very frustrating. Of course, there are legitimate reasons to decline a review. You may not know much about this area (so my mistake of inviting you in the first place). Don't review if you are not qualified to comment on the research. You may have personal or ethical reasons for not reviewing certain papers. This is all fine. An editor can understand this and appreciates a quick email stating these reasons. But on the other hand, there are more and more pressures from universities and institutes to publish. Reviewing is sometimes seen as a waste of time and a nuisance by some academics. I know colleagues who proudly confess that they have never reviewed a paper. I think this is unacceptable. If we all behaved like this, the process would break down.
Get engaged. Help with shaping research. Get inspired with new ideas (after all, you are seeing research as it is unfolding in its final stages). Do it! We need you! 

7.5 .... and submit your review on time 



Have fun reviewing
:)

Thursday, October 11, 2012

How to run a Conditional ANOVA


Today is a wee bit heavier on the stats side again. If you are interested in Differential Item Functioning and how to do it with an easy to use tool, this is for you...

Aim: Identify differential item functioning in numerical scores across groups in order to decide whether the items are unbiased and can be used for cross-cultural comparisons.

General approach: Van de Vijver and Leung (1997) describe a conditional technique which can be used if you use Likert-type scales. It uses traditional ANOVA techniques. The independent variables are (1) the groups to be compared and (2) score levels on the total score (across all items) as an indicator of the true observed or ‘latent’ trait (please note that technically it is not a latent variable). The dependent variable is the score for each individual item. Since we are using the total score (divided into score levels) as an IV, the analysis is called ‘conditional’.

Advantages of Conditional ANOVA: It can be easily run in standard programmes such as SPSS. It is simple. It highlights some key issues and principles of differential item functioning. One particular advantage is that working through these procedures, you can easily find out whether score distributions are similar or different (e.g., is an item bias analysis warranted and even possible?).

Disadvantages of Conditional ANOVA: There are many arbitrary choices in splitting variables and score groups (see below) that can make big differences. It is not very elegant. Better approaches that circumvent some of these problems and that can be implemented in SPSS and other standard programmes include Logistic Regression. Check out Bruno Zumbo’s website and manual. I will also try and put up some notes on this soon.

What do we look for? There are three effects that we look for.
First, a significant main effect of score level would indicate that individuals with low score overall also show a lower score on the respective item. This would be expected and therefore is generally not of theoretical interest (think of it as equivalent to a significant factor loading of the item on the ‘latent’ factor).
Second, a significant main effect of country or sample would indicate that scores on this item for at least one group are significantly higher or lower, independent of the true variable score. This indicates ‘uniform DIF’. (Note: this type of item bias can NOT be detected in Exploratory Factor Analysis with Procrustean Rotation).
Third, a significant interaction between country and score level on the item mean indicates that the item discriminates differently across groups. This indicates ‘non-uniform DIF’. The item is differently related to the true ‘latent’ variable across groups. For example, think of an item of extroversion. In one group (let’s say New Yorkers), ‘being the centre of attention at a cocktail party’ is a good indicator of extroversion, whereas for a group of Muslim youth from Mogadishu in Somalia it is not a relevant item of extroversion (since they are not allowed to drink alcohol and probably have never been at a cocktail party, for obvious reasons).
Note: Such biases MAY be detected through Procrustean Rotation, if examining differentially loading items.

Important: What is our criterion for deciding whether an item shows DIF or not?

Statistical Procedure:

The procedure requires in most cases at least four steps.
Step 1: Calculate the sum score of your variable. For example, if you have an extraversion scale with ten items measured on a scale from 1 to 5, you should create the total sum. This can obviously vary between 10 and 50 for any individual. Use the syntax used in class.

For example:

Compute extroversion=sum(extraversion1, extraversion2,…,extraversion10).

Step 2: You need to create score levels. You would like to group equal numbers of individuals into groups according to their overall extroversion score.
Van de Vijver and Leung (1997) recommend having at least 50 individuals per score group and sample. For example, if you have 100 individuals in each group, you can maximally form 2 groups. If you have 5,000 individuals in each of your cultural samples, you could theoretically form up to 100 score levels (well actually not, because you would have only 40 meaningful groups in this example since the difference between maximum and minimum possible score is 40). Therefore, it is up to you how many score levels you create. Having more levels will obviously allow more fine-grained analyses (you can make finer distinctions between extroversion levels in both groups) and probably more powerful (you are more likely to detect DIF). However, because you have fewer people in your analysis, it might also be less stable. Hence, there is a clear trade-off, but don’t despair. If an item is strongly biased, it should show up in your analysis independent of you have fewer or more score levels. If the bias is less severe, analyses might change across different options.

One issue is that if you have less than 50 people in each score group and cultural sample, the results might become quite unstable and you may find interactions that are hard to interpret. In any case, it important to consider both statistical significance as well as effect sizes when interpreting item bias.

A simple way of getting the desired number of equal groups is to use the rank cases option. You find this under ‘Transform’ -> ‘Rank cases’. Transfer your sum score into the variables box. Click on ‘Rank types’. First, unclick ‘Rank’ (it will rank your sample, but this is something that you do not need). Second, click on ‘Ntiles’ and specify the number of groups you want to create. For example, if you have 200 individuals, you could create 4 groups. If you have larger samples, the discussion from above applies (you have to decide about the number of levels, facing the before-mentioned trade-off in terms of power versus stability).

As discussed above, it is strongly advisable to interpret effect sizes (how big is the effect) in addition to statistical significance levels. This is particularly important if you have large sample sizes in which often minute differences can become significant. SPSS gives you partial eta-squared values routinely (if you click on ‘effect sizes’ under the ‘options’). Cohen (1988) differentiated between small  (0.01), medium (0.06), and large effect size (0.14) for eta-squared. Please note that SPSS gives you partial eta-squared values (which is the variance due to the effect, independent of the effect of other effects), whereas eta-squared does not take the other effects take into account. Partial eta-squared values are often larger than the traditional eta-squared values (overestimating the effect), but at the same time there is much to be recommended for using partial instead of traditional eta-squared values (see Pierce, Block & Aguinis, 2004, in Educational and Psychological Measurement).

Step 3:  Run your ANOVA for each item separately. The IV’s are country/sample and score level (the variable created using ranking procedures). Transfer your IV’s into the ‘Fixed Factor’ boxes. As described above, the important stuff to look out for is the significant main effect of country/sample (indicating uniform DIF) and/or the significant interaction between country/sample x score level (indicating non-uniform DIF). You can use plots produced by SPSS to identify that nature and direction of the bias (under plots, transfer your score level to the ‘horizontal axis’ and the country/sample to ‘separate lines’, click ‘add’ and then ‘continue’). Van de Vijver and Leung (box 4.3) describe a different way of plotting the results. However, the results are the same, only different way of visualising the main effect and/or interaction.
This little figure for example shows evidence of both uniform and nonuniform bias. The item is overall easier for the East German sample and it does not discriminate equally well across all score levels. Among higher score levels, it does not differentiate well for the UK sample. 



Step 4: Ideally, you would not like to have DIF. However, it is likely that you will encounter some biased items. I would run all analyses first and identify the most biased items. If all items are biased, you are in trouble (well, unless you are a cultural psychologist, in which case you rejoice and party). In this case, there is probably little you can do at this point except trying to understand the mechanisms underlying the processes (how do people understand these questions, what does this say about the culture of both groups, etc.).
If you have only a few biased items, remove them (you can either remove the item with the strongest partial eta-square or all of the DIF items in a single swoop – I would recommend the former procedure though) and recompute the sum score (step 1). Go through step 2 and 3 again to see whether your scale is working better now. You may need to repeat this analysis various times, since different items may show up as biased at each iteration of your analysis.

Problems:

My factor analysis showed that one factor is not working in at least one sample: In this case, there is no point in running the conditional ANOVA with that sample included. You are interested in identifying those items that are problematic in measuring the latent score. You therefore assume that the factor is working in all groups included in the analysis.

My overall latent scores do not overlap: This will lead to situations where the latent scores are so dramatically different that you can not find score levels with at least 50 participants in each sample. In this case, your attempt to identify Differential ITEM functioning is problematic, since something else is happening. One option is to increase score levels (make the groups larger – obviously this involves a loss of sensitivity and power to detect effects). Sometimes, even this might not be possible.
At a theoretical level, it could be that you have a situation where you have generalized uniform item bias in at least one sample (for example because one group gives acquiescent answers that are consistently higher or lower). It also might indicate method bias (for example, translation problems that make all items significantly easier in one group compared to the others) or construct bias (for example, you might have tapped into some religious or cultural practices that are more common in one group than in another – in this case your items might load on the intended factor but conceptually the factor is measuring something different across cultural groups). Of course, it can also indicate a true differences. Any number of explanations (construct or method bias or substantive effects that lead to different cultural scores) could be possible.

What happens if most items are biased and only a few unbiased items remain? In this situation you run into the paradox that you can not actually determine whether your biased items are actually unbiased or unbiased items are biased. This type of analysis only functions properly if you have a small number of biased items, up to probably half the number of items in your latent variable. Once you move beyond this, it means that there is a problem with your construct. If you mainly find uniform bias, but no interactions, you can still compare correlations or patterns of scores (since your instrument most likely satisfies metric equivalence). If you have interactions, you do not satisfy metric equivalence and you may need to investigate the structure and function of your theoretical and/or operationalized construct (functional and structural equivalence). 

Any questions? Email me ;) 

Thursday, September 6, 2012

Tales from the field: The last day on the African coast


Time is definitely a rare commodity. I touched down in Capetown, South Africa on a Friday, the 13. Although it feels it was only yesterday, I have one more day and I am back on my way to Europe. It has been a wild and tumultuous two months, stopping in 4 countries and covering thousands of kilometres while having hammered away on my laptop in many random places. I plunged straight into the IACCP winter school business, met 37 bright and eager young minds and had a pretty full-on 4 days of workshop discussions, debates, laughter and a few hours out and about Stellenbosch with this new crew of cross-cultural researchers. The winter school experience was amazing, mind-blowing and exceptionally tiring, mainly the little organizational details that cropped up here and there and everywhere. But it was very cool and emotional to see the presentations of the groups at the end. Even those that had not gone as far as some other projects showed clear signs of research in progress and those intellectual struggles that good research is all about. Sometimes the process is more educational than the polished outcome. So it was a brilliant couple of days. One point though for next time: I am not sure I would have really needed the fridge – sorry, my super luxuriously prized room in the brilliantly spartanic student hostel, but hey, it was good to have a place to change your clothes…

After the winter school, there was one short day to breathe, which I spent with an excursion to the amazing cape peninsula. The weather was stunningly beautiful (as were the prices that we were charged, but only found out later for what kind of ride we were taken).  Then we went straight on with the IACCP conference. One lesson that I still not have learned is how to say ‘NO’. As a consequence of this minor learning deficiency of mine, I was giving non-stop talks most days or was in sessions to attend friends’, colleagues’ or students’ work, and still missed out on so many other interesting looking papers. I am a conference junky, I openly admit and plead happily guilty on this charge, I had a ball and intellectual feast to last me a while. But my sleep deprivation was also quite severe, so at the end of it I was a little nervous wreck. I apologise to everyone who may have thought I was a strange weirdo (well… ). All in all, the conference passed way too quickly and I wished I had more time to talk to friends and colleagues that I had not seen in a long time.

From there things got a bit more complicated. The whole itinerary of the trip was abruptly changed shortly after it started. My plan was to make it from South Africa to Kenya overland through Mozambique and Tanzania. Unfortunately, a few days into this adventure my right eye decided to develop a fairly painful inflammation. Finding an ophthalmologist in semi-rural Mozambique turned into an anthropological field study on the health care system in one of the poorest countries in the world. We finally found a doctor who had some adequate equipment. The original diagnose was scary enough for me to change my schedule and return to South Africa. When I got back there 4 days later, the local eye specialist could not find any evidence of the ulcer that was supposedly and permanently threatening my eye sight.  It was still to be treated with care, but nowhere as serious as it seemed originally. I was devastated. I do not like to lose, especially not against your own body. The medical retreat to Kenya via airplane and a necessary stopover in Dar El Salaam were psychologically painful.

Instead of exploring one of the last frontiers of Africa, I pretty immediately got my nose into the work we were supposed to be doing here in Kenya. Yet, things never quite work out as you had planned them. In addition to all the usual chaos in places like this, we also had some minor political upheavals (luckily I was always a bit far from the action) and the teacher strike right now made a big cross through our field work plans. Instead I was given some sessions on research on field staff and students and beyond that was stuck behind my laptop crunching tons of health related data that we had collected already. Once you get beyond the abstractness of numbers and start connecting the immaterial statistical facts to the images of poverty and inequality that surround you on a daily basis, it can become a pretty heart-breaking and depressing reality. The colourful cloths of women carrying huge loads on their head and little kids on their back, the exotic mud houses with their dark single inner room and picturesque little villages wane when you get closer.

Poverty is only beautiful from a distance.

I am not sure how much of this reality is really being noticed by the European and North American tourists in their flash resorts. At some stage, I could not stand anymore some of the comments and conversations I overheard from fellow Western travellers when I had breakfast at some of the hotels I was staying at (yes, I am damn lucky, this time I was mainly staying in pretty amazing and flash places – thanks to sugar mammi Amina ;). This disparity between what you see outside the bubble of western consumerism in the villages and the streets of the Kenyan coast and what you perceive when being at your tourist hotel, with limited excursions into the harsh reality through well-organized and expensive safaris was probably equally depressing.

So it was a lot of work, full on days and long nights, but with widely ranging schedules. Often I would start working around 8am and only finish at 10 or 11pm, sometimes it could also be 2am before the computer was switched off. Internet was mainly through USB modems and access could be described as patchy at best. I was amazed though how smooth the skype connection worked for the Postgrad Committee meeting. There would be breaks for lunch or dinner (well, first it was Ramadan, so food and lunch breaks were fairly random and hard to come by). Sometimes we would go to a fancy hotel (thanks Jackie), more often though in one of the few local eateries. It was really funny to see how quickly you adapt, a colleague from Zambia is just visiting and he was commenting on the things that I do not notice anymore, how dirty some of these places were, with lots of flies and mosquitos and a dark, crammed atmosphere. A funny incidence happened at lunch today, when a big mama sat down behind me and squeezed me into the table that I was sitting at. As this opulent woman also had a very animated eating and conversation style, my chair also bounced up and down with each wiggle of her impressive body. As I was firmly squeezed into the chair and pressed against the table, I moved in unison with my chair directed by that lady. The extraction process – this is me trying to get up and out – was equally funny. The waitress had to intervene and help the lady move so that I could get out.

A trip full of experiences, smells, colours, impressions and learnings. Some of the things I learned are pretty random, like that you should avoid crossing land borders after night fall, even if it is friendly territory. We again had to oil the wheels of informal markets dressed in neat uniforms – well, this is actually a bit of a lie, the main recipients of unaccounted amounts of cash were well-fed plain clothes officers in nice leather jackets with big plastic ID cards hanging around their neck. Or that most of the staple on the coast (cassava, pineapple, maize) were all introduced from South America by the Portuguese. Or that the East African coast had sophisticated civilizations and extensive trade connections all the way to India and China when Europeans had only a vague idea what was behind the horizon of the Mediterranean. African ivory was a prized possession in China and was carved into the most exquisite sculptures. Chinese porcelain needed African raw materials. In turn, the traders and town folk among the Swahili were sipping their spice tea from elegant Chinese tea cups.

Globalization was long here before we knew it…

I also learned a new appreciation for washing machines and street lights. The former because it would not force me to wash my clothes every second day or so by hand. I got some good instructions from one of the cleaning ladies and demonstrated them by publicly washing my underwear, to the amusement of the department. The latter because it is pitch black without the moon as a street light substitute. Try walking on a path when you are not a cat or forgot to bring your magic battery-powered torch. A few days ago I wandered straight into a half metre deep hole and nose dived into the sand. Just at that moment a guy had to pass and witness my blind struggles with the African soil. Speaking of moving on roads, I will definitely miss the traffic here and its fluid interpretation of a few basic road rules. I have the perception that people consensually agree to drive on the left (being obedient ex-colonial subjects). Speed limits are not a major worry because of the state of the roads – as a digression, I actually wonder whether speed bumps are the more regular inverse of  random pot holes. But this mathematical law may only apply to the rare occurrence of paved roads. -  However, beyond that it is a country of infinite freedom of driving. This is refreshing after our experiences in Mozambique where we had to contribute to African development (a.k.a police officers’ salaries) on an unpleasantly regular basis. But this state of ultimate freedom (and lack of well maintained roads) also meant for example that an excursion along the Kenyan coast that was supposed to last about 30min turned into a 4 hour odyssey, with the car stuck in sand, random smooth talking politicians hitching rides and us ditching the car in the end because it got stuck on some hug slab of coral stone in the middle of the well-worn sand track.

I should be packing now, but instead are procrastinating writing down these random observations as I am trying to squeeze too much stuff into my little back. Looking back to two eventful and busy months, I wish I had slept less, gone out more and met more people. But then hey… can somebody genetically engineer my body to not feel tired? We got lots of projects done, but still there is so much more to do.

I will miss the amazing tropical fruit, the heat, the sun and the fresh breeze off the Indian Ocean. I will miss the sleepy hectic, the tranquil everyday chaos of cars, motorbikes and people, smells, noises and shy gazes towards that strange muzungu. But I also look forward to seeing familiar friends and family. Next stop is Germany, will be nice to see my folks and friends I have not seen in ages. Will be good to have a clean table, clean cloths, home-cooked food and a clean house with no power cuts and no dust and sand covering your body in the evening.

I think I need a proper holiday… But work is not stopping, I need to finish our book with Peter, Viv and Michael and need to write that chapter for the Advances series. Also look forward to the talks and workshop in Bremen and seeing Diana and Katja and their new academic environment.

May the adventure continue….