Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Saturday, September 13, 2008

The Psychology of Music Preferences

Apple has just released a version of iTunes with a new feature called Genius. Genius makes custom playlists (of either music you own or music you might want to by) based on a secret algorithm. If you opt-in, iTunes sends all of your listening history data (e.g., track names, artist names, playcounts, skipcounts) to a central server. The algorithm then looks for patterns in worldwide listening trends. To use Genius you right-click a particular song, choose 'start Genius', and BANG you've got a list of 'similar' songs. I'm loving it. It's helped me rediscover some music that Dan had given me but that I hadn't listened to much.

This got me thinking about how one might statistically look for trends in music preferences. I wondered if there'd ever been a factor analysis of music preferences. A factor analysis is a statical technique for finding trends amongst different variables. It's often used in personality research. You ask a large sample of volunteers a whole stack of questions (e.g., "On a scale of 1 to 10 how much do you like parties?", "How much do you like being the center of attention?" etc.) and look for shared variance in the responses. I've written about factor analyses of personality related data before here.

Anyway, a quick literature search turned up this article in the Journal of Personality and Social Psychology:

Rentfrow, P.J., Gosling, S.D. (2003). The Do Re Mi’s of Everyday Life: The Structure and Personality Correlates of Music Preferences. Journal of Personality and Social Psychology, 84(6), 136-1256.

Among the studies reported is a factor analytic study of music preferences. 1,704 students from the University of Texas were asked to rate each of 14 music genres on a scale from 1 ('I don't like it at all') to 7 ('I like it a great deal'). The genres were: alternative, blues, classical, country, electronica/dance, folk, heavy metal, rap/hip-hop, jazz, pop, religious, rock, soul/funk, and sound tracks.

The analysis revealed 4 main dimensions (factors) that captured 59% of the total variance. The names given to these factors and the genres associated with them are as follows:

- Reflective and complex (blues, jazz, classical, and folk)
- Intense and rebellious (rock, alternative, heavy metal)
- Upbeat and Conventional (country, sound track, religious, and pop)
- Energetic and Rhytmic (rap/hip-hop, soul/funk, and electronica/dance)

These dimensions are reasonably independent of each other (1). People who like reflective and complex music are just as likely to enjoy intense and rebellious music as they are to not. What these factors mean is that if someone likes a genre related to a particular dimension (e.g., blues) then they'll probably also like the other genres on that dimension (e.g., jazz). The same goes for disliking a genre.

One limitation of this study is that peoples' understanding of genre terms may vary. I might think that I don't like folk music and yet like many songs that others would categorise as folk. It would be great to see an analysis done on song by song ratings, rather than just genres.

Another analysis, which was really interesting, involved looking for relationships between e musical preferences and differences in personality and cognitive ability. They found all sorts of relationships, although most of them were quite small (.2ish). The largest one (.4ish) was between a preference for Reflective and Complex music and the personality characteristic Openness to Experience. Interestingly, there was a small (.2ish) relationship between verbal IQ and liking of Reflective and Complex, Intense and Rebellious, or Upbeat and Conventional music (2).

Very interesting stuff.

I'd love to see these researchers team up with Apple and analyse the iTunes Genius data.

---------

(1) Upbeat and Conventional and Energetic and Rythmic correlate .5 if allowed to covary.

(2) And no, I don't think this is evidence that music makes you smarter (can you guess why?).

Monday, July 21, 2008

The rock-off fairness fallacy (SOI cross-post)

I have a new blog! Dan and I are doing a joint blog over at thesubjectsofinterest.blogspot.com

Here's my latest post from over there.

-------------
I've noticed that there is an increasing trend for people to resolve disputes or allocate resources using the game of Rock Paper Scissors. The procedure is sometimes called a 'rock-off', as in "let's rock-off for the last slice of pizza"

This procedure is fine in a two person game (assuming no one cheats), but I often see people happily submitting to three-way rock-offs. In these arrangements two people rock-off, then the third person plays the winner, and the winner of that second rock-off is declared the overall winner.

But this procedure is inherently unfair!

Imagine that John, Fred, and Mary are rocking-off for a slice of pizza. John and Fred play first. Mary plays the winner.

For John to win overall he has to win the first encounter against Fred and then a subsequent encounter against Mary. John has a .5 probability of winning the first time and .5 probability of winning the second time. .5 x .5 = .25, so he has a 25% chance of winning the pizza.

The same applies for Fred. He has to win first against John and then against Mary. Both times he has a .5 probability of winning, .5 x .5 = .25, so he too has a 25% chance of winning the pizza.

But Mary, she gets it easy. No matter what happens she only has to play once. Regardless of whether she's playing John or Fred she has a .5 probability of winning that rock off. So her chance of winning the pizza is 50%.

Mary has double the chance of winning overall because she only plays the winner. Great if you're Mary, bad if you're John or Fred.

Better to just draw straws.

Saturday, April 26, 2008

Mysteries of the Bell Curve Revealed

Behold the famous Bell Curve (aka 'the normal distribution'), loved by some, loathed by many, but indispensable to social sciences like psychology. The Bell Curve is what you get when you graph the distribution of things like people's height, weight, mood, IQ, extroversion, exam scores, affinity for chocolate, willingness to vote Labor, or indeed almost anything in which people vary. Again and again the same pattern appears: some people are on one extreme (e.g., extremely tall), some people are on the other extreme (e.g., extremely short), but most people are relatively average.

Why should this be? Why should such a disparate range of variables all distribute in this way? Is this some sort of divine signal? Or perhaps a government conspiracy? How bizarre that exactly the same shape should come up again and again!

Well, there's actually a good reason for the ubiquity of the Bell Curve. I only realised this a couple of years ago when I saw a nifty little exhibit at the Questacon National Science and Technology Center in Canberra. The exhibit was quite simple. Mounted on a wall, inside a glass case, were a series of pegs arranged vertically in a triangle. Visitors to Questacon were asked to drop a ball into a small shoot at the top of the case, just above the top most peg. The ball would hit the peg and bounce either left or right, then fall and hit another peg and bounce either left or right, and so on, ricocheting all the way to the bottom.

The pegs were carefully arranged so that at each level the ball had a 50/50 chance of falling either to the left or to the right. And so, each time a ball was dropped, it would take a different path through the pyramid of pegs. When the ball reached the bottom, it would fall into one of several slots lined up along the bottom; sensors in these slots, wired up to a computer, recorded the end point of each ball's journey.

The computer kept track of the outcomes. A running tally of the number of balls that had fallen into each slot was presented on a little screen in the form of a bar graph: the greater the tally for a given slot, the higher its bar.

What kind of shape do you think this graph showed after tens of thousands of ball drops? That's right, a bell curve!


Slots 1 and 9 had the smallest tallies, slots 2 and 8 had slightly more, 3 and 7 more again, 4 and 6 had even more, but slot 5 had he largest tally of all. And from this exhibit it's not dificult to see why.

For the ball to make it to slot 9 everything has to go right...literally! (On each peg the ball has to fall to the right). Thus, there's only one path to slot 9. And similarly, for the ball to reach slot 1 everything has to go left, so there's only one path to slot 1. But for slots closer to the center there are multiple routes that the ball can take. In fact, the closer a slot is to the center, the greater the number of possible routes, and the more frequently the ball reaches it. The consequence of this is a lovely bell curve.

So how does this relate to other variables?

Consider the case of an exam. How well a given student does on the exam depends on many different factors. For example:

- How much effort was put into studying
- How intelligent the student is
- How confident the student is on the day
- How much sleep the student got the night before
- Whether the student gets to the exam on time
- Whether the student is sitting next to someone who will distract them in the exam

(etc.)

For a student to get the highest possible score on the exam everything has to go right. That is, all of the factors that influence exam score have to go favorably: the student has studied hard, is intelligent, is confident, arrived on time, etc. And conversely, for a given student to get the worst possible score everything has to go wrong. For most students, however, some things will go right and some things will go wrong (in various combinations). Thus, most students will obtain a relatively average score.

So with exams, as with the pegs, we can see that there are more paths to being average than to being extreme, resulting in a bell curve. The same is true for other variables. For example, there are many different things that influence height (genes, nutrition, etc.) and so there are many more paths to being of average height than there are to being either extremely tall or short. Similarly, there are many different influences on people's love of chocolate (past experiences, advertising etc.) and so ratings of chocolate admiration should also distribute as a bell curve.

In other words, the bell curve is the distribution you get when there are multiple independent influences. And so the ubiquity of the Bell Curve in social science is not really that mysterious after all. The bell curve, my friends, is simply the signature of complexity. No wonder it pops up everywhere.

Homework

Here's a wee experiment you can do at home. Take 5 coins and toss them all together. Count up the number of heads and write it down. Repeat this 30 or more times, each time writing down the number of heads that come up.

When you're done, count up the number of times you got 5 out of 5, 4 out of 5, 3 out of 5, 2 out of 5, and 1 out of 5. Now draw a bar graph. What does the shape of the graph look like?

If you're brave, try doing this with 10, 20, or 30, coins at a time. If you're smart, just do it in Excel.