1.8 - Quantitative Data Analysis
Qualitative and quantitative data
Data in psychological research can be broadly divided into two types: quantitative and qualitative.
Quantitative data
Quantitative data involves measuring phenomena using numbers, such as counting the number of words recalled in a memory experiment or the number of participants who obeyed in a study like Milgram's obedience experiment.
Strengths and limitations of quantitative data:
- Strengths - Quantitative methods are often stricter and more controlled, leading to data that can be easily replicated. This enhances reliability and supports the scientific status of the findings.
- Limitations - Can be criticised for being narrow and unrealistic, as it often focuses only on small fragments of behaviour.
Qualitative data
Qualitative data involves describing experiences, meanings, and explanations, such as exploring how memory works through interviews or asking participants why they administered high levels of shocks in Milgram's experiment.
Strengths and limitations of qualitative data:
- Strengths - Qualitative approaches typically involve less control and are conducted in more natural settings, which can produce more valid data.
- Limitations - Often has low reliability due to difficulties in replication, and can be subjective. Subjectivity occurs when researchers must organise and select from the information.
Measures of central tendency
Measures of central tendency are descriptive statistics that summarise a set of data by identifying a single value that represents the centre of the distribution. The three main types are the mean, median, and mode.
Mean
The mean is the arithmetic average of a set of scores. To calculate it, add up all the scores in a condition and divide by the number of participants or scores.
Formula for mean:
Where:
- = Mean
- = Sum of all scores
- = Number of scores
Median
The median is the middle value in a set of scores when arranged in ascending order. For an odd number of scores, it is the central value. For an even number, it is the average of the two central values.
Mode
The mode is the most frequently occurring score in a dataset. A dataset can have one mode, multiple modes, or no mode if all scores are unique.
Worked example - Calculating the mean
Four test results are: 16, 19, 21, 24. Calculate the mean score.
Step 1: Identify the values
- Scores: 16, 19, 21, 24
Step 2: Sum the scores
Sum = 16 + 19 + 21 + 24 = 80
Step 3: Apply the formula
Worked example - Calculating the median and mode
Find the median and mode for the scores: 5, 8, 8, 12, 15, 17, 22.
Step 1: Arrange the scores
Ordered scores: 5, 8, 8, 12, 15, 17, 22
Step 2: Calculate the median
With seven scores (odd number), the median is the fourth value: 12
Step 3: Identify the mode
The score 8 appears twice, more than any other, so the mode is 8
Measures of dispersion
Measures of dispersion describe how spread out the scores in a dataset are, indicating variability around the central tendency. This will tell us whether our scores are clustered closely round the mean or are widely scattered.
Range
The range is a simple measure of dispersion calculated by subtracting the lowest score from the highest score.
Standard deviation
Standard deviation (often abbreviated as SD) is a more precise measure of dispersion that accounts for every score in the dataset. It gives us an idea of how much, on average, scores in a distribution differ from the mean. A small SD shows the mean to be a good representation of the data as a whole, while a large SD compared to this might be showing that the mean is not representative of how the group scored as a whole.
Formula for standard deviation:
Where:
- = Standard deviation
- = Each score
- = Mean
- = Number of scores
- = Sum of
Worked example - Calculating the range
Calculate the range for the scores: 5, 15, 20, 25, 30, 35, 40, 45, 50, 60.
Step 1: Identify the extremes
- Lowest score = 5
- Highest score = 60
Step 2: Apply the calculation
Range = 60 - 5 = 55
Worked example - Calculating standard deviation
Calculate the standard deviation for the scores: 3, 6, 9, 10, 11, 14, 17.
Step 1: Calculate the mean
Sum = 3 + 6 + 9 + 10 + 11 + 14 + 17 = 70
Step 2: Calculate deviations and squares
| x | x - | |
|---|---|---|
| 3 | -7 | 49 |
| 6 | -4 | 16 |
| 9 | -1 | 1 |
| 10 | 0 | 0 |
| 11 | 1 | 1 |
| 14 | 4 | 16 |
| 17 | 7 | 49 |
Sum of squares = 132
Step 3: Apply the formula
Step 4: Interpretation
An SD of approximately 4.7 shows moderate spread around the mean of 10, indicating the scores are somewhat varied but not extremely so.
The normal distribution
The normal distribution is a symmetrical, bell-shaped pattern that often appears when measuring naturally occurring phenomena, such as heights, weights, or IQ scores in large samples.
Characteristics of the normal distribution
- The mean, median, and mode all coincide at the highest point of the curve.
- It is symmetrical, with the pattern of scores identical on both sides of the centre.
- A large number of scores fall relatively close to the mean on either side. As the distance from the mean increases, the scores become fewer.
- The tails extend towards infinity but never touch the horizontal axis.
Skewed distributions
Not all distributions are normal; some are skewed, meaning they are asymmetrical with a tail extending in one direction. Skewed distributions often occur with small or biased samples, and in them, the mean, median, and mode have different values.
Types of skewed distributions
- Positively skewed - The tail extends to the right (positive direction), with more low scores and a few high outliers.
- Negatively skewed - The tail extends to the left (negative direction), with more high scores and a few low outliers.