2.8 - Inferential Statistics
The purpose of inferential statistics
Inferential statistics are tools used in psychological research to draw conclusions about a larger population based on data from a sample. Unlike descriptive statistics, which simply summarise data (such as means or percentages), inferential statistics help researchers test hypotheses and determine whether observed differences or patterns are likely to reflect real effects rather than random variation.
This process allows psychologists to infer whether results from a study can be generalised beyond the participants involved. For example, if a study examines how stress affects memory in a group of students, inferential statistics help decide if these findings apply to the wider population of students.
Key aims of inferential statistics
- Testing predictions - To evaluate whether the independent variable (the factor manipulated in an experiment, such as exposure to stress) has a genuine impact on the dependent variable (the factor measured, such as memory performance).
- Assessing chance - To calculate the probability that results occurred due to random factors rather than the independent variable.
- Generalising findings - To determine if sample results are valid for the entire target population, supporting evidence-based conclusions in psychology.
Probability and significance levels
Probability in inferential statistics refers to the likelihood that study results are due to chance rather than the effect of the independent variable. Psychologists use significance levels to set a threshold for deciding if results are meaningful.
A common significance level is 0.05 (or 5%), meaning results are considered significant if there is a 5% or lower chance they occurred randomly. This provides at least 95% confidence that the independent variable caused the observed effects.
Understanding probability notation
- p - Represents the probability that results are due to chance.
- p > 0.05 - Indicates the probability of chance is greater than 5%, so results are not significant.
- p ≤ 0.05 - Indicates the probability of chance is 5% or less, so results are significant.
In some cases, stricter levels are used:
- 0.01 (1%) - For research challenging established theories, providing 99% confidence.
- 0.001 (0.1%) - For high-stakes studies, like drug trials, offering even greater certainty.
These adjustments help ensure robust conclusions but can introduce risks, as explored in later sections.
Deciding to accept or reject hypotheses
Once data is analysed using inferential statistics, researchers decide whether to accept or reject their hypotheses. The experimental hypothesis (also called the alternative hypothesis) predicts a relationship between variables, while the null hypothesis assumes no real effect, attributing any differences to chance.
Steps in hypothesis decision-making
- Calculate the probability (p-value) using a statistical test.
- Compare the p-value to the chosen significance level.
- If p ≤ significance level (e.g., 0.05), reject the null hypothesis and accept the experimental hypothesis – the results are significant, suggesting the independent variable influenced the dependent variable.
- If p > significance level, accept the null hypothesis and reject the experimental hypothesis – the results are not significant, likely due to chance.
For instance, in a study testing if attending a psychology revision conference improves exam scores, a significant result (p ≤ 0.05) would mean rejecting the null hypothesis and concluding the conference had a real impact.
Type 1 and Type 2 errors
Errors can occur when deciding hypotheses, even with careful statistical analysis. These are known as Type 1 and Type 2 errors, and they arise from inappropriate significance levels.
Type 1 error
- Definition - Rejecting the null hypothesis when it is actually true, meaning results are wrongly attributed to the independent variable instead of chance.
- Common cause - Using a significance level that is too lenient, such as 0.10 (10%), which increases the risk of false positives.
- Implication - This can lead to accepting flawed conclusions, such as claiming a treatment works when it does not.
Type 2 error
- Definition - Accepting the null hypothesis when it is actually false, meaning a real effect is missed and attributed to chance.
- Common cause - Using a significance level that is too strict, such as 0.001 (0.1%), which increases the risk of false negatives.
- Implication - This might overlook important findings, such as failing to detect that a therapy is effective.
Choosing the right significance level involves balancing these risks based on the study's context.
Comparing observed and critical values in statistical tests
Statistical tests generate an observed value (calculated from the study's data) that must be compared to a critical value (a threshold from statistical tables) to determine significance. The comparison rule depends on the specific test used.
Key concepts in value comparison
- Observed value - The result from applying the statistical test to your data.
- Critical value - A benchmark value based on the significance level, sample size, and test type; it indicates the cutoff for significance.
- Decision rules - For some tests (e.g., t-tests or chi-square), the observed value must be equal to or greater than the critical value to be significant. For others (e.g., Mann-Whitney U or Wilcoxon), it must be equal to or less than the critical value.
If the comparison shows significance, you can be confident the results are not due to chance.