4.2 - Operant Conditioning
Principles of operant conditioning
Operant conditioning is a learning process where voluntary behaviours are strengthened or weakened through consequences, such as rewards or punishments. This occurs because individuals learn to associate their actions with outcomes, leading them to repeat behaviours that bring positive results and avoid those that do not. As a result, this form of conditioning focuses on shaping deliberate actions rather than automatic responses.
Distinction from classical conditioning
Classical conditioning and operant conditioning are both forms of associative learning, but they differ in key ways. Classical conditioning involves pairing a neutral stimulus with an unconditioned stimulus to elicit an involuntary reflex response, such as salivation in Pavlov's dogs. In contrast, operant conditioning requires the subject to perform a voluntary behaviour first, which is then followed by a consequence that influences whether the behaviour is repeated.
Key features of operant conditioning
- Voluntary behaviours - These are actions that the subject chooses to perform, unlike the automatic reflexes in classical conditioning.
- Consequences drive learning - A behaviour (called an operant) is followed by a reinforcer, which increases the likelihood of repetition; for example, a dog fetching a stick might receive a treat, encouraging it to fetch again.
- Timing of events - The subject's action comes before the consequence, whereas in classical conditioning, the stimulus precedes the response.
- Flexibility - Almost any behaviour can be conditioned through operant methods, not just reflexes.
Skinner's research on operant conditioning
B.F. Skinner, a key psychologist, investigated operant conditioning using rats to demonstrate how behaviours can be shaped through reinforcement. His work highlighted that motivation, such as hunger, influences how quickly learning occurs.
Method
Skinner used a device called a Skinner box, which contained a lever and a food tray. Pressing the lever released a food pellet. Rats were placed in the box and observed as they explored and learned.
Results
Rats quickly learnt to press the lever to obtain food, especially when hungry. Well-fed rats showed less interest, indicating that the reinforcer's value affects learning speed. Skinner also found that stopping pain (e.g., through a response that ends discomfort) acted as a strong reinforcer, leading to rapid learning and slow forgetting.
Conclusions
These findings showed that reinforcers increase the repetition of behaviours, and operant conditioning can apply to a wide range of actions.
Types of reinforcement and punishment
Reinforcement increases the likelihood of a behaviour being repeated, while punishment decreases it. Reinforcers can be positive or negative, and some are primary (directly satisfying basic needs) or secondary (gained through exchange for primary ones). Understanding these helps explain how behaviours are maintained or changed.
Positive reinforcement
This involves adding a desirable stimulus after a behaviour to encourage its repetition. For example, giving money for completing a task acts as a powerful reinforcer in many societies.
Negative reinforcement
This strengthens a behaviour by removing an unpleasant stimulus. Negative reinforcement is distinct from punishment: it rewards by eliminating something bad, whereas punishment adds something bad to discourage behaviour.
Secondary reinforcement
These are items or tokens that can be exchanged for primary reinforcers.
Punishment
This involves applying an unpleasant consequence to reduce a behaviour's occurrence. Severe punishment can deter actions.
Schedules of reinforcement
Schedules of reinforcement refer to the patterns or timings of delivering reinforcers, which affect how quickly behaviours are learnt and how resistant they are to extinction (stopping when reinforcement ends). Skinner identified five main schedules, showing that intermittent reinforcement often leads to more persistent behaviours than constant rewards.
Continuous reinforcement
The behaviour is reinforced every time it occurs. This leads to quick learning but fast extinction if reinforcement stops, as the subject expects a reward each time.
Fixed-ratio reinforcement
Reinforcement is given after a set number of correct responses, such as every sixth or ninth time. This produces high response rates.
Variable-ratio reinforcement
The number of responses needed for reinforcement varies unpredictably. This creates very persistent behaviours, as the subject keeps responding in hope of the next reward.
Fixed-interval reinforcement
Reinforcement is provided after a fixed time period, provided at least one correct response occurs during that interval.
Variable-interval reinforcement
The time between reinforcements changes unpredictably.
Switching from continuous to intermittent schedules makes behaviours harder to extinguish, as the subject does not expect a reward every time and continues responding longer.
Applications in behaviour modification
Behaviour modification applies operant conditioning principles to change unwanted behaviours or encourage desired ones, using reinforcement strategically. This is common in settings like schools and psychiatric institutions, where it promotes positive actions.
Behaviour shaping
Behaviour shaping involves gradually reinforcing successive approximations of a target behaviour until the full behaviour is achieved. This breaks complex skills into steps, making learning manageable. For example, teaching a child to tidy their room might start with reinforcing putting toys in a designated box, then adding placing books on a shelf, and finally requiring the bed to be made before a reward. Each step builds on the last, requiring improvement for reinforcement.
Token economy
A token economy is a system where tokens are given as secondary reinforcers for appropriate behaviours, which can be exchanged for primary rewards like sweets or privileges. This is often used in psychiatric wards to encourage positive actions.
Key features of token economies:
- Method - Tokens are awarded for desired behaviours and collected to trade for goods or activities.
- Effectiveness - They successfully change behaviours while in use, but effects may fade without tokens unless variable-ratio schedules are introduced to prolong responses.
- Limitations - Token economies treat symptoms rather than root causes, and behaviours may revert when the system ends.
In schools, similar principles apply: praise reinforces good behaviour, while consequences weaken disruptions.