The short answer: Reinforcement is anything that follows a behavior and makes it more likely in the future. Positive reinforcement adds something the learner wants; negative reinforcement removes something unpleasant, and both strengthen behavior. Schedules of reinforcement (fixed and variable ratios and intervals) control how steady and persistent the behavior stays. Over time you thin reinforcement so the behavior survives on the natural payoffs of real life.
Positive vs negative: the most tested distinction
This is the single most confused pair on the exam, so lock it in: both positive and negative reinforcement strengthen behavior. The difference is only in what happens. Positive reinforcement adds something desirable: praise, a token, a favorite toy. Negative reinforcement removes something aversive: the loud noise stops when the learner covers their ears, the difficult worksheet is taken away when they ask for a break appropriately.
Negative reinforcement is not punishment. Punishment weakens behavior; negative reinforcement strengthens it by letting the learner escape or avoid something unpleasant. If the behavior increases, it was reinforced, period. Ask yourself: did something get added or removed, and did the behavior go up or down? That two-question check answers nearly every reinforcement question.
Types of reinforcers
- Primary (unconditioned): naturally reinforcing with no learning needed, like food and water.
- Conditioned (secondary): learned reinforcers, like tokens, praise, or stickers, that work because they have been paired with other reinforcers.
- Generalized conditioned: paired with many different reinforcers, like money or tokens exchangeable for lots of things. Hard to satiate, which makes them session gold.
The Premack principle says a high-probability behavior can reinforce a low-probability one: "First finish your worksheet, then you can play the iPad." Grandma's rule, now with a fancy name, and the exam asks about it regularly.
Schedules: the slot machine science
- Continuous (FR1): every correct response earns reinforcement. Best for brand-new skills.
- Fixed ratio (FR): reinforcement after a set number of responses (every 5th). Produces high rates with a brief pause after each payoff.
- Variable ratio (VR): reinforcement after an unpredictable number of responses (on average every 5th). The slot machine schedule: highest, steadiest response rates, very resistant to extinction.
- Fixed interval (FI): reinforcement for the first response after a set time passes. Produces the scallop pattern: pausing, then responding faster as the time approaches.
- Variable interval (VI): reinforcement for the first response after an unpredictable time. Steady, moderate responding, like checking your phone for messages.
Thinning: from constant treats to real life
Nobody gets a sticker for every email they write as an adult. Thinning gradually reduces how often reinforcement is delivered, moving from continuous to intermittent schedules, so the behavior survives on natural, real-world payoffs. Thin gradually based on data. Thin too fast and the behavior collapses; that collapse is a clue to slow down, not a sign the learner is broken.
Motivating operations: why the same reinforcer flops sometimes
Ever wonder why the iPad works magic on Monday and gets ignored on Friday? Motivating operations change how valuable a reinforcer is right now. Deprivation makes it more powerful: no iPad all morning means the iPad is gold by afternoon. Satiation makes it worthless: twenty minutes of iPad and offering more earns a shrug. Smart RBTs watch for satiation and rotate reinforcers before they go stale, and they use deprivation ethically, never by withholding necessities, just by timing. The exam frames this as the establishing operation (value up) versus the abolishing operation (value down). If a scenario says a previously great reinforcer stopped working, check satiation first.
Quick-fire review
- Both positive and negative reinforcement strengthen behavior
- Ratio = count, interval = time; variable beats fixed for steadiness
- VR = highest steady rate; FI = scalloped pattern
- Premack = favorite activity reinforces less favorite task
RBT Exam Prep