GAME THEORY / INCENTIVES / NASH
Why two people who want less jail time both get more
Each prisoner wants a shorter sentence. When both reason privately, they can still land in a cell that is worse for both than mutual silence—unless repetition or credible punishment rewires the game.
Prisoner's dilemma overview
Two rational choices can lock both into a worse cell
The prisoner's dilemma is a 2×2 game: each player independently chooses Cooperate or Defect. Payoffs here are years in jail—lower is better for each person.
When both stay silent (Cooperate), each gets a short sentence. When one confesses (Defect) while the other stays silent, the confessor walks free and the silent one gets the longest term. When both confess, both get medium terms—worse for the pair than mutual silence.
Keep these three ideas
- Payoff matrix first
The whole story lives in four cells: CC, CD, DC, DD.
- Dominant move
Given the other's choice, Defect often shortens your own sentence.
- Nash ≠ social best
Mutual Defect can be stable even though mutual Cooperate is better for both.
Step 01
Read the payoff matrix before blaming anyone
Rows are your move; columns are your partner's. Each cell shows (your years, partner's years). The numbers invent a crisp conflict: Defect is privately tempting, Cooperate is jointly kinder.
Nothing mystical happens yet—only a map of incentives. Later steps ask what a self-interested chooser does when they cannot bind the other player's hand.
What to notice
- Lower is better
These payoffs are jail years, not utility points.
- Asymmetry of temptation
Unilateral Defect beats Cooperate against a cooperator.
- Symmetric trap
Against a defector, Defect still hurts you less than Cooperate.
Step 02
Fix the other's move—then pick what shortens your sentence
Hold the partner's choice fixed and compare your two options. If they Cooperate, Defect cuts your term from 1 to 0. If they Defect, Defect cuts it from 3 to 2.
That is a dominant strategy: whatever they do, Defect privately beats Cooperate. Rationality here means 'minimize my years,' not 'maximize our joint calm.'
What changes here
- Conditional comparison
Dominance is checked column by column (or row by row).
- No mind-reading required
You do not need to know their move to prefer Defect.
- Still one-shot
This step is a single round with no future punishment.
Step 03
When both reason the same way, they meet in the worse cell
If Defect is dominant for you and for them, the predicted one-shot outcome is mutual Defect: 2 years each. Neither player wants to switch alone—switching to Cooperate would raise your term from 2 to 3.
That is a Nash equilibrium: a pair of strategies where no one benefits from a unilateral change. Stability is not the same as joint happiness.
What changes here
- Nash = no unilateral escape
From DD, a lone switch to Cooperate hurts the switcher.
- CC is unstable alone
From CC, a lone Defect pays the tempter—so CC needs trust or enforcement.
- Symmetric reasoning
The trap tightens when both use the same dominance logic.
Step 04
Add the years: the stable cell is not the kindest total
Sum both sentences. CC totals 2 years; DD totals 4. The Nash outcome wastes two extra years of freedom that mutual silence would have preserved.
This is the system's punchline: individually rational best replies can stack into a collectively worse ledger. The gap is not a math error—it is the conflict between private and joint objectives.
What changes here
- Welfare ≠ Nash
Joint total years ranks CC above DD even though DD is stable.
- Pareto idea
Moving from DD to CC can help both without hurting either.
- Needs a mechanism
Getting to CC usually needs commitment, reputation, or punishment—not wishful thinking.
Step 05
Repeat the game—and credible punishment can protect Cooperate
In a one-shot game, Defect dominates. In a repeated game with a future, strategies like tit-for-tat (copy the other's last move) or grim trigger (punish forever after a Defect) can make Cooperate incentive-compatible.
The key is credibility: the threatened punishment must be something you would actually carry out, and the shadow of the future must be long enough that today's Temptation is not worth tomorrow's Punishment.
What changes here
- Shadow of the future
Continuation value can outweigh one-round temptation.
- Tit-for-tat intuition
Start nice; mirror the last move; keep punishment proportional.
- Not magic
Finite known endings, noise, and weak monitoring can reopen the trap.
Put the system back together
Private best replies can stack into a shared loss—unless a future (or a rule) rewires the matrix
The prisoner's dilemma is not a story about stupid players. It is a story about a payoff matrix where Defect is privately safer in a one-shot meeting, so mutual Defect becomes the stable prediction—even though mutual Cooperate wastes fewer years.
Mechanisms matter: repetition, reputation, contracts, and credible punishment change the effective game. When those are weak, the Nash trap remains.
Reader checklist
- Name the four cells and whose years appear where.
- Check dominance: does Defect beat Cooperate against both partner moves?
- Separate Nash stability from joint totals.
- Ask whether the interaction is one-shot or repeated.
- If Cooperation appears, identify the enforcement or shadow of the future that supports it.
Four moves of the trap
- MatrixWhat are the numbers?
Four cells define temptation, reward, sucker, and punishment.
- DominanceWhat is privately safer?
In the classic one-shot PD, Defect beats Cooperate either way.
- NashWhere do best replies meet?
Mutual Defect can be stable even when jointly worse.
- RepeatWhat rewires incentives?
Future rounds and credible punishment can protect Cooperate.
The dilemma is not that people want harm—it is that separately sensible choices can compose into shared harm.
Teaching note: payoffs are stylized jail years (CC 1,1; DC 0,3; CD 3,0; DD 2,2). Real legal systems, noise, and incomplete monitoring change details; the qualitative conflict between private dominance and joint efficiency is the point.