An in-depth analysis of Loot Boxes
Dopamine is a neuromodulator meaning that instead of conveying a specific sensory signal, it changes how the rest of the brain processes and responds to information. It is of two types — tonic and phasic. Tonic dopamine is the slow, steady baseline level of dopamine in synapses. It represents how willing you are, in general, to pursue goals. High tonic dopamine means you’re energised and motivated whereas low tonic dopamine means apathy and disengagement. Phasic dopamine is the rapid, high amplitude burst in dopamine that occurs in response to unexpected rewards or reward-predicting stimuli; It is the learning signal. Phasic spikes are felt most strongly relative to the tonic baseline. Loot boxes exploit this mechanism by making unwanted and worthless or low value items be most prevalent, setting a low tonic baseline of the players. This makes every rare item generate a much stronger phasic spike than if medium rarity items were most common. This “drought” mechanism isn’t a flaw in game design but rather a deliberate manipulation of the environment in which a normal learning signal operates. Reward prediction error (RPE) is the mechanism that makes the brain learn from experience. The mathematical form is: RPE = Outcome - Expected value. A 3 case framework can be developed as follows:
- Better than expected: positive RPE → phasic dopamine spike above baseline
- Exactly as expected: zero RPE → no deviation from tonic baseline
- Worse than expected: negative RPE → phasic dopamine dips below baseline
This mechanism of the brain is exploited through the use of a variable ratio schedule — where reward timing is unpredictable. Under a variable ratio schedule, the brain can never fully predict when reward will come. This means RPE never reaches zero and the signal fires continuously. It is the uncertainty itself that generates the dopamine response during waiting. Continuing the discussion on variable ratio schedules, Skinner established in the 1950s that variable ratio (VR) schedules produce higher and more persistent rates of response than any other schedule — they are the schedule that maximises RPE. Under VR schedules, the next response is always potentially the one that produces reward. There is no rational point at which to stop — every response has the same expected value as the last. The brain cannot form a stable expected value because the timing is truly random. This causes dopamine to fire in anticipation on every trial, rendering the waiting period itself rewarding. Furthermore, behaviors reinforced on VR schedules are dramatically more resistant to extinction than those reinforced continuously. A child who receives a sweet every time he asks will stop asking quickly when sweets are withdrawn. But a child who receives sweets unpredictably will persist for far longer. The system has trained the brain that “no reward this time” is simply the normal precondition for eventual reward and this extinction resistance is exactly what makes loot boxes difficult to disengage from. Moving on from VR schedules, the anterior cingulate (ACC) is a key node in the mPFC and sits at the interface between cognition and emotion. Its general function is conflict monitoring and error detection; it signals when the expected and actual states of the world don’t match, and when competing response options are in conflict. Luke Clark et al. showed that near-misses — outcomes just below the win threshold — activated the ACC almost identically to actual wins, and more strongly than clear losses. Additionally, near-misses increased participants’ urge to play, even though objectively they were no better than losses. When applied to loot boxes, if a drop rate distribution clusters outcomes just below the rare item threshold — eg. “super rare” items at 1% but “near-super-rare” items at 5% — each near-miss fires the ACC as if reward just occurred. This mechanism is yet another example of a game design that structurally exploits the developing brain of a child. The insula is also an important region of the brain for this paper as it is the primary interoceptive cortex — it processes internal body states. We will however be focusing on its role in the pain of paying: the aversive signal that creates friction in spending decisions. Knutson et al. ran fMRI studies on consumer purchase decisions and found that insula activation before a purchase decision predicted whether participants would decline to buy. The insula acts as the neural source of buyer hesitation. In a well-functioning market, this is a useful circuit as it keeps expenditure roughly calibrated to perceived value. Loot boxes however, suppress the insula in the following ways:
- Premium currency abstraction: Real money is converted into in-game currency like gems or coins. This adds a cognitive conversion step in the process of purchasing. Working memory must retrieve the exchange rate, compute the real cost, and maintain this during the purchase decision — under time pressure and in a socially charged environment. The insula responds to real money and is slower to respond to abstract tokens. This step of cognitive conversion reduces the salience of the monetary cost.
- Small incremental purchases: Individual loot box prices are intentionally calibrated to stay just below the threshold at which insula activation becomes consciously aversive. £2 seems like a small amount but the aggregate spending of £200 is never shown.
- Social normalisation: When spending on loot boxes is portrayed as a common and expected activity within a peer group, the insula’s unfairness signal is suppressed by social proof.
The goal of suppression of the insula is to make spending feel like less than it actually is. Now turning to the vicious cycle created by loot boxes, the lateral habenula (LHb) is a small structure in the epithalamus. It acts as the brain’s anti-reward centre and is of most importance when explaining why loot boxes don’t let you stop. Matsumoto and Hikosaka showed that LHb neurons fire in response to stimuli predicting absence of reward or punishment — and when they fire, they inhibit VTA (Ventral Tegmental Area) dopamine neurons which is the primary dopaminergic pathway. The LHb is the mechanism by which persistent failure disengages behavior by suppressing dopamine. It acts as a rational quit signal. However, the lateral habenula presents an analogous case as this is only in a well-functioning environment and it has already been established that loot boxes do not satisfy that condition. The LHb fires on persistent negative RPE but near-misses prevent this. Due to the ACC registering near-misses as almost-wins, the RPE associated with a near-miss is far less negative than a clear loss. Therefore, near-miss distributions keep the experienced RPE from going deeply negative long enough to trigger sustained LHb firing. This effect compounded with the sustained engagement from the ACC forms the vicious cycle of spending. The cycle runs as follows: the variable ratio schedule fires RPE continuously; near-miss outcomes sustain ACC-driven motivation; those same near-miss distributions prevent RPE from falling negative enough to trigger habenula disengagement; the player cannot stop; another box is opened; and the cycle repeats. Beyond near-miss engineering, the orbitofrontal cortex plays a distinct role in value construction. Its primary role is subjective value computation: it integrates multiple sources of information to assign a single motivational value to an outcome or object. Nonetheless, this value is not objective and is highly subjective to contextual manipulation. Plassmann et al. demonstrated that OFC activity during a tasting task increased when participants believed a wine was more expensive — even when it was the same wine. Instead of providing an objective analysis, the OFC was constructing subjective value from available cues. This construction process is the exact vulnerability that loot boxes exploit. If you can manipulate the cues the OFC uses to assign value before an outcome is received, you can inflate the subjective value of that outcome without changing its objective properties. When applied to loot boxes, the opening sequence — the slow reveal, the golden beam of light, the particle effects, the sound design — is a salience intervention targeted at the OFC. It inflates the perceived value of whatever item is to be revealed before the reveal actually occurs. By the time the item appears, the OFC has already assigned it an elevated value based on the ceremony of its arrival. This is why the same cosmetic skin feels worth more when obtained from a loot box than when purchased directly. The OFC also processes social value and status. When rare items function as visible status symbols in the game — worn by high-status players, displayed in kill screens, shown off in lobbies — the OFC assigns them a value that extends beyond their functional properties. This mechanism of salience engineering manipulates the available cues to inflate OFC’s assigned value before reward evaluation begins. Lastly, the prefrontal cortex (PFC) is the brain’s executive function centre — the primary neural structure that can override impulses generated by the reward system. It is also the last brain region to reach full structural and functional maturity. The PFC — particularly the lPFC and the vmPFC — mediates impulse control, working memory, delay discounting, and metacognitive monitoring. In the context of loot boxes, it is the only structure that can evaluate the full financial context, recall previous losses, override the insula suppression and decide to stop despite ongoing RPE-driven motivation. When it works: you open a loot box, feel the dopamine pull to open another, pause, remember you’ve spent £40 already, and stop. When it’s overridden by stress, fatigue, social pressure or immaturity, the subcortical impulse wins. When immaturity is at issue, Casey et al. and Giedd et al. established through longitudinal neuroimaging that PFC myelination continues into the mid-twenties. Critics may consider this as a minor difference but that is not the case as an adolescent has fewer synaptic connections, less efficient white matter connectivity, and weaker GABAergic inhibition of limbic circuits. The structural capacity for impulse control, working memory, and metacognitive monitoring is measurably lower than in adults. The Gambling Act 2005 draws a binary line at 18. Below it: no gambling. But a 16-year-old has a measurably less capable version of the neural structure that would allow them to resist the exploitation stack described in this paper. The regulatory framework assumes a cognitive architecture that adolescents do not yet have. No single mechanism in this paper is unique to loot boxes. What makes loot boxes distinctively harmful is that all mechanisms operate simultaneously, in a socially embedded environment, with zero friction, targeting people whose primary neural defence is not yet fully developed. This stack effect sits at the core of the argument for mechanism targeted regulation.
(this is an excerpt from my policy white paper titled: A Neuroeconomic Analysis of Loot Box Regulation in the United Kingdom)