Dopamine signals for reward value and riskbasic and recent data
If every decision you make is secretly a probability calculation, and if your brain is constantly comparing what it expected to get against what it actually got, then there must be a signal that carries that comparison — a correction signal, firing moment to moment, updating the ledger. Neuroscientists found it. It lives in a cluster of cells in the midbrain, it fires in bursts lasting milliseconds, and its currency is dopamine. The story of how we know this begins with a surprisingly blunt set of early experiments. Lesion studies, electrical self-stimulation work, and drug addiction research all pointed the same direction: damage or hijack the midbrain dopamine systems, and reward-seeking behavior collapses or spirals. These were suggestive findings, but they couldn't tell you what individual neurons were actually doing. For that, Wolfram Schultz and colleagues turned to a more precise tool — single-neuron recording in awake, behaving monkeys. Small electrodes placed near identified dopamine neurons measured the electrical spikes each cell emits while monkeys experienced cues, rewards, and other events in carefully designed tasks. One cell at a time. That's the method. And what it revealed changed how we think about learning itself.
The foundational finding is this: about seventy-five to eighty percent of midbrain dopamine neurons produce brief, stereotyped bursts — onset latencies under one hundred milliseconds, durations under two hundred milliseconds — following temporally unpredicted rewards. But these bursts don't simply announce "reward." They announce deviation from expectation. A reward better than predicted triggers a burst, a phasic activation. A reward that arrives exactly as predicted triggers essentially nothing. A reward that was expected but doesn't come triggers a depression — activity drops below baseline. Three outcomes, three responses: activation, silence, suppression. That's a bidirectional signal, and bidirectionality is precisely what learning theory requires. Schultz connects this explicitly to the Rescorla-Wagner model and to temporal-difference reinforcement learning — two frameworks, one from classical conditioning research and one from computer science, that both posit a teaching signal equal to the difference between what was predicted and what was obtained. The dopamine response is that signal, implemented in biology. Several experimental tests confirm it.
In blocking experiments — where a stimulus is paired with a reward that is already fully predicted by another cue — the added stimulus fails to acquire any predictive power, and dopamine neurons show no response to it. In conditioned inhibition tests, a stimulus that predicts the absence of reward fails to produce a depression when reward is indeed omitted because no error occurred. These aren't incidental findings. They're the formal fingerprints of a prediction-error code. The signal is also quantitatively rich. Phasic activations to reward-predicting cues scale monotonically with reward probability and with reward magnitude. When different combinations of probability and magnitude yield the same expected value — the product of the two — the neuronal responses are indistinguishable. The neurons are computing expected value, not just tracking one dimension of it. Schultz also reports a striking normalization effect: delivering a larger reward across several different binary distributions produced the same dopamine activation despite a tenfold range in absolute magnitudes across those distributions. The neuron wasn't reporting the raw difference between obtained and expected reward.
It was reporting that difference divided by the standard deviation of the predicted reward distribution — the prediction error measured in units of variability. An outcome two standard deviations above expectation triggers a similar response whether the absolute rewards are large or small. The neuron is asking: how surprising was this, relative to how variable this situation usually is? The timing of the response matters as much as its magnitude. Early in training, dopamine neurons fire when the reward arrives. As learning proceeds, that activation migrates back to the predictive cue. By the time the cue reliably predicts the reward, the response to the reward itself has largely disappeared — the prediction is met, so no error is generated. If the reward is then delayed, firing drops at the moment it was expected and rises again when it finally arrives. This temporal shift matches the principal characteristics of temporal-difference models, where prediction errors propagate back through time as associations are learned. Now complicate the picture. The clean prediction-error story is real, but it's not the whole story of dopamine signaling. Schultz documents at least two other distinct signals carried by smaller fractions of neurons.
Physically intense stimuli — loud sounds, bright flashes — produce phasic activations in substantial proportions of dopamine neurons, especially when those stimuli are novel. This response can persist for months if the intensity remains high. But Schultz is careful to note that this physically salient response is distinct from the reward value signal. It doesn't behave like a general alerting signal either, because many genuinely attention-grabbing events — reward omission, conditioned inhibitors — produce depressions rather than activations. Salience and reward are related, but they are not the same thing. Aversive events tell a different story again. Primary aversive stimuli — air puffs, hypertonic saline, shock — do produce excitatory responses, but only in a small fraction of neurons: across different studies, figures range from roughly eleven to twenty-nine percent. Conditioned aversive stimuli show similar numbers. The majority of neurons exposed to aversive events are either suppressed or unaffected. In population averages, the rare activations are often cancelled by the more common depressions. Dopamine's answer to punishment is predominantly silence, not a shout.
This is a critical point: if dopamine were simply an arousal or attention signal, punishers — which are highly attention-grabbing — should drive robust activations. They largely don't. Schultz also notes frequent brief activations to non-rewarding stimuli in general, but these are shorter than reward responses and often followed by depressions. Their prevalence depends heavily on task context — how many of the stimuli in a session are rewarded, whether appetitive and aversive cues share the same sensory modality — pointing to stimulus generalization and pseudoconditioning as contributors. The second distinct signal Schultz describes concerns risk. In decision theory, a reward distribution has two key properties: its expected value and its variance — how spread out the possible outcomes are. When reward probability varies from zero to one with fixed magnitude, risk follows an inverted-U function that peaks at probability 0.5, where gains and misses are equally likely and uncertainty is maximal. About one third of dopamine neurons show a slow, moderate activation that rises during the interval between a predictive cue and the reward, varies monotonically with risk, and peaks around that 0.5 probability point. This signal has a latency of roughly one second and a much slower time course than the fast phasic value response. It appears to be a distinct signal, not an echo of the prediction-error code.
Put these two signals together and something striking emerges. Dopamine neurons carry both of the classical variables in decision theory — expected value and risk — but on different timescales. The fast phasic signal, present in roughly three-quarters of neurons, encodes expected value and prediction error, scaled by variance. The slower risk signal, present in about a third of neurons, encodes uncertainty itself. Schultz also notes a possible receptor-level implication: the lower-magnitude risk signal would tend to produce lower dopamine concentrations at synaptic sites, preferentially engaging high-affinity D2 receptors, while the larger phasic value response would recruit lower-affinity D1 receptors. The two signals may literally speak to different downstream audiences. The implications reach well beyond the monkey laboratory. Disruption of dopamine burst firing impairs appetitive learning and fear conditioning. Temporal discounting — the tendency to value sooner rewards more than later ones — is mirrored in dopamine responses: choice indifference values for delays of four, eight, and sixteen seconds fall by roughly twenty-five, fifty, and seventy-five percent relative to a two-second delay, and dopamine responses decrease in the same hyperbolic pattern. The neurons are not just tracking reward in the abstract; they are tracking subjective value as it changes with time.
What Schultz's review ultimately demonstrates is that the midbrain dopamine system is running something close to the algorithms that economists and computer scientists developed independently to describe rational decision-making under uncertainty. Expected value, prediction error, variance scaling, temporal discounting — the mathematics is there, implemented in cells firing for a fraction of a second in a brain the size of a walnut. The method was humble: one electrode, one neuron, one monkey, one trial at a time. The answer it returned was not humble at all. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Stress and substance use disorders: risk, relapse, and treatment outcomes
- Induction of a common microglia gene expression signature by aging and neurodegenerative conditions: a co-expression meta-analysis
- Spatial Relational Memory Requires Hippocampal Adult Neurogenesis
- Hyperactive neuronal autophagy depletes BDNF and impairs adult hippocampal neurogenesis in a corticosterone-induced mouse model of depression
- Microglial cell dysregulation in brain aging and neurodegeneration
- Harmonics of Circadian Gene Transcription in Mammals