Urban Scaling and Its DeviationsRevealing the Structure of Wealth, Innovation and Crime across Cities

Luís M. A. Bettencourt, José Lobo, Deborah Strumsky, Geoffrey B. WestView original
OverviewBalancedalloy voice
If you've ever scrolled through a list of "best cities" and thought, huh, that just looks like a ranking of the biggest places, you're not imagining it. Per-capita scores treat a city of fifty thousand and a city of five million as if the only difference is how many people you divide by. But big cities don't just have more people—they change how life works. As Luís Bettencourt and colleagues showed, many things we care about in cities—wealth, innovation, even crime—don't rise one-to-one with population. They rise faster. In plain terms, if you double the population, you don't just double the output; you get a bonus. For numerous socioeconomic measures, that bonus works out to something like fifteen percent more per person with each doubling. Gross Metropolitan Product, or the total economic output of a metro area, is a good example. When they fit a simple relationship of the form "urban quantity equals a constant times population raised to a power," the power for output is about 1.13. Specifically, it is 1.126 with a tight confidence band. Bigger cities, on average, are more productive per person. That idea sounds abstract, so let’s say it out loud the way the paper does. Take any city metric, call it Y. Take population, call it N. The data align well with a power law: Y equals some baseline level times N raised to b, that exponent I just mentioned. If b is greater than one, you have superlinear scaling—each extra person adds a bit more than their share. If b were exactly one, you'd get a flat per-capita picture. But for measures like income or total output, b isn’t one. That's why per-capita comparisons warp our view. They combine two effects: the universal bump that comes from agglomeration and the true local performance that's special to a place. So, Bettencourt's team changes the baseline. Instead of normalizing by population, they first estimate the size effect and then ask what's left. The unit of analysis is the U.S. metropolitan statistical area—the standardized regions that the Office of Management and Budget draws around an urban core and its commuter belt. They pull output and personal income from the Bureau of Economic Analysis. They gather violent crime data from the FBI's Uniform Crime Reports, and patents from the U.S. Patent and Trademark Office. Then they do something simple and powerful. They fit the logarithm of each quantity against the logarithm of population for the same year, across all metros, using ordinary least squares. The line that emerges is the size-dependent "expected" value. The vertical deviation for each city—how far above or below that line the city sits—that's the scale-adjusted metropolitan indicator, or SAMI. It's a clean, dimensionless measure of local over- or under-performance that isn’t just a proxy for being big. Now, what do those residuals look like? Here’s a neat statistical twist. If you plot the SAMIs across all cities for a given measure in a specific year, they don’t make a smooth bell curve. They have sharper peaks and fatter tails, resembling an exponential. Statisticians call that a Laplace distribution. The spread of that distribution differs by what you're measuring. For innovation, the width is wide—on the order of four-tenths on their scale—which indicates that cities vary significantly in patents beyond what size predicts. For violent crime, the width is around two-tenths. For income and total output, it’s narrower, roughly a tenth. That spread aligns with a larger point about how much population size explains. If you ask, "How much of the ups and downs across cities can I predict from population alone?" the answer is, for some measures, almost everything. Income and total output are about ninety-five percent predictable from size. Violent crime is a little looser, roughly in the mid-eighties. Patents are the outlier; only about two-thirds of the variance is accounted for by size, which means local factors loom much larger for innovation. Once you have SAMIs, rankings become fascinating. The lists no longer tilt toward the giants. On a size-adjusted basis, New York looks average in income and output, and only modestly inventive. San Francisco shines in income and patents but doesn’t appear outlandishly safe or dangerous after accounting for size. Smaller places that never top per-capita lists suddenly show up: Bridgeport for income, Corvallis and San Jose for patents, and towns like Logan or Bangor for safety. It's the urban equivalent of grading on a curve—but the curve is shaped by the physics of cities: how interaction, networks, and infrastructure change with scale. The real magic kicks in when you observe these SAMIs over time. The team builds decades-long time series—income data from the late 1960s through two thousand six, and patents from the mid-1970s through two thousand five. Then they ask a deceptively simple question: if a city is doing unusually well relative to its size this year, does that advantage stick? The answer is yes, and for a long time. When they compute autocorrelation—the tendency for this year's deviation to resemble next year's and the year after's—they discover long decay times. Income retains its local edge or drag for something like thirty-five years. Patents maintain their signature for about nineteen years. This means that "local character," in the statistical sense, has memory measured in generations, not election cycles. Space adds another layer. You might expect neighboring cities to rise and fall together, and up close they do. When you take equal-time similarities in SAMIs and plot them against the distance between cities, correlations are strongest nearby and decrease with distance. But by the time you're examining pairs more than about two hundred kilometers apart, the average correlation is essentially nonexistent. The spread around that average is significant—individual pairs can still be quite coordinated or quite opposed. But overall, geography is not destiny beyond a couple of hours’ drive. If geography isn't the glue, what is? The authors focus on the idea of kindred trajectories. They compute how similar pairs of cities are in their SAMI histories—patents with patents, income with income—and translate that similarity into a "decorrelation distance." That number is small when two cities have moved together over time and larger when they haven’t. Then they cluster. Five broad families emerge for income among the largest metros, and the groupings make intuitive sense. There’s a high-tech arc—San Francisco, San Jose, Minneapolis, Denver, Seattle—that shares an innovation-forward model. There’s a market and transportation hub family—think Pittsburgh, Cincinnati, Memphis, Birmingham—with a different rhythm. The takeaway isn’t that every city in a cluster is the same; it’s that their size-adjusted fortunes have pulsed in sync for decades, likely tied to shared industrial structures and history more than simple proximity. Do the classical relationships among wealth, innovation, and crime survive after you strip out the size effect? Weakened, yes; erased, no. Cities that exceed expectations on income tend to exceed them on patents as well. Places that lag on both generally exhibit higher violent crime than their size would suggest. But the strength of those connections is modest compared to the significant role that population plays. You can feel that intuitively: some cities excel at generating ideas but haven’t translated that into broad income gains yet. Others create wealth without producing patents at the same pace. The SAMI perspective allows you to see those mismatches clearly. A quick note on infrastructure, because it completes the picture. Not everything in cities grows superlinearly. Roads, electrical cables, the infrastructure of the place—those tend to show economies of scale. The same power-law form applies, but the exponent is below one. That indicates you need proportionally less infrastructure per person as you grow larger. Together, that constitutes the urban double dividend: more output and innovation per person, and less infrastructure per person. But again, that’s just the baseline. The interesting part is who beats the baseline, by how much, and for how long. There’s also some caution in how you should use these metrics. Bettencourt and colleagues are explicit that SAMIs aren't magic policy grades. They are deviations from a statistical baseline. And those deviations can encompass many local realities—things like industrial mix, housing costs, climate, spatial layout—that don’t reduce to simple levers. The fact that SAMIs for income and patents revert only over decades means a single year's snapshot won’t tell you if last summer's program "worked." It will reveal whether your city sits above or below where a city of your size typically falls and whether that position is part of a long arc. All of this rests on a very standard foundation: U.S. metro definitions from the Office of Management and Budget; output and income data from the Bureau of Economic Analysis; violent crime statistics from the FBI; and patents from the U.S. Patent and Trademark Office. Log-log regressions are used to estimate the scaling line, with residuals serving as the SAMIs. The technique is transparent. The patterns are robust across years. And the diagnostic of Laplace-shaped residuals—a sharp peak with heavy tails—appears repeatedly. It indicates that exceptional over- and under-performance is more common than a bell curve would predict. Let me circle back to the motivating gripe: the per-capita league table. Once you see the size effect, you can’t unsee it. Per-capita rankings almost inevitably crown the giants because they incorporate the superlinear bonus of being big. Switching to SAMIs doesn’t just reshuffle the list; it changes the question from "Who has the highest number?" to "Who's outpacing what cities of this size normally do?" That reframing is subtle and profound. It reveals Bridgeport's unusual income position without pretending it’s New York. It highlights San Jose's inventive streak without attributing that to sheer headcount. And it stops penalizing dense places for having more of everything, including the bad. That is precisely what superlinear scaling predicts. The bigger story here is that cities share a universal grammar. Double the population and certain metrics—output, wages, inventions, even crime—get longer in a remarkably predictable way. Within that grammar, each city writes its own paragraphs. The SAMIs provide a way to read those paragraphs without being dazzled or misled by word count. They are empirical, they are comparable, and, because of that long memory, they represent the beginning of a diagnosis, not the end. Where does this go next? The team hints at obvious extensions. Take the same method to other countries, new decades, more indicators, and see which pieces of the story travel. Use the clustering to choose peers that make sense for benchmarking, even if they’re in a different time zone. And, cautiously, combine SAMIs with process data—on migration, industry churn, housing supply. Use that to pry apart which local levers move which deviations. But even before that, there's a very practical mental shift available to anyone reading a city chart tomorrow. Ask what the size baseline would imply, and then determine who's truly above or below it. That’s the sound of the universal and the local, harmonizing—and sometimes clashing—inside the urban chorus.

If you've ever scrolled through a list of "best cities" and thought, huh, that just looks like a ranking of the biggest places, you're not imagining it. Per-capita scores treat a city of fifty thousand and a city of five million as if the only difference is how many people you divide by. But big cities don't just have more people—they change how life works.

As Luís Bettencourt and colleagues showed, many things we care about in cities—wealth, innovation, even crime—don't rise one-to-one with population. They rise faster. In plain terms, if you double the population, you don't just double the output; you get a bonus.

For numerous socioeconomic measures, that bonus works out to something like fifteen percent more per person with each doubling. Gross Metropolitan Product, or the total economic output of a metro area, is a good example. When they fit a simple relationship of the form "urban quantity equals a constant times population raised to a power," the power for output is about 1.13.

Specifically, it is 1.126 with a tight confidence band. Bigger cities, on average, are more productive per person.

That idea sounds abstract, so let’s say it out loud the way the paper does. Take any city metric, call it Y. Take population, call it N.

The data align well with a power law: Y equals some baseline level times N raised to b, that exponent I just mentioned. If b is greater than one, you have superlinear scaling—each extra person adds a bit more than their share. If b were exactly one, you'd get a flat per-capita picture.

But for measures like income or total output, b isn’t one. That's why per-capita comparisons warp our view. They combine two effects: the universal bump that comes from agglomeration and the true local performance that's special to a place.

So, Bettencourt's team changes the baseline. Instead of normalizing by population, they first estimate the size effect and then ask what's left. The unit of analysis is the U.S. metropolitan statistical area—the standardized regions that the Office of Management and Budget draws around an urban core and its commuter belt.

They pull output and personal income from the Bureau of Economic Analysis. They gather violent crime data from the FBI's Uniform Crime Reports, and patents from the U.S. Patent and Trademark Office.

Then they do something simple and powerful. They fit the logarithm of each quantity against the logarithm of population for the same year, across all metros, using ordinary least squares. The line that emerges is the size-dependent "expected" value.

The vertical deviation for each city—how far above or below that line the city sits—that's the scale-adjusted metropolitan indicator, or SAMI. It's a clean, dimensionless measure of local over- or under-performance that isn’t just a proxy for being big.

Now, what do those residuals look like? Here’s a neat statistical twist. If you plot the SAMIs across all cities for a given measure in a specific year, they don’t make a smooth bell curve.

They have sharper peaks and fatter tails, resembling an exponential. Statisticians call that a Laplace distribution. The spread of that distribution differs by what you're measuring.

For innovation, the width is wide—on the order of four-tenths on their scale—which indicates that cities vary significantly in patents beyond what size predicts. For violent crime, the width is around two-tenths. For income and total output, it’s narrower, roughly a tenth.

That spread aligns with a larger point about how much population size explains. If you ask, "How much of the ups and downs across cities can I predict from population alone?" the answer is, for some measures, almost everything. Income and total output are about ninety-five percent predictable from size.

Violent crime is a little looser, roughly in the mid-eighties. Patents are the outlier; only about two-thirds of the variance is accounted for by size, which means local factors loom much larger for innovation.

Once you have SAMIs, rankings become fascinating. The lists no longer tilt toward the giants. On a size-adjusted basis, New York looks average in income and output, and only modestly inventive.

San Francisco shines in income and patents but doesn’t appear outlandishly safe or dangerous after accounting for size. Smaller places that never top per-capita lists suddenly show up: Bridgeport for income, Corvallis and San Jose for patents, and towns like Logan or Bangor for safety. It's the urban equivalent of grading on a curve—but the curve is shaped by the physics of cities: how interaction, networks, and infrastructure change with scale.

The real magic kicks in when you observe these SAMIs over time. The team builds decades-long time series—income data from the late 1960s through two thousand six, and patents from the mid-1970s through two thousand five. Then they ask a deceptively simple question: if a city is doing unusually well relative to its size this year, does that advantage stick?

The answer is yes, and for a long time. When they compute autocorrelation—the tendency for this year's deviation to resemble next year's and the year after's—they discover long decay times. Income retains its local edge or drag for something like thirty-five years.

Patents maintain their signature for about nineteen years. This means that "local character," in the statistical sense, has memory measured in generations, not election cycles.

Space adds another layer. You might expect neighboring cities to rise and fall together, and up close they do. When you take equal-time similarities in SAMIs and plot them against the distance between cities, correlations are strongest nearby and decrease with distance.

But by the time you're examining pairs more than about two hundred kilometers apart, the average correlation is essentially nonexistent. The spread around that average is significant—individual pairs can still be quite coordinated or quite opposed. But overall, geography is not destiny beyond a couple of hours’ drive.

If geography isn't the glue, what is? The authors focus on the idea of kindred trajectories. They compute how similar pairs of cities are in their SAMI histories—patents with patents, income with income—and translate that similarity into a "decorrelation distance." That number is small when two cities have moved together over time and larger when they haven’t.

Then they cluster. Five broad families emerge for income among the largest metros, and the groupings make intuitive sense. There’s a high-tech arc—San Francisco, San Jose, Minneapolis, Denver, Seattle—that shares an innovation-forward model.

There’s a market and transportation hub family—think Pittsburgh, Cincinnati, Memphis, Birmingham—with a different rhythm. The takeaway isn’t that every city in a cluster is the same; it’s that their size-adjusted fortunes have pulsed in sync for decades, likely tied to shared industrial structures and history more than simple proximity.

Do the classical relationships among wealth, innovation, and crime survive after you strip out the size effect? Weakened, yes; erased, no. Cities that exceed expectations on income tend to exceed them on patents as well.

Places that lag on both generally exhibit higher violent crime than their size would suggest. But the strength of those connections is modest compared to the significant role that population plays. You can feel that intuitively: some cities excel at generating ideas but haven’t translated that into broad income gains yet.

Others create wealth without producing patents at the same pace. The SAMI perspective allows you to see those mismatches clearly.

A quick note on infrastructure, because it completes the picture. Not everything in cities grows superlinearly. Roads, electrical cables, the infrastructure of the place—those tend to show economies of scale.

The same power-law form applies, but the exponent is below one. That indicates you need proportionally less infrastructure per person as you grow larger. Together, that constitutes the urban double dividend: more output and innovation per person, and less infrastructure per person.

But again, that’s just the baseline. The interesting part is who beats the baseline, by how much, and for how long.

There’s also some caution in how you should use these metrics. Bettencourt and colleagues are explicit that SAMIs aren't magic policy grades. They are deviations from a statistical baseline.

And those deviations can encompass many local realities—things like industrial mix, housing costs, climate, spatial layout—that don’t reduce to simple levers. The fact that SAMIs for income and patents revert only over decades means a single year's snapshot won’t tell you if last summer's program "worked." It will reveal whether your city sits above or below where a city of your size typically falls and whether that position is part of a long arc.

All of this rests on a very standard foundation: U.S. metro definitions from the Office of Management and Budget; output and income data from the Bureau of Economic Analysis; violent crime statistics from the FBI; and patents from the U.S. Patent and Trademark Office. Log-log regressions are used to estimate the scaling line, with residuals serving as the SAMIs.

The technique is transparent. The patterns are robust across years. And the diagnostic of Laplace-shaped residuals—a sharp peak with heavy tails—appears repeatedly.

It indicates that exceptional over- and under-performance is more common than a bell curve would predict.

Let me circle back to the motivating gripe: the per-capita league table. Once you see the size effect, you can’t unsee it. Per-capita rankings almost inevitably crown the giants because they incorporate the superlinear bonus of being big.

Switching to SAMIs doesn’t just reshuffle the list; it changes the question from "Who has the highest number?" to "Who's outpacing what cities of this size normally do?" That reframing is subtle and profound. It reveals Bridgeport's unusual income position without pretending it’s New York. It highlights San Jose's inventive streak without attributing that to sheer headcount.

And it stops penalizing dense places for having more of everything, including the bad. That is precisely what superlinear scaling predicts.

The bigger story here is that cities share a universal grammar. Double the population and certain metrics—output, wages, inventions, even crime—get longer in a remarkably predictable way. Within that grammar, each city writes its own paragraphs.

The SAMIs provide a way to read those paragraphs without being dazzled or misled by word count. They are empirical, they are comparable, and, because of that long memory, they represent the beginning of a diagnosis, not the end.

Where does this go next? The team hints at obvious extensions. Take the same method to other countries, new decades, more indicators, and see which pieces of the story travel.

Use the clustering to choose peers that make sense for benchmarking, even if they’re in a different time zone. And, cautiously, combine SAMIs with process data—on migration, industry churn, housing supply. Use that to pry apart which local levers move which deviations.

But even before that, there's a very practical mental shift available to anyone reading a city chart tomorrow. Ask what the size baseline would imply, and then determine who's truly above or below it. That’s the sound of the universal and the local, harmonizing—and sometimes clashing—inside the urban chorus.

More in Economics, Econometrics and Finance