Statistical Basis for Predicting Technological Progress
A government energy minister sits down in 2010 to decide whether to bet billions on solar panels. Her question is brutally practical: will panels get cheap enough, fast enough, to matter for the grid? Multiple theories existed to answer that question, but none had been rigorously tested against each other. That was the gap Nagy, Farmer, Bui, and Trancik set out to close. The competing frameworks go back decades. Theodore Wright proposed in 1936 that unit cost falls as a power law of cumulative production. This means that every time you double the total amount of a thing ever made, cost drops by a fixed percentage. Gordon Moore's observation in 1965 about integrated circuits was generalized into a broader claim that technologies improve exponentially with time, regardless of how much is actually being produced. Then there are alternatives, such as Goddard's critique of learning-curve interpretations, the Sinclair-Klepper-Cohen analysis of where cost reductions actually come from, and Nordhaus's cautionary work on the limits of learning-based models. Six hypotheses exist, all plausible, yet none have been put through the same rigorous test at the same time on the same data.
So, Nagy and colleagues built the test. They compiled what they call the most expansive dataset of its kind: sixty-two different technologies drawn from chemical, hardware, energy, and other categories, assembled into a Performance Curve Database. Each dataset spans between ten and thirty-nine years, sampled annually. The common metric across all sixty-two is inflation-adjusted cost per unit. This is a deliberately crude aggregate since the precise meaning of a "unit" can drift over time, which makes any consistent patterns all the more striking when they appear. The evaluation method is called hindcasting. In plain language, this means pretending you’re standing in some past year, fitting your model only to data available up to that point, projecting forward, then scoring your forecast against what actually happened. This is the cleanest way to test a forecasting law because you are grading predictions against real history rather than the data you already used to build the model. The authors ran this across all six hypotheses, generating a total of thirty-seven thousand seven hundred and forty-five hindcast error points. To rank the laws fairly, they built a mixed-effects statistical model that addresses two tricky features of this kind of data. First, heteroscedasticity means the variance in errors isn't constant; it grows with the error itself. Second, temporal correlation indicates that errors close together in time tend to be correlated.
The response variable they settled on, after a likelihood search across transformations, was the square root of the absolute logarithmic forecasting error. The core model states that this quantity equals a hypothesis-specific intercept plus a slope multiplied by the forecasting horizon — how many years out you’re trying to predict. Dataset-specific random adjustments account for the fact that some technologies are simply more predictable than others. The entire model was fit using maximum likelihood, reducing hundreds of technology-specific parameters down to just sixteen free parameters total. Now for the results. Wright’s law wins. Across the full dataset, Wright’s cumulative production formulation produces the best hindcast forecasts. However, Moore’s law is close behind — close enough that the difference in slope between the two sits right at the boundary of statistical significance. The other laws fall further back. Goddard does noticeably worse at short horizons, indicating that cumulative production contains forecasting information that annual production alone doesn’t capture. The Sinclair-Klepper-Cohen variant and Moore do worse at longer horizons, partly because their extra parameters invite overfitting. Nordhaus overfitted so severely that it was dropped from the main comparison entirely.
The universal finding across all surviving hypotheses is this: the square root of the logarithmic error grows linearly with forecast horizon at about two and a half percent per year. The fitted intercepts cluster tightly between zero point sixteen and zero point nineteen, while the slopes range between zero point zero twenty-four and zero point zero twenty-eight. What this means in practice is that forecasting uncertainty increases with lead time in a predictable, quantifiable way. You cannot predict the future perfectly. But you can predict it, and you can know how wrong you’re likely to be. That second part is what was previously missing. The paper makes this concrete with photovoltaics. Using the Moore functional form and their error model, the expected residential-scale solar cost in 2020 is six cents per kilowatt-hour, with a projected range of three to twelve cents. By 2030, the expected cost is two cents per kilowatt-hour, with a range of zero point four to eleven cents. For reference, coal at the plant — the cheapest conventional alternative at that time — was running about five cents per kilowatt-hour and was not expected to fall. So the forecast isn’t just an academic exercise; it’s a concrete statement about when solar crosses coal.
Now here's where the paper takes an unexpected turn. After establishing that Wright beats Moore — barely — the authors ask why the race is so close. The answer traces back to a conjecture made by Sahal decades earlier, which had never been empirically confirmed on a large dataset. Sahal pointed out a mathematical relationship: if production grows exponentially over time, then Wright’s law and Moore’s law become equivalent. You cannot distinguish them from cost and production data alone. The reason lies in their structures. Wright states that cost falls as a power of cumulative production; Moore states that cost falls exponentially with time. If cumulative production itself grows exponentially — which it will if annual production grows exponentially — then both descriptions are tracking the same underlying curve, just from different angles. The three key parameters lock together: the Wright exponent approximately equals the Moore decay rate divided by the production growth rate. Nagy and colleagues test this directly. They fit exponential models for production across all sixty-two datasets and then compare measured Wright exponents to the values derived by taking each technology’s Moore rate and dividing by its production growth rate. When plotted against each other, the points cluster along the identity line. Sahal’s conjecture is correct.
The empirical consequence is striking. For most of these sixty-two technologies, production does in fact grow roughly exponentially. That means every doubling of cumulative production tends to happen in roughly the same amount of time — and that regularity is sufficient to make Wright and Moore behave almost identically on the data. This is exactly why the horse race ends in a photo finish. The paper is careful about what this does and doesn’t mean. Wright and Moore are not the same theory; they carry different causal stories about what drives cost reduction. One points to learning through accumulated production experience, while the other points to some kind of autonomous progress through time. You can construct scenarios that pull them apart. Nagy and colleagues offer one: if the rate of photovoltaic production growth suddenly accelerated, Wright’s law predicts costs would fall faster because you’d be racing through cumulative production doublings more quickly. Moore’s law predicts no change because only time matters. Distinguishing the theories requires cases where production deviates from steady exponential growth. The current database does not provide enough of those. Thus, the theories remain separable in principle and nearly indistinguishable in practice.
What this study ultimately delivers is not a single winning equation. It’s a validated framework for forecasting with known uncertainty. Because the error model pools information across sixty-two technologies, it can tell you not just a central trajectory but how that trajectory typically diverges from reality as the horizon lengthens. The dominant advance lies in pairing any reasonable law — whether Wright or Moore, as they are nearly equivalent — with a statistically grounded error structure. For climate policy, the implications are direct and contained in the evidence. Cost trajectories for technologies like solar photovoltaics can be forecasted with quantified uncertainty bounds. A minister betting on solar in 2010 doesn’t need a perfect prediction. She needs a central estimate and honest error bars. Nagy, Farmer, Bui, and Trancik showed that both are achievable and that technological progress, which can feel like it arrives by surprise, is actually predictably imperfect. That predictability is itself a resource for planning. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Consumer response to corporate irresponsible behavior: Moral emotions and virtues
- Simpson's paradox in psychological science: a practical guide
- What Is the Subjective Cost of Cognitive Effort? Load, Trait, and Aging Effects Revealed by Economic Preference
- ‘Predatory’ open access: a longitudinal study of article volumes and market characteristics
- Do Altmetrics Work? Twitter and Ten Other Social Web Services
- Do Pressures to Publish Increase Scientists' Bias? An Empirical Support from US States Data