Technical noteInherent benchmark or not? Comparing Nash–Sutcliffe and Kling–Gupta efficiency scores

Wouter Knoben, Jim Freer, Ross WoodsView original
OverviewBalancedmaya voice
Zero and negative 0.41. In hydrology, both of those numbers are supposed to mean the same thing: your model is no better than guessing. The problem is they do not. They come from two different metrics, and conflating them has been quietly distorting how the field judges its own models. This is the story of how that confusion happened, and the surprisingly simple arithmetic that resolves it. Hydrologists need a single number to summarize model performance. The traditional answer has been the Nash–Sutcliffe efficiency, introduced by Nash and Sutcliffe in 1970. Nash–Sutcliffe efficiency, or NSE, is built on a clean idea: normalize your model's error against the natural variability of the river itself. Algebraically, NSE is one minus the ratio of the model's squared errors to the variance of the observed flows. Say that out loud: one minus the sum of squared differences between simulated and observed flow, divided by the sum of squared differences between observed flow and the mean observed flow. The benchmark meaning falls right out of that form. If your model's errors are exactly as large as the river's own variability around its mean, the ratio is one and NSE is zero. NSE equals zero means you are doing no better than just predicting the mean flow every single time step. NSE below zero means you are actually worse than that. It's intuitive. Practitioners have relied on NSE equals zero as their pass-fail line for decades. Then Gupta and colleagues proposed the Kling–Gupta efficiency, or KGE, in 2009, motivated by a real limitation in NSE. NSE is dominated by large flow events — a few big floods can swamp the score — and it does not tell you what kind of error you are making. KGE fixes that by decomposing performance into three separate components: the correlation between observed and simulated flows, a variability ratio that compares the standard deviation of simulations to observations, and a bias ratio that compares the means. Perfect agreement gives KGE equals one on all three components simultaneously. The formula is then one minus the Euclidean distance of those three components from their ideal values — one minus the square root of the sum of three squared differences: correlation minus one squared, plus variability ratio minus one squared, plus bias ratio minus one squared. That decomposition is genuinely useful. When a model fails, you can see whether it is a timing problem, a variability problem, or a mean bias problem. For those reasons, KGE has been increasingly adopted. Here is where the trouble starts. Because NSE equals zero is such a familiar benchmark, many practitioners carried that intuition directly into KGE. Negative KGE is bad, positive KGE is good, and KGE equals zero is the line. Knoben, Freer, and Woods showed in 2019 that this transfer of intuition is simply wrong — and the demonstration requires nothing more than arithmetic. Ask what happens when you plug the mean-flow predictor into the KGE formula. The mean-flow predictor is a flat line: every time step gets the same value, the observed mean. Walk through each component. The bias ratio equals one — the mean of your simulation equals the mean of observations, exactly. But the variability ratio collapses to zero because a constant has no standard deviation. And the correlation between a flat line and a variable hydrograph is formally undefined, but Knoben and colleagues argue it makes intuitive sense to assign the correlation a value of zero since there is no co-variation between the two. Now substitute those three values into the KGE formula: one minus the square root of zero minus one squared, plus zero minus one squared, plus one minus one squared. That simplifies to one minus the square root of one plus one plus zero. One minus the square root of two. The square root of two is approximately 1.41, so the mean-flow benchmark scores KGE equals one minus 1.41, which is approximately negative 0.41. Not zero. Negative 0.41. That single result reframes everything. It means a model with a KGE of negative 0.20 is actually beating the mean-flow benchmark — even though its score is negative. The mean-flow line sits at negative 0.41, not at zero. Any model that scores above negative 0.41 is an improvement over trivially predicting the mean. The field has been rejecting models that were, by the metric's own geometry, genuinely informative. Knoben and colleagues documented this directly. They identified a string of published studies — including work by Rogelis and colleagues, Schönfelder and colleagues, Andersson and colleagues, Fowler and colleagues, Siqueira and colleagues, Sutanudjaja and colleagues, and Towner and colleagues — that treat negative KGE as unambiguously bad model performance. In each case, the implicit assumption is that KGE and NSE share a zero benchmark. They do not. KGE has no inherent benchmark at zero. The number zero on the KGE scale has no special meaning. It is just a number. The second major finding from Knoben and colleagues is that NSE and KGE are not interchangeable even when you understand the benchmark correction. Their relationship is non-unique. It depends on the coefficient of variation of the observed streamflow — how variable the river is relative to its mean. A flashy, spiky river with high coefficient of variation will map errors to KGE and NSE in a completely different way than a smooth, stable river with a low coefficient of variation. The paper illustrates this with synthetic experiments across catchments spanning coefficients of variation of 0.28, 2.06, and 5.00. The results are striking. A simulation with a mean bias of roughly positive 39 percent scores KGE equals 0.61 and would be accepted by a reasonable KGE threshold, but scores poorly enough on NSE to be rejected. Another case with bias near 97 percent scores NSE equals 0.96 — easily accepted on NSE grounds — but KGE equals only 0.03, which would be rejected under a KGE threshold. These are not edge cases manufactured to make a point. They reflect the structural difference between the two metrics: NSE is dominated by high-flow performance, while KGE spreads weight more evenly across correlation, variability, and bias. The same model, on the same river, can pass one test and fail the other — and which test it passes or fails tells you something different about the error. They are measuring different things. This means a modeller who switches from NSE to KGE and simply keeps using zero as the acceptability cutoff is making two errors simultaneously. First, they are applying the wrong benchmark — the mean flow is at negative 0.41, not zero. Second, they are treating a score that means one thing on NSE as if it means the same thing on KGE, when the relationship between the two scores shifts depending on how variable the river is. Knoben and colleagues recommend two practical fixes. The first is to use explicit benchmarks rather than treating zero as an inherited threshold. The mean flow is one defensible choice, but it is not the only one — depending on the application, a climatological predictor or a simpler model might be a more appropriate yardstick. Whatever the benchmark is, state it explicitly, compute the model's skill relative to it, and report that. A skill score normalized against the benchmark — positive means better than the benchmark, negative means worse — is far more interpretable than a raw KGE value against an implicit zero. The second fix is to stop treating KGE as a single verdict and start reading its components. When a model fails, knowing whether it is the correlation, the variability ratio, or the bias ratio that is driving the poor score tells you something actionable. A composite number discards that information. The broader argument the paper makes is worth sitting with. The field has drifted toward using aggregated efficiency metrics as universal pass-fail tests, detached from the specific purpose of the model being evaluated. Knoben and colleagues argue for a framework built around purpose-dependent metrics and explicit benchmarks — where the choice of what to measure and what to compare against is determined by the question being asked, not inherited from a different metric's historical convention. The mean-flow benchmark made sense as the NSE reference point because NSE's algebra builds it in. Carrying that convention into KGE, where it does not belong, is the kind of error that propagates silently through a literature for years. The fix, as Knoben and colleagues show, is not complicated. It's negative 0.41. Know what that number means. Build your comparisons from there. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Zero and negative 0.41. In hydrology, both of those numbers are supposed to mean the same thing: your model is no better than guessing. The problem is they do not. They come from two different metrics, and conflating them has been quietly distorting how the field judges its own models. This is the story of how that confusion happened, and the surprisingly simple arithmetic that resolves it. Hydrologists need a single number to summarize model performance. The traditional answer has been the Nash–Sutcliffe efficiency, introduced by Nash and Sutcliffe in 1970. Nash–Sutcliffe efficiency, or NSE, is built on a clean idea: normalize your model's error against the natural variability of the river itself. Algebraically, NSE is one minus the ratio of the model's squared errors to the variance of the observed flows. Say that out loud: one minus the sum of squared differences between simulated and observed flow, divided by the sum of squared differences between observed flow and the mean observed flow. The benchmark meaning falls right out of that form. If your model's errors are exactly as large as the river's own variability around its mean, the ratio is one and NSE is zero. NSE equals zero means you are doing no better than just predicting the mean flow every single time step. NSE below zero means you are actually worse than that. It's intuitive. Practitioners have relied on NSE equals zero as their pass-fail line for decades.

Then Gupta and colleagues proposed the Kling–Gupta efficiency, or KGE, in 2009, motivated by a real limitation in NSE. NSE is dominated by large flow events — a few big floods can swamp the score — and it does not tell you what kind of error you are making. KGE fixes that by decomposing performance into three separate components: the correlation between observed and simulated flows, a variability ratio that compares the standard deviation of simulations to observations, and a bias ratio that compares the means. Perfect agreement gives KGE equals one on all three components simultaneously. The formula is then one minus the Euclidean distance of those three components from their ideal values — one minus the square root of the sum of three squared differences: correlation minus one squared, plus variability ratio minus one squared, plus bias ratio minus one squared. That decomposition is genuinely useful. When a model fails, you can see whether it is a timing problem, a variability problem, or a mean bias problem. For those reasons, KGE has been increasingly adopted. Here is where the trouble starts. Because NSE equals zero is such a familiar benchmark, many practitioners carried that intuition directly into KGE. Negative KGE is bad, positive KGE is good, and KGE equals zero is the line. Knoben, Freer, and Woods showed in 2019 that this transfer of intuition is simply wrong — and the demonstration requires nothing more than arithmetic.

Ask what happens when you plug the mean-flow predictor into the KGE formula. The mean-flow predictor is a flat line: every time step gets the same value, the observed mean. Walk through each component. The bias ratio equals one — the mean of your simulation equals the mean of observations, exactly. But the variability ratio collapses to zero because a constant has no standard deviation. And the correlation between a flat line and a variable hydrograph is formally undefined, but Knoben and colleagues argue it makes intuitive sense to assign the correlation a value of zero since there is no co-variation between the two. Now substitute those three values into the KGE formula: one minus the square root of zero minus one squared, plus zero minus one squared, plus one minus one squared. That simplifies to one minus the square root of one plus one plus zero. One minus the square root of two. The square root of two is approximately 1.41, so the mean-flow benchmark scores KGE equals one minus 1.41, which is approximately negative 0.41. Not zero. Negative 0.41. That single result reframes everything. It means a model with a KGE of negative 0.20 is actually beating the mean-flow benchmark — even though its score is negative. The mean-flow line sits at negative 0.41, not at zero. Any model that scores above negative 0.41 is an improvement over trivially predicting the mean. The field has been rejecting models that were, by the metric's own geometry, genuinely informative.

Knoben and colleagues documented this directly. They identified a string of published studies — including work by Rogelis and colleagues, Schönfelder and colleagues, Andersson and colleagues, Fowler and colleagues, Siqueira and colleagues, Sutanudjaja and colleagues, and Towner and colleagues — that treat negative KGE as unambiguously bad model performance. In each case, the implicit assumption is that KGE and NSE share a zero benchmark. They do not. KGE has no inherent benchmark at zero. The number zero on the KGE scale has no special meaning. It is just a number. The second major finding from Knoben and colleagues is that NSE and KGE are not interchangeable even when you understand the benchmark correction. Their relationship is non-unique. It depends on the coefficient of variation of the observed streamflow — how variable the river is relative to its mean. A flashy, spiky river with high coefficient of variation will map errors to KGE and NSE in a completely different way than a smooth, stable river with a low coefficient of variation. The paper illustrates this with synthetic experiments across catchments spanning coefficients of variation of 0.28, 2.06, and 5.00. The results are striking. A simulation with a mean bias of roughly positive 39 percent scores KGE equals 0.61 and would be accepted by a reasonable KGE threshold, but scores poorly enough on NSE to be rejected.

Another case with bias near 97 percent scores NSE equals 0.96 — easily accepted on NSE grounds — but KGE equals only 0.03, which would be rejected under a KGE threshold. These are not edge cases manufactured to make a point. They reflect the structural difference between the two metrics: NSE is dominated by high-flow performance, while KGE spreads weight more evenly across correlation, variability, and bias. The same model, on the same river, can pass one test and fail the other — and which test it passes or fails tells you something different about the error. They are measuring different things. This means a modeller who switches from NSE to KGE and simply keeps using zero as the acceptability cutoff is making two errors simultaneously. First, they are applying the wrong benchmark — the mean flow is at negative 0.41, not zero. Second, they are treating a score that means one thing on NSE as if it means the same thing on KGE, when the relationship between the two scores shifts depending on how variable the river is. Knoben and colleagues recommend two practical fixes. The first is to use explicit benchmarks rather than treating zero as an inherited threshold. The mean flow is one defensible choice, but it is not the only one — depending on the application, a climatological predictor or a simpler model might be a more appropriate yardstick.

Whatever the benchmark is, state it explicitly, compute the model's skill relative to it, and report that. A skill score normalized against the benchmark — positive means better than the benchmark, negative means worse — is far more interpretable than a raw KGE value against an implicit zero. The second fix is to stop treating KGE as a single verdict and start reading its components. When a model fails, knowing whether it is the correlation, the variability ratio, or the bias ratio that is driving the poor score tells you something actionable. A composite number discards that information. The broader argument the paper makes is worth sitting with. The field has drifted toward using aggregated efficiency metrics as universal pass-fail tests, detached from the specific purpose of the model being evaluated. Knoben and colleagues argue for a framework built around purpose-dependent metrics and explicit benchmarks — where the choice of what to measure and what to compare against is determined by the question being asked, not inherited from a different metric's historical convention. The mean-flow benchmark made sense as the NSE reference point because NSE's algebra builds it in. Carrying that convention into KGE, where it does not belong, is the kind of error that propagates silently through a literature for years. The fix, as Knoben and colleagues show, is not complicated. It's negative 0.41. Know what that number means. Build your comparisons from there.

This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Environmental Science