Quantitative ESG disclosure and divergence of ESG ratings
Two reputable agencies look at the same company. Both have access to its Environmental, Social, and Governance, or ESG reports, which are intended to inform investors how responsible this firm is. They produce scores so far apart that they might as well be rating different companies. That gap is real, it's documented, and it has a practical consequence: if you're building a sustainable portfolio, which rating do you trust? Min Liu's 2022 study set out to answer a specific version of that question — does providing agencies with more hard numbers to work with actually close that gap? The short answer is no. The reason why reveals something important about how information works in markets. Liu's study draws on Chinese A-share nonfinancial listed firms from 2014 to 2020, collecting ESG scores from six major Chinese rating providers: SynTao Green Finance, Sino-Securities Index, CASVI, WIND ESG, FTSE Russell, and Rankins. That's up to six independent ratings for the same firm in the same year, rescaled to a common zero to ten range. To measure disagreement, Liu calculates the standard deviation of a firm's ratings across those providers. The wider the spread, the more agencies disagree. The full dataset covers four thousand nine hundred sixty-six firm-year observations. The mean divergence in that sample is about 2.05 on a ten-point scale, with a standard deviation of 0.86. That's not a rounding error. That's meaningful disagreement about the same underlying reality.
And it's not just noise. Prior work by Gibson and colleagues and by Avramov and colleagues links higher ESG rating disagreement to a higher cost of equity — investors demand a premium for the extra uncertainty. Liu's own data confirm the downstream consequences: firms with wider rating divergence today tend to have lower average ESG ratings in the next period. The coefficient in her preferred specification is negative 0.078, significant at the one-percent level. Disagreement, in other words, is a signal — not just an artifact of imperfect measurement. So what causes it? The intuitive answer is a lack of data. If agencies are working with different information, of course they'll disagree. If you give them more numbers — more quantitative ESG metrics — they should converge. This is roughly how financial disclosure works: standardized earnings reports reduce analyst dispersion. Liu challenges that intuition by drawing on the sociology of valuation. Convergence requires two things to be true at once. First, raters need to share a basic theory of what ESG covers — what counts, what matters. Second, their measures need to be commensurable — comparable in the sense that the same metric means the same thing across providers. The problem is that ESG disclosure in most markets is neither standardized nor consistently defined. Kotsantonis and Serafeim showed that fifty Fortune firms used more than twenty distinct metrics just for employee health and safety.
When firms disclose a lot of numbers and none of those numbers map onto a shared framework, agencies fill in the gaps with their own judgment. More numbers mean more surface area for interpretation. And more interpretation means more disagreement. This logic generates two competing hypotheses. Either more quantitative ESG disclosure reduces disagreement — because numbers constrain discretion — or it increases disagreement, because non-standardized numbers amplify interpretive variation. Liu calls these H1b and H1a respectively. A third hypothesis, H2, adds the crucial qualifier: standardized quantitative disclosure — the kind governed by a specific reporting guide — should reduce disagreement even if generic numerical disclosure does not. The data back H1a. In the main regressions, the coefficient on ESG_Qmetrics — that's the natural log of the number of quantified ESG indicators disclosed — is positive and statistically significant across every specification. In the most controlled model, with firm, industry, year, and agency fixed effects, the coefficient is 0.061 with a t-statistic of 3.10. Given that the sample standard deviation of divergence is 0.86, the paper describes this as economically substantial. More quantitative disclosure, not less, is associated with wider disagreement.
The pillar-level analysis sharpens this. When Liu separates environmental, social, and governance disclosure into distinct variables, the environmental and social pillars drive the effect. Their coefficients are 0.058 and 0.057 respectively, both statistically significant. The governance pillar's coefficient of 0.048 is not significant. This makes intuitive sense: environmental and social metrics — carbon emissions, water usage, employee turnover, supply chain conditions — are highly context-dependent and industry-specific. Different agencies apply different benchmarks. Governance metrics, by contrast, tend to be more legible and more comparable across firms. Robustness checks using alternative measures of divergence replicate the finding. Across those alternative specifications, the coefficients on ESG_Qmetrics range from 0.009 to 0.075, all positive, all significant. But here's where the paper gets structurally interesting. If non-standardized numbers make things worse, what happens when you mandate standardized numbers? Liu tests this using the Hong Kong Exchanges and Clearing's Environmental, Social, and Governance Reporting Guide, which took effect in July 2020 and required cross-listed firms to report ESG data in specific, comparable formats.
This is a quasi-natural experiment — firms cross-listed on the Hong Kong Exchange are the treatment group; firms listed only on Shanghai and Shenzhen are the control. A difference-in-differences design compares how rating divergence evolves across these two groups before and after the rule change. The result flips. Under the Hong Kong Exchange framework — where disclosure is standardized and directly comparable — additional numerical ESG information reduces agencies' rating disagreement rather than increasing it. The interaction term in the difference-in-differences model is negative. The underlying mechanism is commensurability: when everyone is filling in the same template with the same definitions, there's less room for idiosyncratic interpretation. The data become a shared language rather than raw material for competing judgments. This is the paper's central insight, and it deserves a beat. The problem isn't that firms are hiding data. The problem is that the data they're releasing don't speak a common language. More disclosure without standardization is like asking six translators to interpret a speech, each using a different dictionary. You get six different readings — and the more complex the speech, the wider the gap.
Two additional findings extend this picture. First, the disagreement effect isn't uniform across firms. Liu splits the sample by business diversification — measured using a Herfindahl index of operating income — and finds the coefficient on ESG_Qmetrics is 0.144 for single-business firms but only 0.051 for diversified firms. That difference is significant at the five-percent level. Single-business firms tend to disclose highly industry-specific metrics; those metrics are harder for agencies to evaluate without deep sector expertise, so raters rely more heavily on their own frameworks and diverge more. Diversified firms present a more averaged, less specialized ESG profile, which is easier to assess consistently. The effect also concentrates among poor ESG performers. Splitting by average ESG rating, the coefficient on ESG_Qmetrics is 0.087 for below-average performers and only 0.025 for above-average ones. Liu's explanation is that agencies generally agree when ESG performance is clearly good — positive events are easy to credit. But when performance is poor, agencies differ in how severely to penalize bad outcomes. One agency might assign its lowest grade for a major violation; another might dock fewer points. That severity judgment varies across providers, and more disclosure gives each agency more material to apply its own severity calculus.
Second, the consequence of disagreement matters for how investors should interpret it. As noted earlier, wider divergence in the current period predicts lower average ESG ratings in the next. The coefficient is negative 0.078 in the primary specification and negative 0.041 when firm fixed effects absorb stable firm characteristics — both negative, both significant. Rating disagreement isn't symmetric uncertainty about an unknowable truth. It tracks underlying ESG weakness. When agencies can't agree on a firm's score, that uncertainty tends to resolve downward. The practical implications follow directly from the findings. For investors, the quantity of ESG numbers in a company's report is not a quality signal. A thick sustainability disclosure document full of metrics does not mean those metrics are interpretable or comparable. What matters is whether those numbers were produced within a standardized framework that constrains how agencies can use them. For regulators, the Hong Kong Exchange experiment points toward a specific intervention: mandate comparable formats and defined key performance indicators, not just increased disclosure volume. The goal isn't more data — it's commensurable data.
Liu's study has acknowledged limits: it focuses on the Chinese market, draws on six agencies, and measures disclosure primarily from ESG and corporate social responsibility reports. How these dynamics play out in other regulatory contexts, with different agency ecosystems, remains an open question. But the core finding is portable: the credibility problem facing ESG ratings isn't fundamentally about data availability. It's a measurement design problem. And more of the wrong kind of data makes it worse. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Inferring Labor Income Risk and Partial Insurance From Economic Choices
- The stock market's reaction to quality certification: Empirical evidence from Spain
- Responding to Globalization: Impacts of Certification on Colombian Small-Scale Coffee Growers
- The Molecular Genetic Architecture of Self-Employment
- Initial Coin Offerings
- Trajectories of brand hate