Assessing corporate sustainability with large language models: evidence from Europe
1 More Paper · Full Reading

About this paper
A full audio edition of this paper.
Authors: K. Forster, L. Keil, V. Wagner, M.A. Müller, T. Sellhorn, S. Feuerriegel
Publication date: 2026
Read the paper: https://doi.org/10.1038/s41467-026-75160-z
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “Assessing corporate sustainability with large language models: evidence from Europe,” by K. Forster and colleagues. Published in 2026.
1,2, Lucas Keil 3,4, Victor Wagner 1,4,5, Maximilian A. Müller3,4, Kerstin Forster Received: 10 September 2025 Thorsten Sellhorn1,4 & Stefan Feuerriegel 1,2 Accepted: 19 June 2026 Companies play a crucial role in achieving global sustainability goals, yet evi-Check for updates dence on their progress across environmental, social, and governance (ESG) dimensions remains limited. We develop a machine learning framework to systematically extract ESG indicators from corporate reports. Applying this approach to annual and sustainability reports of 600 large European firms (2014–2023), we construct a dataset of 2.9 million ESG observations across environmental, social, and governance topics. We assess ESG transparency based on disclosures aligned with the European Sustainability Reporting Standards (ESRS) and evaluate ESG performance using extracted numerical indicators.
Results reveal a pronounced transparency gap: firms in the top ESG rating decile disclose 22% more indicators than those in the bottom decile, although this gap narrows over time. Performance trends are uneven: while most social indicators remain largely stagnant, except for gains in gender equality, environmental indicators show some improvement. Reported scope 3 emissions increase sharply, largely reflecting improved disclosure. Our open-source framework enables systematic tracking of corporate ESG efforts.
Private-sector businesses importantly shape the world’s progress toward environmental, social, and governance (ESG) goals1–5. A total of 169 companies are responsible for 80 percent of industrial GHG emissions, of which 56 are European6. Yet, despite widespread cor-porate pledges to reduce environmental footprints7,8, many compa-nies still lag behind their environmental goals2,3. Similarly, corporate efforts to improve workplace diversity also continue to fall short9,10. Clearly, corporate ESG efforts need close monitoring, if performance improvements are to be achieved.
However, comprehensive evidence about corporate progress along ESG dimensions is limited, primarily for five reasons. First, many studies focus narrowly on a few ESG dimensions such as carbon emissions, energy use, or water withdrawal11–16, while neglecting other environmental (e.g., biodiversity) as well as social and governance indicators (e.g., employee turnover). Second, others are confined to specific sectors4,17, countries18,19, or time periods20–22, reducing the ability to track broader cross-sectoral trends or identify regional dis-parities. Third, several studies analyze the narrative in corporate sustainability reports qualitatively23–26, lacking quantifiable data nee-ded for tracking progress in ESG performance.
Fourth, while there is a large body of research in finance, accounting, and legal studies using natural language processing to analyze regulatory filings (see over-views in refs. 27 and28), primarily for the purpose of text classification (e.g., sentiment analysis, readability analysis), monitoring ESG performance is a different task and requires the extraction of structured quantitative information. Fifth, the most comprehensive ESG data is compiled by commercial providers (see Supplementary Informa-tion S2); however, because no unified reporting standard exists, these datasets are costly, focused on investors’ information needs, and often report inconsistent measures—unless these measures come directly from the companies themselves29.
Given these limitations, a compre-hensive, quantitative assessment of corporate ESG transparency and performance across granular ESG dimensions and over a multi-year trajectory is missing. Here, we aim to fill this gap.
Targeted transparency—i.e., disclosure requirements aimed at empowering civil society to hold companies accountable30—allows for tracking progress along ESG dimensions31 and can even act as a catalyst for change32,33. Policy-makers rely on ESG-related transparency to inform regulatory frameworks and monitor compliance with sustain-ability standards. For example, the EU’s recent Corporate Sustain-(CSRD)34 ability Reporting Directive mandates the most comprehensive sustainability reporting requirements globally. It introduced a set of European Sustainability Reporting Standards (ESRS)35 to unify ESG reporting for large, listed companies (see Fig. 1a and Supplementary Information S1).
Financial stakeholders, including banks, insurers, and asset managers, increasingly seek to align their portfolios with sustainability goals36 and thus depend on reliable ESG information to assess and manage risks and returns. Societal actors, such as consumers, non-governmental organizations (NGOs), and employees, rely on such information to hold corporations accountable for their ESG performance37. While transparency alone does not guarantee improved ESG outcomes, comprehensive and gap-free ESG reporting is often a necessary first step to help align corporate prac-tices and global ESG goals13.
Here, we develop a machine learning (ML) framework to construct a time series of ESG indicators from corporate reports based on retrieval-augmented generation (RAG) (Fig. 1). We then track the values of these ESG indicators along environmental (e.g., scope 1, 2, and 3 GHG emissions, water consumption, waste), social (e.g., employee turnover, women in top management, gender pay gap), and governance (e.g., lobbying expenses) topics as specified by ESRS. Our analysis covers the 600 largest listed corporations in Europe, which represent nearly 90% of the investable equity market in Europe based on market capitalization, over the 2014–2023 period. Overall, we prompt our ML framework for Ntrans = 2, 880, 249 ESG indicators, of which Nperf = 847, 835 return a numerical value because the corre-sponding ESG indicator was disclosed by the company and can be extracted from the corporate report.
Notably, while we only extract values for 29% of the indicators, this does not imply that 71% are missing; companies might not disclose data on certain indicators because they refer to a topic that is not deemed relevant (i.e., material) to the sector or entity.
We use this dataset for two key analyses: First, we assess ESG-related transparency as the availability of ESG disclosures defined by ESRS. Second, we analyze ESG performance over time and across industries based on the values of these indicators. To demonstrate the scientific value of our generated dataset for downstream analysis, we report descriptive evidence but refrain from causal claims about the drivers of transparency or performance. We make both our ML fra-mework and our dataset publicly available to allow policy-makers, investors, and other societal actors to systematically monitor and compare corporate ESG efforts—including among industry peers and against companies’ publicly stated targets. Doing so should help drive corporate progress toward global sustainability goals.
Results.
ML framework overview
Our end-to-end ML framework operates as follows (Fig. 1b; see the “Methods” and the Supplementary Information for full technical details). For each company–year (i.e., each company observed in a given year), we collect the corresponding annual and sustainability reports (as PDFs), extract and clean the text from the PDFs, split the text into overlapping chunks, and index the chunks in a vector data-base. Within each company–year, we then query the indexed text separately for each of the 501 ESG indicators, retrieve and re-rank candidate chunks, and use a state-of-the-art LLM to extract the requested numeric value (or abstain if absent). Finally, we standardize units and currencies and store the resulting output in a structured format (Supplementary Fig. S4). We apply this framework to STOXX Europe 600 index constituents (as of 2023) over the 2014–2023 per-iod.
Key methodological limitations are discussed in the Supplemen-tary Information S4.
We validated the outputs against (i) a proprietary benchmark dataset and (ii) expert human annotations on a subset. The agreement is strong against both comparisons (Supplementary Fig. S1–S2), which implies that the generated dataset is reliable for the descriptive ana-lyses below. The full validation design and error analysis are described in Methods section under Validation.
While our framework enables large-scale, automated ESG indi-cator extraction, the reliability of its outputs depends on the quality of extraction at each stage of the pipeline. We therefore quantify how observations progress through each stage of the ML pipeline and company–year report the results at both the level and the indicator–company–year level, which allows us to identify at which stage information enters—or drops out of—the process (Supplemen-tary Fig. S6). Building on this decomposition, an indicator may fail to be retrieved for two reasons: (i) it may not be disclosed (e.g., due to non-materiality or because the company reports a different indicator or uses a different reporting variant), or(ii) the extraction may fail(e.g., due to ambiguous reporting).
To assess whether the extraction may further vary across indicators, we additionally report per-indicator disclosure detection and standardization rates (Supplementary Table S6).
ESG-related transparency
We use our ML framework to track ESG indicators from both the annual and sustainability reports of the 600 major European companies listed in the STOXX Europe 600 (as of 2023) for the time period 2014 through 2023 (Fig. 1b). These companies cover nearly 90% of the investable equity market in Europe and span 16 European countries, including the UK (139 companies), France, and Ger-many (Fig. 1c). The full list of all companies is provided in Sup-plementary Table S1. From each report, we extract reported ESG indicators from among the 501 indicators defined by European Sus-tainability Reporting Standards (ESRS)35 (see Supplementary Informa-tion S1 for background). ESRS are structured along over-arching topics (e.g., climate change, own workforce, business conduct), each man-dating a granular set of quantitative ESG indicators that collectively constitute the 501 indicators.
We track these indicators to determine: whether the indicator is present or absent from the corporate report (transparency), and, if present, its numerical value (performance). Accordingly, this enables systematic benchmarking across firms, sectors, and time, which helps generate descriptive evidence, but without causal claims, to inform hypothesis development and demonstrate the scientific value of the dataset.
Our analysis shows an overall trend toward increased transpar-ency (Fig. 2a). The average number of disclosed ESG indicators increased from 117.8 (2014) to 179.7 (2023; + 52.5%), with more pro-nounced increases in specific topics such as climate change (+ 125.0%), water (+ 83.0%), and circular economy (+ 73.0%). Notably, the number of reported ESG indicators varies across industry sectors due to dif-ferences in business model, resource intensity, and regulatory expo-sure. But even beyond structural sector differences, our analysis points to a notable transparency gap (Fig. 2b, c). In 2023, companies with a top − 10% ESG rating (based on external data providers38) disclosed on average 186.5 ESG indicators, compared to an average of 174.7 indi-cators for companies that are lagging behind, corresponding to a 6.8% difference.
This gap was substantially wider in 2014, when companies in the upper decile reported on average 144.7 indicators versus just 103.8 for the bottom decile (difference: 39.4%). This narrowing of the transparency gap over the last decade coincides with a general increase in average ESG ratings over this period38, suggesting an overall strengthening and alignment of data availability about corporate sustainability.
We find substantial disparities in ESG-related transparency across topics and industries, as measured by a transparency score defined as the relative number of disclosed indicators out of all 501 indicators (Fig. 3a, b). This could be explained by differences in the materiality of topics, that is, certain sectors or entities may not be affected by a given topic and therefore do not disclose detailed information, or different scope and depth of disclosure frameworks prior to the ESRS. Overall, transparency scores are unequally distributed across topics (Fig. 3a). In 2023, topics with particularly high transparency scores were own workforce (54.7%), governance (48.9%), and circular economy (42.3%), whereas transparency scores were particularly low for the topics pol-lution (5.6%) and biodiversity (8.3%).
Overall, several topics show increasing transparency scores between 2014 and 2023, such as cir-cular economy (change in transparency score: + 17.8 percentage points [p.p.]), climate change (+ 15.5 p.p.), and own workforce (+ 13.1 p.p.). There is also some heterogeneous variation in overall as well as topic-specific transparency scores across industries (Fig. 3b). Variation in overall transparency may reflect sector-specific disclosure traditions and peer benchmarking effects, whereby firms adopt similar practices as their industry peers to meet investor and stakeholder expectations. Topic-specific variation might reflect their relative importance across industries. For example, whereas topics like own workforce and gov-ernance consistently score highest, arguably due to being universally material and often regulated, the relative importance of other topics varies.
One striking example is the financial industry, which displays comparatively lower transparency across environmental topics in contrast with social and governance, consistent with its lower direct ecological footprint but greater exposure to organizational and ethical issues. Still, the relative ranking of topics shows broadly similar char-acteristics across industries, with one possible explanation for that being a preference of companies to report on commonly accepted, measurable metrics requested by data providers and rating agencies. Further context on ESG transparency trends and the role of rating agencies is provided in Supplementary Information S2.
To identify potential drivers behind ESG transparency, we analyze transparency scores across different company characteristics. For this analysis, we compare companies in the top 10% and bottom 10% by market capitalization and ESG rating. In general, larger companies are expected to be more transparent, as they have more resources and stronger incentives (e.g., public exposure) for ESG reporting. Compa-nies with higher ESG ratings are typically more advanced in their sus-tainability practices and should thus tend to disclose more. Overall, transparency is indeed lower among smaller companies and those with lower ESG ratings (Fig. 4a,b).
Specifically, companies in the top 10% by market capitalization and ESG rating have significantly higher transparency scores than the middle 80% by + 3.1 p.p. (p < 0.001) and + 3.7 p.p. (p < 0.001), respectively. In contrast, transparency scores are significantly lower for companies in the bottom 10% of market capitalization (− 3.9 p.p.; p < 0.001) and ESG rating (− 2.5 p.p.; p < 0.001) compared to the middle 80%.
Further, we analyze the ESG controversies score, a proprietary metric capturing exposure to negative ESG-related events and news coverage39 (Fig. 4c). Companies with higher (i.e., worse) ESG con-troversies scores are those whose ESG performance has been more contentious in the past. Here, we find that companies in the bottom 10% of the ESG controversies score (i.e., the least controversial) have significantly lower transparency scores than the middle 80% (− 2.4 p.p.; p < 0.001), while the difference for companies in the top 10% is small and not statistically significant (− 0.6 p.p.). This suggests that companies with low controversy exposure also tend to disclose less, possibly because they face less external pressure to report compre-hensively on ESG matters.
Overall, these differences translate into substantial transparency gaps between companies, as confirmed in a regression analysis (see Table 1). For example, companies in the top − 10% of ESG ratings have, on average, 22% higher transparency scores than companies in the bottom − 10%, highlighting the sub-stantial transparency gap between the two groups.
rated and c lower-rated in terms of sustainability performance. Specifically, we compare against companies for which the ESG rating (based on lagged MSCI ESG ratings38) ranks in the top − 10% and bottom − 10%, respectively. Deciles are cal-culated on a yearly basis. This analysis shows how top- and bottom-rated compa-nies compare in ESG disclosure relative to the full sample.
In addition to assessing corporate ESG transparency, our frame-work enables large-scale, granular analysis of ESG performance—that is, the numerical values of disclosed ESG indicators over time. We focus our analysis on a selected set of indicators that are particularly relevant to current public debates and recent policy frameworks (e.g., the Corporate Sustainability Due Diligence Directive40, and the Zero Pollution Action Plan41). The complete set of indicators is included in our dataset (see Data availability statement) and is accessible through an interactive dashboard at (the linked source jUlW6L1X8u8).
Environmental performance
Overall, progress toward reducing corporate emissions is mixed (Fig. 5a–d). The graphs show the development of selected indicators over time and provide information on percentiles as well as reporting intensity (i.e., the number of firms reporting on a particular indicator in a given year). For example, median scope 1 emissions (i.e., direct GHG emissions from sources owned or controlled by the company such as production facilities or company vehicles) have declined by 66.8% since 2014 (Fig. 5a), suggesting a gradual shift toward lower direct emissions. Similarly, median scope 2 emissions (i.e., indirect emissions resulting from purchased electricity, steam, heating, or cooling) have declined by 76.4% (Fig. 5b).
In contrast, total scope 3 emissions (i.e., indirect emissions across the companies’ upstream and downstream value chain), remained largely stagnant between 2014 and 2020, before increasing by a factor of 5.6 through 2023 (Fig. 5c).
In the specific case of scope 3 emissions, however, we argue that this sharp increase is mainly attributable to higher transparency levels: companies have started to report emissions data for a wider range of previously untracked categories, which mechanically raises total scope 3 emissions. To substantiate this, we inspect the 15 scope 3 categories as defined in the Corporate Value Chain Accounting and Reporting Standard by the GHG Protocol. These exhibit no clear upward trends in individual categories (see Fig. 6) but a stark increase in the number of categories companies report emissions for (see Supplementary Fig. S21)42. Additionally, we confirm through regression analysis that the increase in total scope 3 emissions is partly driven by the increase in reported categories (see Supplementary Table S4).
As expected, we observe a sharp drop in scope 3 emissions from travel during the COVID-19 pandemic (Fig. 5d). Although travel-related emissions have rebounded by a factor of 2.8 since 2021, they remain below pre-pandemic levels in 2023, which suggests that many companies have either revised their travel policies or continue to embrace virtual meetings as part of a post-pandemic shift in workplace practices. This analysis highlights that the perceptions of corporate ESG performance are shaped by data availability, that is, ESG transparency. By enabling the tracking of ESG indicators, our ML approach can contribute to the ability of the public to monitor corporate ESG performance.
Beyond emissions, progress across other environmental indica-tors is uneven, with some showing considerable changes, while others have stagnated (Fig. 5e–i). For example, the increasing share of renewable energy sources (Fig. 5f), accompanied by a decline in fossil-based energy (not reported here), indicates growing alignment with global decarbonization efforts and a gradual transition toward more sustainable energy sources. However, companies in the highest decile of energy intensity—measured as energy consumed per EUR of rev-enue—continue to rely predominantly on fossil fuels. Furthermore, energy (Fig. 5e) and water consumption (Fig. 5g) have declined by 37.7% and 38.1%, respectively, over the observation period. Reductions in total waste generated (Fig. 5h), and non-recycled waste (Fig. 5i) have been modest.
To account for the underlying trends in economic activity, we analyze inflation-adjusted median revenues across the sample period from the Worldscope database43. Overall, we find only a modest increase of 16.7% over the full sample period (see Supplementary to the Sustainable Industry Classification System (SICS), developed by the Sus-tainability Accounting Standards Board (SASB) to classify companies into sectors with comparable exposure to sustainability-related risks and opportunities68. Industries in (b) are sorted by their overall transparency score (descending).
Table S3), suggesting that macroeconomic growth alone is unlikely to fully explain the trends in ESG performance. By analyzing intensity ratios, we find that the companies in our sample seem to have made real adjustments in some areas (e.g., reducing scope 1 and 2 emission intensities), while, in other areas, the intensity ratios remain stagnant (e.g., energy consumption) or increase (e.g., indirect emissions) (see Supplementary Fig. S13).
To further explore heterogeneity across companies, we stratify our analysis of corporate ESG performance by different company characteristics, namely, market capitalization, ESG ratings, and ESG controversies scores, comparing companies in the top and bottom 10% of each characteristic (see Supplementary Fig. S22–S24). Here, scope 1 emissions are substantially lower among smaller companies, compa-nies with lower ESG ratings, and those associated with fewer ESG controversies. We observe similar results for most of the other indicators.
The observed differences suggest that, in line with expectations, environmental performance is lower for larger companies and companies with more controversies around ESG practices. Notably, companies with higher ESG ratings alsoexhibit lower performance, and this pattern holds when analyzing intensities instead of absolute performance levels (as larger companies tend to have higher ESG ratings), highlighting the need to differentiate between sustainability trans-parency and performance (see Supplementary Fig. S14). There are two possible explanations for why companies with higher ratings exhibit lower environmental performance. First, ESG ratings as those by MSCI do not solely assess risks arising from negative impacts but also emphasize how these risks and opportunities are managed, which can favor larger firms with more advanced governance structures.
Second, the MSCI rating is determined relative to peers within the same industry, meaning that companies operating in resource-intensive sectors (e.g., extractives & minerals processing) may still receive comparatively high ratings despite sizable absolute impacts if they outperform their sector counterparts38.
Finally, to account for changes in the number and composition of reporting companies over time, we further examine whether trends differ between early and late adopters of ESG-related reporting. Spe-cifically, we conduct a separate analysis comparing firms that began ESG reporting early in the sample period to those that started later (see Supplementary Fig. S15). We find that overall trends in performance remain consistent across both groups, indicating that the main find-ings are not driven by sample composition. Restricting the sample to companies that report a given metric in all sample years yields closely similar trends (see Supplementary Fig. S18–S20).
Social performance
Performance across social indicators is mixed (Fig. 7). Notably, for instance, employee turnover has increased by 2.6 p.p. since 2014 (Fig. 7a), suggesting growing challenges in workforce retention. At the same time, the share of female employees in top management has increased steadily by 9.2 p.p. (Fig. 7e), which reflects ongoing efforts to promote gender equality in corporate leadership. The gender pay gap has narrowed by 4.7 p.p. since 2014 but has widened again by 0.5 p.p. since 2021 (Fig. 7f), suggesting potential setbacks in corporate efforts to promote equal pay. In 2023, the amount of fines, penalties, and compensation for damages as a result of incidents and complaints has controversies score such that companies with more controversies appear in the top 10%.
Shown are violin plots, which represent the kernel probability density of the data, together with boxplots displaying the median and interquartile range. Reported below each group is the number of firm-year observations (n). Reported above the brackets is the difference (diff) in percentage points (p.p.). Statistical comparisons are based on two-sided t-tests (with corresponding to the 0.1% significance level). Whiskers indicate the range of non-outlier values, extending to 1.5 times the interquartile range beyond the first and third quartiles.
decreased by 69.5% compared to 2014 levels. In contrast, the number of training hours per employee (Fig. 7c) remains stagnant over the observation period.
However, other indicators even exhibit decreases over time. The share of employees covered by collective bargaining agreements has decreased by 7.1 p.p., (Fig. 7b), indicating a weakening of employees’ bargaining power. Similarly, the annual remuneration ratio (i.e., the ratio of total annual compensation of the highest-paid individual to the median annual total remuneration for all employees, excluding the highest-paid individual) has increased by 1325.9% since 2014 (Fig. 7g), pointing to a widening gap between executive compensation relative to employee pay. The number of days lost to work-related injuries, ill health, and fatalities among employees has increased by 40.4% (Fig. 7d), mirroring an increase in the number of complaints filed by own workforce by 39.1% (Fig. 7h).
Consistent with our findings for environmental performance, companies that began ESG reporting earlier do not exhibit sub-stantially different trends in social performance (see Supplementary Fig. S16), and trends among constant reporters are closely similar (see Supplementary Fig. S19). As expected, social indicators measured in absolute terms exhibit lower levels for smaller companies while rela-tive metrics display comparable performance, with the exception of those pertaining to gender equality, where larger companies seem to perform better, possibly due to higher public exposure (see Supple-mentary Fig. S25). For ESG ratings and controversies, the results are mixed (see Supplementary Fig. S26–S27).
Governance performance
To understand governance practices (Fig. 8), we examine the share of independent board members, which remains consistently high, reaching 75.0% in 2023 (Fig. 8a), reflecting previous governance
Table 1 | Determinants of transparency
Notes: This table reports regression estimates where the dependent variable is the company-level transparency score, calculated as the number of reported ESG indicators divided by all indicators listed in ESRS. Each panel uses top − 10% and bottom − 10% dummies of the focal determinant as regressors. Deciles are calculated on a yearly basis. Panel A reports results using lagged market capitalization as the main independent variable. Panel B uses the lagged ESG rating from MSCI, and panel C uses the lagged ESG controversies score from Refinitiv (inverse-coded, i.e., higher scores indicate higher controversy exposure). Model includes only the focal determinant dummies and no additional controls. Models - additionally include the remaining two determinants as continuous controls.
The models vary by the set of included fixed effects (FE): Model and Model include no fixed effects; Model includes year fixed effects; Model includes sector fixed effects; Model includes year-by-sector fixed effects;and Model includes company fixed effects. Standard errors are clustered by company and reported in parentheses. Significance levels:, reforms aimed at strengthening board autonomy. In contrast, lobbying expenses have increased by 747.6% since 2019 (Fig. 8b). In the case of lobbying expenses, the earlier drop around 2014/15 could be attrib-uted to limited or strategic disclosure by early adopters of ESG-related reporting (see Supplementary Fig. S17); trends among constant reporters are broadly consistent (see Supplementary Fig. S20). How-ever, transparency on lobbying expenses remains generally low, resulting in fewer observations compared to other indicators shown here.
Lobbying expenses are also higher for larger companies and highlight variation between top and bottom-performing companies. For each panel, dot sizes indicate reporting intensity (i.e., the number of companies dis-closing values for each indicator in a given year), n denotes the total number of company–year observations for the indicator, and Δ denotes the percentage change in indicator value between 2014 and 2023. Note that logarithmic axes are used to better visualize ESG indicators that span several orders of magnitude. For indicators recorded in percent, we use linear scales. In d we also indicate the time period of the COVID-19 pandemic, referring to the World Health Organization’s definition of COVID-19 as a public health emergency of international concern91.
, correspond to 0.1%, 1%, and 5%, respectively. Reported significance levels are based on two-sided t-tests. No multiple-comparison adjustments were applied. Estimates are based on ordinary
least squares (OLS) regression. We find that companies in the top − 10% (bottom 10%) of market capitalization, ESG rating, and controversies score have significantly higher (lower) levels of
transparency than the average company. For example, according to Model, companies in the top − 10% versus bottom − 10% of ESG ratings have transparency scores that are ~22% higher,
calculated by comparing fitted transparency levels from the regression model.
companies with higher ratings, as well as more controversies around their ESG practices. For other governance indicators, the results are mixed, while larger companies tend to have a higher degree of inde-pendence within their board structure (see Supplementary Fig. S28–S30).
Discussion.
We extract and analyze granular, quantitative ESG indicators aligned with emerging sustainability reporting standards (that is, ESRS35).
reporting companies (solid line, 50th percentile), the interquartile range (shaded band, bounded by the 25th and 75th percentiles, inner dashed lines), and the 10th and 90th percentiles (outer dashed lines) to capture the distribution and highlight variation between top- and bottom-performing companies. For each panel, dot sizes indicate reporting intensity (i.e., the number of companies disclosing values for each indicator in a given year), and n denotes the total number of company–year observations for the indicator. Note that logarithmic axes are used to better visualize different orders of magnitude.
Although our findings point to an overall trend toward increased corporate ESG transparency, disclosure practices remain uneven, highlighting structural disparities shaped by topic- and industry-specific reporting priorities. Regarding ESG performance, our analysis shows heterogeneous trends across key indicators, including large increases in some measures (e.g., remuneration ratios). We note, however, that some of these observed changes coincide with expan-ded disclosure, highlighting how evolving transparency can affect the interpretation of reported ESG performance.
Our analysis is based on a comprehensive dataset on corporate ESG-related transparency and performance, which comprises 501 quantitative ESG indicators across nine ESRS topics. The main purpose of our analysis is to show the scientific value of the dataset for hypothesis generation and subsequent theory building, by providing descriptive insights that can subsequently be examined using causal inference methods. To construct our dataset, we leverage large lan-guage models as state-of-the-art techniques from machine learning. Whereas such machine learning models have been used in sustain-ability research17,19,21,22,44–50, for example, to analyze sustainability narratives23–26, we fill evidence gaps regarding transparency and per-formance by analyzing quantitative ESG indicators.
We also address key limitations of existing ESG analyses that are based on manual data collection and thus restricted to a few, selected indicators (e.g., carbon emissions, energy use, or water withdrawal12–16) or limited to certain countries18,19, sectors4,17, or short time periods20–22. In contrast, our framework enables tracking heterogeneity in trends across broad samples, e.g., to identify cross-industry differences or regional dis-parities. Furthermore, our open science approach democratizes data access.
Our analysis of corporate transparency reveals persistent gaps in ESG disclosures, which suggests that the transparency of current reporting practices is unevenly distributed across companies, sectors, and topics51. In interpreting these patterns, we accordingly refrain from making causal claims but highlight several cases where the findings are consistent with theoretical expectations.
The overall increase in indicator transparency is consistent with anticipated growth in regulatory constraints, investor demand, com-petitive pressure from peers, and expanding rating agency coverage12,52–55. While better extraction performance in more recent reports could, in principle, contribute to this trend, we observe similar patterns when analyzing a subset of high-quality reports, suggesting that the trend is not an artifact of our extraction pipeline.
The observation that the transparency gap between higher-rated and lower-rated companies becomes narrower over time is consistent with a convergence dynamic: as sustainability reporting transitioned from a voluntary differentiator to a common and eventually manda-tory practice, laggards increasingly adopted disclosure norms established by early leaders54,55. A regression analysis that interacts rating group dummies with a linear time trend confirms this pattern, indi-cating that lower-rated companies display relatively stronger increases in transparency over time, though the differential trends attenuate once company-level controls and fixed effects are included (Supple-mentary Table S2).
Notably, ESG ratings themselves increased over the sample period, and, because ratings may, at least partially, reflect transparency, including a one-year lag mitigates—but does not fully eliminate—the concern that the association between ratings and transparency is partly mechanical54,56,57.
Higher transparency scores among larger companies are con-sistent with greater resource availability, stronger regulatory expo-sure, more intense stakeholder scrutiny, and greater exposure to a diverse set of ESG topics51,58,59. An alternative explanation is that our extraction pipeline may perform better on reports from larger companies if these reports have higher document text quality; however, we find similar associations when analyzing a subset of reports with high document text quality (see Supplementary Fig. S7–S12). Notably, the size-transparency relationship becomes statistically insignificant once we control for ESG ratings and controversy exposure (see Table 1, Panel A), and remains insignificant across all subsequent specifications including sector and company fixed effects.
This suggests that the univariate association between size and transparency largely reflects other company characteristics correlated with size, such as rating coverage and controversy exposure, rather than size per se.
We further analyze the association between ESG controversy exposure, measured through the Refinitiv ESG controversies score60, and transparency (see Table 1, Panel C). We find that companies with lower controversy exposure have significantly lower transparency scores, while the difference for companies with higher controversy exposure is small and not statistically significant. This pattern admits several interpretations. Companies with less controversy exposure may face weaker external pressure to disclose comprehensively, or, conversely, companies that disclose more may expose themselves to greater scrutiny, creating a mechanical link between transparency and measured controversy exposure58,59,61.
Once we control for market capitalization and ESG ratings (see Table 1, Panel C, Model 2–6), the association becomes statistically insignificant, suggesting that the univariate association between controversy exposure and transparency largely captures other company characteristics corre-lated with controversy exposure, consistent with recent research on the company characteristics that shape ESG transparency, such as size, being third-party rated, analyst coverage, and institutional holdings12,51.
In addition, the observed disparities in transparency across sec-tors and topics highlight the need for regulatory frameworks that mandate the standardized disclosure of material ESG indicators to ensure consistency, comprehensiveness, and comparability in corpo-rate ESG reporting55,62. Cross-sectional variation in topic-level trans-parency—i.e., highest for own workforce and governance, lowest for pollution and biodiversity—likely reflects the maturity of established reporting frameworks and the degree of stakeholder attention that different topics receive63,64. Pollution and biodiversity disclosures may lag because these topics lack standardized metrics and are material to fewer companies65.
Alternatively, our pipeline may perform better on topics with well-codified terminology than on less standardized areas, or these topics may be genuinely non-material for many companies in our sample.
Our assessment of environmental performance reveals several trends in direct environmental indicators (e.g., scope 1 GHG emissions and renewable energy adoption). The decline in median scope 1 and scope 2 emissions is consistent with a reduction in direct emissions (e.g., through operational efficiencies) and a shift toward lower-carbon energy sources for direct operations, which is corroborated by the concurrent increase in renewable energy shares we document. How-ever, these reductions may also partly reflect divestiture of high-emitting assets, offshoring of production, or reclassification of emis-sions across scopes rather than genuine abatement. We examine inflation-adjusted median revenues across the sample period and find only a modest increase (16.7%), and do not document a downturn in overall economic activity that might independently explain the emis-sions decline.
We also present company-level intensity measures in support of our findings (see Supplementary Fig. S13).
Moreover, the concurrent increase in scope 3 GHG emissions shows the challenges companies face in managing emissions beyond their immediate operations—particularly across complex supply chains13. In light of widespread net-zero pledges7,8 and the legally binding goal set by the EU to reach climate neutrality by 205066, the current trajectory of emissions suggests that many companies remain considerably off-track. The new EU Corporate Sustainability Due Dili-gence Directive (CSDDD)40 may alter this trajectory by requiring that companies identify and address environmental risks along their value chains, thus strengthening accountability for scope 3 GHG emissions.
Notably, trends in reported scope 3 GHG emissions may reflect not only changes in underlying activities but also increased transpar-ency, including greater granularity, broader scope of coverage, and improved measurement practices. Furthermore, because scope 3 emissions encompass indirect emissions across value chains and often reporting companies (solid line, 50th percentile), the interquartile range (shaded band, bounded by the 25th and 75th percentiles, inner dashed lines), and the 10th and 90th percentiles (outer dashed lines) to capture the distribution and highlight variation between top and bottom-performing companies.
For each panel, dot sizes indicate reporting intensity (i.e., the number of companies disclosing values for each indicator in a given year), n denotes the total number of company–year observations for the indicator, and Δ denotes the percentage change in indicator value between 2014 and 2023. Note that logarithmic axes are used to better visualize different orders of magnitude. For indicators recorded in percent, we use linear scales.
involve overlapping reporting boundaries, we emphasize the need for policymakers to interpret scope 3 emission data cautiously to avoid double-counting and therefore misinformed conclusions or misguided policy responses. Additionally, our results show that major emitters continue to rely heavily on fossil energy sources, with limited adoption of renewables, implying that, despite the technical feasibility of dec-arbonizing industrial processes, current market incentives and reg-ulatory pressures may be insufficient to drive large-scale transition in carbon-intensive sectors. Further, several other environmental dimensions—such as water or waste management—exhibit less pro-gress in terms of performance, but often also regarding transparency (e.g., few companies disclose biodiversity-related indicators).
These findings highlight the importance of regulatory incentives to strengthen environmental performance holistically across entire value chains.
Similarly, uneven performance across social indicators reflects continued efforts to improve gender diversity in leadership, but lim-ited progress toward greater pay and income equality. The steady increase in female representation in top management is consistent with growing regulatory pressure (e.g., such as Directive (EU) 2022/ 2381 on gender balance on corporate boards) and evolving stake-holder expectations, though it may also partly reflect expanded defi-nitions of top management or selective reporting of favorable metrics. The gender pay gap narrowed overall but widened slightly after 2021, which could reflect pandemic-related workforce disruptions and return-to-office policies that differentially affected women.
The large increase in remuneration ratios suggests that executive compensation outpaced median wages, potentially driven by equity-based pay structures and market performance during the sample period. Country-level regulatory differences, industry trends, and macro-economic conditions may further shape these drifts.
Governance indicators show an increase in board independence and a decrease in the number of days to pay invoices, echoing Directive (EU) 2017/828 on shareholder voting rights regarding executive pay, or Directive 2011/7/EU on harmonized payment terms to curb abusive payment practices towards small-medium enterprises. In contrast, increasing lobbying expenditures since 2019 point to the need for stronger disclosure and oversight. Overall, these patterns highlight the limits of voluntary corporate action and underscore the role of policy in addressing persistent transparency gaps across ESG dimensions.
Our ML framework has several strengths. First, it generates a highly granular dataset that covers 501 different ESG indicators across between top- and bottom-performing companies. For each panel, dot sizes indicate reporting intensity (i.e., the number of companies disclosing values for each indi-cator in a given year), n denotes the total number of company–year observations for the indicator, and Δ denotes the percentage change in indicator value between 2014 and 2023. Note that logarithmic axes are used to better visualize different orders of magnitude. For indicators recorded in percent, we use linear scales.
nine ESRS topics. Consequently, our dataset is considerably more comprehensive than other available datasets, including those from commercial providers (e.g., Refinitiv60). Second, our framework is scalable, which enables automated ESG analysis across time, indus-tries, and countries. The scalability also enables regular updates of our dataset, as well as extending the dataset to emerging ESG frameworks and other geographical regions. Third, we open-source both our dataset and framework to democratize access to ESG data. So far, ESG data is often buried in lengthy reports or needs to be purchased at high cost from commercial providers that invest in the costly process of manual extraction, or that use their own, often diverging, definitions of ESG indicators.
By providing open-source access, we make ESG data freely accessible to stakeholders including policy-makers, investors, and societal actors.
Our work is subject to the following limitations. First, the observed patterns are consistent with multiple causes, yet we refrain from such causal claims and interpret the results as descriptive, opening avenues for future research using our dataset. Second, as we apply the ESRS framework retrospectively to reports published before ESRS became mandatory, documented data gaps may partly reflect differences in the reporting frameworks under which companies report. Transparency scores should therefore be interpreted as align-ment with current ESRS requirements. However, ESRS are largely based on previously established frameworks, especially GRI, and thus overlap substantially in underlying indicators, making our findings relevant beyond ESRS-specific requirements. Third, sustainability reports are self-reported and, during our sample period, subject to limited assurance.
While our framework can detect what is disclosed, it cannot identify strategic non-disclosure. Nevertheless, such selective disclosures should be mitigated through future regulations, including ESRS, with mandatory audits of such data being rolled out. Impor-tantly, outright falsification of financial or ESRS reports constitutes a legal offence under existing capital market regulations, which may deter systematic data manipulation. Moreover, non-retrieval of an indicator can reflect either non-disclosure or imperfect extraction (e.g., due to ambiguous information in the report); hence, absence of an extracted value does not necessarily represent the absence of reporting. Fourth, our sample comprises the 600 largest listed
European companies, placing it at the upper tail of company size, visibility, and institutional coverage. Our findings may not generalize to smaller or private companies, or to non-European markets where regulatory pressure and stakeholder expectations differ substantially. Nevertheless, our ML framework is scalable, enabling future research to extend the analysis to other regions and company populations. Fifth, the accuracy of the RAG pipeline depends on the structure, clarity, completeness, and consistency of the corporate reports ana-lyzed. Variations in data quality—such as missing or ambiguous infor-mation—can affect extraction accuracy, and LLM-based extraction may be prone to errors arising from hallucinations and the opaque rea-soning processes inherent in LLMs. We thus validated extracted values against two external datasets (see Supplementary Fig.
S1) to confirm the reliability of our framework.
In sum, targeted transparency through high-quality ESG dis-closures empowers policy-makers, financial stakeholders, and societal actors to hold companies accountable for their ESG efforts—in turn, advancing global sustainability goals. Our framework allows these stakeholders to systematically track corporate ESG-related transpar-ency and performance at scale. It thus provides a basis for evaluating the effectiveness of emerging regulatory initiatives on corporate ESG reporting and progress toward global sustainability goals.
Methods.
Data
Our sample consists of all companies listed in the STOXX Europe 600 stock index in 2023. The index constituents come from a broad range of countries and industries, covering nearly 90% of the invest-able market in Europe67. As a result, our dataset includes companies from 16 European countries and 11 industry sectors. The sectors are classified according to the Sustainable Industry Classification System (SICS), which was developed by the Sustainability Accounting Stan-dards Board (SASB) to group companies into sectors with comparable exposure to sustainability-related risks and opportunities68. Company fundamentals, such as market capitalization and sales, are sourced from the Worldscope database, a service by the London Stock Exchange Group43. An overview of the companies included in the dataset is provided in Supplementary Table S1.
Coverage across industries and countries over time is shown in Supplementary Fig. S5.
For all companies, we collect both annual reports and sustain-ability reports published as PDF files (overall 9,173 documents) between 2014 and 2023. The publicly available reports are sourced from intermediaries (i.e., the linked source the linked source) as well as directly from companies’ web-sites. The distribution of annual and sustainability reports over time is shown in Supplementary Fig. S3. Overall, the PDF files have an average length of 183 pages, which amounts to a total of 1,678,551 pages across the dataset. The word count is 88,965 on average per PDF. In our ML framework, we later concatenate both the annual and sustainability reports to form a panel dataset with company–year observations.
In these corporate reports, we then track 501 quantitative ESG indicators defined by European Sustainability Reporting Standards (ESRS)35. ESRS are described in the Supplementary Information S1. A complete list of the ESG indicators including metadata is provided in a supplementary CSV file.
ML framework
We developed an ML framework to extract quantitative ESG indicators from corporate reports based on retrieval-augmented generation (RAG)69. Our framework consists of five steps (see Supplementary Fig. S4): In step 1, we preprocess the corporate report PDF files and divide the text into meaningful chunks (preprocessing). In step 2, we embed the text chunks and store the vector representations in a vector database (indexing). In step 3, for each ESG indicator description, the top-k most similar report chunks are retrieved and re-ranked
(retrieval). In step 4, each ESG indicator description and the corre-sponding report chunks are passed in a tailored prompt to a pre-trained LLM, which outputs the extracted ESG indicator (generation). Finally, in step 5, the model output is postprocessed by standardizing values and units (postprocessing). Key methodological limitations are discussed in the Supplementary Information S4.
Step 1: preprocessing. First, we used PyMuPDF (v1.24.9)70 to parse the corporate report PDF files. The extracted text was cleaned by normalizing white spaces and resolving character encoding mis-matches. For each company and reporting year, the corresponding annual and sustainability reports were concatenated into a single document, denoted as d. Each document d was then partitioned into a sequence of semantically coherent chunks Cd = fc1, d, c2, d..., cM, d g, where M is the total number of chunks for document d, determined by the document length and the fixed chunk size and overlap parameters. This ensures efficient retrieval and compatibility with the context length constraints of the LLM in step 4.
Specifically, we applied recursive character text splitting71, which is designed to preserve meaningful contextual boundaries, to partition each document d into its corresponding set of chunks Cd. Each chunk ci,d was restricted to a maximum length of 400 characters, with an overlap of 100 characters between consecutive chunks, to ensure the granularity for precise similarity matching while minimizing the risk of sentence fragmentation72.
Step 2: indexing. We transformed each text chunk ci,d into a dense vector representation using the pre-trained embedding model all-MiniLM-L12-v273, implemented via sentence-transformers (v3.1.0). The model all-MiniLM-L12-v2 is an efficient yet semantically cap-able transformer-based model, which makes it well-suited forthe large-scale retrieval task in this work74. The model maps each chunk ci,d to a point in a 384-dimensional vector space, where semantically similar texts are positioned closer together, while dissimilar ones are more distant. The model all-MiniLM-L12-v2 is based on the MiniLM model family, which achieves state-of-the-art performance while improving computational efficiency75.
To improve its ability to encode semantically meaningful representations, all-MiniLM-L12-v2 was fine-tuned on a corpus of more than one billion sentence pairs using a contrastive learning objective. Given a sentence from a training pair, the all-MiniLM-L12-v2 model was trained to identify the associated sentence from a set of randomly sampled alternatives. Formally, cosine similarity is computed for all sentence pairs in a batch, and a cross-entropy loss is applied to maximize similarity for correct pairs while minimizing it for incorrect pairs. This ensures that semantically related chunks ci,d are embedded closer together in the vector space while unrelated ones are pushed apart73.
Step 3: retrieval. The resulting vector representations for all chunks ci,d were stored in a FAISS vector database, implemented using faiss-cpu (v1.8.0.post1), which enables scalable and low-latency retrieval76,77. Next, for each document d and each ESG indicator description, denoted as query q, we retrieved the top-30 most relevant report chunks, forming the set C30 q, d. The indicator description q con-sists of the relevant indicator subtopic and its description according to ESRS (e.g., gross scope 3 greenhouse gas emissions: category 1.1 cloud computing and data center services or percentage of employees at top management level: female).
We used a state-of-the-art hybrid search method that integrates cosine similarity search and keyword-based weighting72, search with equal implemented using faiss-cpu (v1.8.0.post1) and rank-bm25 (v0.2.2), which enables contextual semantic understanding with precise term-based retrieval. The top-30 most relevant chunks C30 q, d were re-ranked using the cross-encoder model Llama-Rank-V178, a fine-tuned variant of Llama3-8B-Instruct79 trained on relevance-annotated data. Unlike bi-encoder models such as all-MiniLM-L12-v2, which encode queries and documents independently, cross-encoders jointly process fine-grained query-chunk pairs (q, ci,d) to capture semantic relationships. Due to the high computational cost of cross-encoders, we employed a two-stage strategy: the efficient bi-encoder all-MiniLM-L12-v2 retrieved a candidate top-30 set, which was then re-ranked by Llama-Rank-V1.
This balances retrieval efficiency with ranking precision. Specifically, Llama-Rank-V1 computed a numeric relevance score for each query-chunk pair (q, ci,d), where ci, d 2 C30 q, d. The top-10 re-ranked chunks C10 q, d were expanded by concatenating each ci, d 2 C10 q, d with its preceding chunk ci−1,d and subsequent chunk ci +1,d, respectively. This then resulted in a set of extended chunks C10 q, d = fðci1, d k ci, d k ci + 1, d Þjci, d 2 C10 q, d g. This approach is known to mitigate the risk of missing important details that may span across chunk boundaries, such as large cohesive tables72. As a result, step 3 returned the most contextually relevant information from the corpus.
Step 4: generation. For each document d and query q, the corre-sponding retrieved report chunks C10 q, d were included in a tailored prompt and passed to a pre-trained LLM. The prompt is based on best practices in LLM prompt design80–82 and instructs the model to refrain from hallucinating if the requested ESG indicator is absent from the corporate report. The exact prompt is stated in Supplementary Information S3. Specifically, the model is asked to return the target value and unit of the requested ESG indicator described by q, if reported, in JSON format. JSON ensures structured, machine-readable output that facilitates downstream processing. We selected Llama-3.1-70B-Instruct83 as inference model, which offers state-of-the-art performance in natural language understanding and generation tasks without the excessive computational cost of larger models.
Llama-3.1-70B-Instruct is an externally developed, instruction-tuned large language model with publicly released weights. We did not update model parameters on our corpus, and all outputs were gen-erated via prompting conditioned on the retrieved report chunks. Llama-3.1-70B-Instruct fine-tuned Furthermore, was for instruction-following tasks79, which makes it better suited for struc-tured information extraction than standard generative models. We intentionally chose Llama-3.1-70B-Instruct over alternative, proprietary models (e.g., GPT-4) for two reasons: first, it gives state-of-the-art performance at the time of writing, and, second, it is open-source, meaning that end-users can easily adopt our framework84. We accessed Llama-3.1-70B-Instruct via the Together AI API85 using the together Python package (v1.3.3).
Informed by best practices for LLM-based research80, we set the temperature to 0, which leads the model to consistently select the token with the highest probability. This ensures deterministic output and, thus, promotes reproducibility. Reporting follows best practice86.
Step 5: postprocessing. Finally, we perform the following post-processing steps. Recall that the model output for each document d and query q contains the ESG indicator value and corresponding unit. The latter was standardized to adhere to the units specified by the European Financial Reporting AdvisoryGroup (EFRAG)35 as follows. We used regular expressions to unify the output strings, and the Pint library (v0.24.4)87 to convert the unified strings to standard units. The currencies were converted to USD according to the annual exchange rates provided by the Federal Reserve88. Thus, we enable robust downstream analysis and ensure a structured dataset for future research. The detailed regular expressions for the standardization process are in our source codes. Eventually, the final dataset is stored in a CSV file.
The pipeline was implemented in Python (v3.12.10). Statistical analyses, tables and figures were generated using custom Python and R (v4.4.0) scripts. Complete package versions and installation require-ments are provided in the public repository listed in the Code Avail-ability statement and in the Reporting Summary.
Validation
We validated our ML framework using two datasets: a proprietary dataset of Nperf = 212, 713 values across 38 ESG indicators, sourced from LSEG Refinitiv, a provider of financial and ESG data60; and a human-annotated subset of Nperf = 210 values across 21 ESG indicators (the raw data is made publicly available in our code base). The pro-prietary dataset was compiled by Refinitiv from publicly available sources, including company disclosures, stock exchange filings, and third-party information such as news websites. However, not all data points are directly cross-verified by the companies themselves. The human-annotated dataset was annotated independently by two PhD-level researchers (K.F. and L.K.) familiar with ESRS reporting, each of whom annotated half of the dataset, such that the entire dataset was covered by expert annotation.
To ensure the reliability of the anno-tations, another annotator—a graduate-level university student with prior experience in corporate ESG analysis—independently annotated the full subset. Then, the inter-rater reliability was assessed. Specifi-cally, we computed the intraclass correlation coefficient (ICC) for each subset using a two-way mixed-effects model with absolute agreement and single measures89. We found consistently high ICC values across both annotator combinations (subset 1: ICC = 0.978, 95% CI [0.968, 0.985]; subset 2: ICC = 0.998, 95% CI [0.997, 0.998]), which indicates excellent agreement and thus strong support for the reliability of the manual annotations.
However, both datasets come with inherent limitations. The proprietary dataset covers only 38 ESG indicators—which is fewer than the 501 indicators defined by ESRS. The human-annotated dataset is constrained by the manual nature of the annotation process, which limits scalability and thus size. Manual annotation is particularly time-intensive due to the length and complexity of sustainability reports, which often span hundreds of pages. As a result, each value took us ~15–20 minutes to annotate, which corresponds to roughly 105–140 h (or around three weeks under regular working hours) of full-time labeling. We initially considered delegating the annotation task to trained university students; however, this approach proved unsuccessful due to insufficient annotation quality. Despite thorough training, stu-dent annotators frequently introduced mismatches in definitions or units.
Due to these issues, we reverted to relying on experienced researchers for annotation. Hence, both validation datasets remain substantially smaller in scope—both in terms of ESG indicators and number of companies—than our machine-learning-based dataset.
We then proceeded as follows. For both datasets, we compared our extracted ESG indicators against the corresponding values from the validation datasets (see Supplementary Fig. S1 for both Refinitiv and the manually annotated validation dataset). In both cases, the correspondence is strong: for the proprietary dataset, the estimated slope from the OLS regression is β = 0.885 (p < 0.001), with an adjusted R2 of 0.91, and, for the manually annotated dataset, the estimated slope from the OLS regression is also β = 0.885 (p < 0.001), with an adjusted R2 of 0.93. Hence, we find a strong agreement between our ML fra-mework and the benchmark datasets. To further validate the per-formance of our framework, we evaluated standardized mean absolute error (sMAE) and standardized root mean squared error (sRMSE).
These metrics normalize absolute and squared errors by the inter-quartile range (IQR) of the human-annotated values for each indicator, which enables comparison across indicators with different scales for better interpretability while providing robustness against outliers. sMAE measures deviations between predictions and human-annotated reference values, while sRMSE additionally penalizes large errors, thus providing a more conservative assessment. Macro-averaged across indicators, the framework achieves an sMAE of 0.21 (95% CI [0.08, 0.37]) and sRMSE of 0.39 (95% CI [0.15, 0.54]) (Supplementary Table S5), demonstrating robust alignment with human-annotated values.
Indicators with undefined denominators (IQR = 0) and one indicator with only three observations, where a single prediction error produced unstable standardized metrics, were excluded from the macro-averages to maintain reliability of the aggregated estimates. Confidence intervals for the aggregated metrics were estimated using a nonparametric percentile bootstrap with 2,000 resamples. To assesspotential heterogeneity acrossESG indicators, we computed the agreement for each ML-extracted ESG indicator with the proprietary dataset from Refinitiv. We consistently find positive and statistically significant correlations for the majority of ESG indicators (Supple-mentary Fig. S2), with particularly high agreement for indicators rela-ted to greenhouse gas emissions, suggesting that our extraction approach is not biased toward specific topics but overall robust.
We focused on the Refinitiv dataset in the heterogeneity analysis because the human-annotated validation set is too small (all n≤10) to yield statistically reliable estimates, whereas the proprietary dataset typi-cally contains n ≫ 100 values per indicator, which thus allows for such a heterogeneity analysis.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Data availability.
The complete dataset of extracted ESG indicators generated in this study has been deposited in the Open Science Framework (OSF) repository under accession code q2jpv (the linked source). The corporate annual and sustainability reports used in this study have been deposited in Harvard Dataverse under accession code DVN/ 84HKPS (the linked source). The remaining third-party datasets used for validation and additional analyses are available under restricted access because they are proprietary and require a subscription or license; access can be obtained from the respective data providers.
Specifically, the LSEG Refinitiv ESG data used for vali-dation are available from LSEG; company fundamentals data are available from LSEG Worldscope Fundamentals (the linked source); ESG controversies scores are avail-able from LSEG ESG Scores (the linked source sustainable-finance/esg-scores); and MSCI ESG ratings are available from MSCI.
Code availability
The code used to reproduce the analyses, tables, and figures reported in this study is available on GitHub at the linked source forsterkerstin/corporate-sustainability-trackerand has been archived on Zenodo90. Llama-3.1-70B-Instruct is available at the linked source.
1. Carbon Disclosure Project.
The Carbon Majors database: launch report. the linked source MajorsLaunchReport.pdf (2024).
2. Jiang, X., Kim, S. & Lu, S.
Limited accountability and awareness of corporate emissions target outcomes. Nat. Clim. Change 15, 279–286 (2025).
3. Cenci, S., Burato, M., Rei, M. & Zollo, M.
The alignment of compa-nies’ sustainability behavior and emissions with global climate tar-gets. Nat. Commun. 14, 7831 (2023).
4. Rekker, S., Ives, M.
C., Wade, B., Webb, L. & Greig, C. Measuring corporate Paris compliance using a strict science-based approach. Nat. Commun. 13, 4441 (2022).
5. Österblom, H., Bebbington, J., Blasiak, R., Sobkowiak, M. & Folke, C.
Transnational corporations, biosphere stewardship, and sustain-able futures. Annu. Rev. Environ. Resour. 47, 609–635 (2022).
6. Climate Action 100+.
Climate Action 100+. the linked source. climateaction100.org/ (2025).
7. The Climate Pledge.
The Climate Pledge. the linked source. theclimatepledge.com/ (2025).
8. Science Based Targets Initiative.
Ambitious corporate climate action. the linked source (2025).
9. Eurostat.
Gender pay gap statistics the linked source statistics-explained/index.php?title=Genderpaygap statistics (2025).
10. McKinsey & Company.
Diversity matters even more: the case for holistic impact. the linked source diversity-and-inclusion/diversity-matters-even-more-the-case-for-holistic-impact#/ (2023).
11. Beck, J. et al.
Addressing data gaps in sustainability reporting: a benchmark dataset for greenhouse gas emission extraction. Sci. Data 12, 1497 (2025).
12. Cohen, S., Kadach, I. & Ormazabal, G.
Institutional investors, cli-mate disclosure, and carbon emissions. J. Account. Econ. 76, 101640 (2023).
13. Klaaßen, L. & Stoll, C.
Harmonizing corporate carbon footprints. Nat. Commun. 12, 6149 (2021).
14. Mastrandrea, R., ter Burg, R., Shan, Y., Hubacek, K. & Ruzzenenti, F.
Assessments of the environmental performance of global compa-nies need to account for company size. Commun. Earth Environ. 5, 42 (2024).
15. Nguyen, Q., Diaz-Rainey, I. & Kuruppuarachchi, D.
Predicting cor-porate carbon footprints for climate finance risk analyses: a machine learning approach. Energy Econ. 95, 105129 (2021).
16. Zhang, Z. et al.
Embodied carbon emissions in the supply chains of multinational enterprises. Nat. Clim. Change 10, 1096–1101 (2020).
17. Arslan, M., Munawar, S. & Sibilla, M.
Sustainable energy decision-making with an RAG-LLM system. In 2024 International Conference on Decision Aid Sciences and Applications (2024).
18. Ardic, O., Ozturk, M.
U., Demirtas, I. & Arslan, S. Information extraction from sustainability reports in Turkish through RAG approach. In 2024 32nd Signal Processing and Communications Applications Conference (2024).
19. Yang, J.-Y. et al.
EcoSmartGuide: language learning model and retrieval-augmented generation-based platform for streamlined environmental, social, and governance information access and report generation. In 2024 IEEE 6th Eurasia Conference on Biome-dical Engineering, Healthcare and Sustainability, 343–347 (2024).
20. Bronzini, M., Nicolini, C., Lepri, B., Passerini, A. & Staiano, J.
Glitter or gold? deriving structured insights from sustainability reports via large language models. EPJ Data Sci. 13, 41 (2024).
21. Colesanti Senni, C., Schimanski, T., Bingler, J., Ni, J. & Leippold, M.
Using AI to assess corporate climate transition disclosures. Environ. Res. Commun. 7, 021010 (2025).
22. Zou, Y. et al.
ESGReveal: an LLM-based approach for extracting
23. Ni, J. et al.
CHATREPORT: democratizing sustainability disclosure analysis through LLM-based tools. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 21–51 (Association for Computational Lin-guistics, Singapore, 2023).
24. Schimanski, T., Ni, J., Martín, R.
S., Ranger, N. & Leippold, M. Clim-
Retrieve: a benchmarking dataset for information retrieval from corporate climate disclosures. In Proceedings of the 2024 Con-ference on Empirical Methods in Natural Language Processing, 17509–17524 (Association for Computational Linguistics, Miami, Florida, USA, 2024).
25. Bingler, J.
A., Kraus, M., Leippold, M. & Webersinke, N. How cheap talk in climate disclosures relates to climate initiatives, corporate emissions, and reputation risk. J. Bank. Financ. 164, 107191 (2024).
26. Schimanski, T., Bingler, J., Kraus, M., Hyslop, C. & Leippold, M.
ClimateBERT-NetZero: detecting and assessing net zero and reduction targets. In Proceedings of the 2023 Conference on
Empirical Methods in Natural Language Processing, 15745–15756 (Association for Computational Linguistics, Singapore, 2023).
27. Bochkay, K., Brown, S.
V., Leone, A. J. & Tucker, J. W. Textual ana-lysis in accounting: What’s next? Contemp. Account. Res. 40, 765–805 (2023).
28. Loughran, T. & McDonald, B.
Textual analysis in accounting and
29. KPMG & Google Cloud.
Closing the disconnect in ESG data. the linked source (2021).
30. Weil, D., Graham, M. & Fung, A.
Targeting transparency. Science 31. Hombach, K. & Sellhorn, T
Shaping corporate actions through targeted transparency regulation: a framework and review of extant evidence. Schmalenbach Bus. Rev. 71, 137–168 (2019).
32. Christensen, H
B., Floyd, E., Liu, L. Y. & Maffett, M. The real effects of mandated information on social responsibility in financial reports: evidence from mine-safety records. J. Account. Econ. 64, 284–304 (2017).
33. Bonetti, P., Leuz, C. & Michelon, G
Internalizing externalities through public pressure: transparency regulation for fracking and water quality. the linked source (2024).
34. European Parliament and the Council of the European Union
Directive (EU) 2022/2464 of the European Parliament and of the Council of 14 December 2022 amending Regulation (EU) No 537/ 2014, Directive 2004/109/EC, Directive 2006/43/EC and Directive 2013/34/EU, as regards corporate sustainability reporting (text with EEA relevance). Official Journal of the European Union L322, 15–80 (2022).
35. European Financial Reporting Advisory Group (EFRAG)
ESRS XBRL taxonomy. the linked source (2024).
36. Organisation for Economic Co-operation and Development
OECD review on aligning finance with climate goals. the linked source b9b7ce49-en.pdf (2024).
37. Greenstone, M., Leuz, C. & Breuer, P
Mandatory disclosure would reveal corporate carbon damages. Science 381, 837–840 (2023).
38. MSCI Inc
ESG ratings. the linked source sustainability-solutions/esg-ratings (2025).
39. London Stock Exchange Group
Environmental, social and gov-ernance scores from LSEG. the linked source data-analytics/enus/documents/methodology/lseg-esg-scores-methodology.pdf (2024).
40. European Commission
Corporate sustainability due diligence. the linked source corporate-sustainability-due-diligenceen (2024).
41. European Commission
Zero pollution action plan. the linked source en (2021).
42. GHG Protocol
Corporate Value Chain (Scope 3) Accounting and Reporting Standard. the linked source (2011).
43. London Stock Exchange Group
Worldscope Fundamentals. the linked source (2025).
44. Callaghan, M. et al
Machine learning map of climate policy litera-ture reveals disparities between scientific attention, policy density, and emissions. npj Clim. Action 4, 7 (2025).
45. Callaghan, M. et al
Machine-learning-based evidence and attribu-tion mapping of 100,000 climate impact studies. Nat. Clim. Change 11, 966–972 (2021).
46. Gehricke, S., Leippold, M., Schimanski, T. & Delgado Fajardo, C
To disclose, or not to disclose: evaluating the effectiveness of man-datory climate-related disclosure. SSRN Electronic J. the linked source (2025).
47. Schimanski, T. et al
Bridging the gap in ESG measurement: using NLP to quantify environmental, social, and governance commu-nication. Financ. Res. Lett. 61, 104979 (2024).
48. Toetzke, M., Probst, B., Feuerriegel, S., Anadon, L
D. & Hoffmann, V. H. Analyzing the dynamics of innovation networks in climate tech-nologies using large language models. the linked source 4810933 (2024).
49. Toetzke, M., Stünzi, A. & Egli, F
Consistent and replicable estima-tion of bilateral climate finance. Nat. Clim. Change 12, 897–900 (2022).
50. Toetzke, M., Banholzer, N. & Feuerriegel, S
Monitoring global development aid with machine learning. Nat. Sustainability 5, 533–541 (2022).
51. Lin, Y., Shen, R., Wang, J. & Yu, Y
J. Global evolution of environ-mental and social disclosure in annual reports. J. Account. Res. 62, 1941–1988 (2024).
52. Ioannou, I. & Serafeim, G
The consequences of mandatory corpo-rate sustainability reporting. In McWilliams, A., Rupp, D. E., Siegel, D. S., Stahl, G. K. & Waldman, D. A. (eds.) The Oxford Handbook of Corporate Social Responsibility: Psychological and Organizational Perspectives.
53. Ali, A., Klasa, S. & Yeung, E
Industry concentration and corporate disclosure policy. J. Account. Econ. 58, 240–264 (2014).
54. Christensen, D
M., Serafeim, G. & Sikochi, A. Why is corporate vir-tue in the eye of the beholder? the case of ESG ratings. Account. Rev. 97, 147–175 (2022).
55. Bochkay, K., Hales, J. & Serafeim, G
Disclosure standards and communication norms: evidence of voluntary sustainability stan-dards as a coordinating device for capital markets. Rev. Account. Stud. 30, 3021–3064 (2025).
56. Berg, F., Kölbel, J
F. & Rigobon, R. Aggregate confusion: the divergence of ESG ratings. Rev. Financ. 26, 1315–1344 (2022).
57. Berg, F., Fabisik, K. & Sautner, Z
Is history repeating itself? The (un) predictable past of ESG ratings. the linked source 3722087 (2021).
58. Ilhan, E., Krueger, P., Sautner, Z. & Starks, L
T. Climate risk dis-closure and institutional investors. Rev. Financial Stud. 36, 2617–2650 (2023).
59. Clarkson, P
M., Li, Y., Richardson, G. D. & Vasvari, F. P. Revisiting the relation between environmental performance and environmental disclosure: an empirical analysis. Account., Organ. Soc. 33, 303–327 (2008).
60. London Stock Exchange Group
ESG data. the linked source en/data-analytics/financial-data/company-data/esg-data (2025). 61. Derrien, F., Krueger, P., Landier, A. & Yao, T. ESG news, future cash flows, and firm value. J. Financ. 80, 3499–3554 (2025).
62. Tamasiga, P., Onyeaka, H., Bakwena, M. & Ouassou, E
H. Beyond compliance: evaluating the role of environmental, social and gov-ernance disclosures in enhancing firm value and performance. SN Bus. Econ. 4, 118 (2024).
63. Garel, A., Romec, A., Sautner, Z. & Wagner, A
F. Do investors care about biodiversity? Rev. Financ. 28, 1151–1186 (2024).
64. Bourveau, T., Chowdhury, M., Le, A. & Rouen, E
Human capital disclosures. SSRN Electronic Journal. the linked source 4138543 (2022).
65. Taskforce on Nature-related Financial Disclosures
Recommenda-tions of the taskforce on nature-related financial disclosures. the linked source Recommendations-of-the-Taskforce-on-Nature-related-Financial-Disclosures.pdf (2023).
66. European Parliament and the Council of the European Union
Reg-ulation (EU) 2021/1119 of the European Parliament and of the Council of 30 June 2021 establishing the framework for achieving climate neutrality and amending Regulations (EC) No 401/2009 and (EU) 2018/1999 ("European Climate Law”). Official Journal of the Eur-opean Union L243, 1-17 (2021).
67. STOXX Ltd
STOXX Europe 600. the linked source sxxp/ (2025).
68. Sustainability Accounting Standards Board
Find your industry. the linked source (2025).
69. Lewis, P. et al
Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems, 33, 9459–9474.
70. Artifex Software
PyMuPDF. the linked source PyMuPDF (2025).
71. LangChain
How to recursively split text by characters. the linked source (2025).
72. Wang, X. et al
Searching for best practices in retrieval-augmented generation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 17716-17736 (Association for Computational Linguistics, Miami, Florida, USA, 2024).
73. Hugging Face. all-MiniLM-L12-v2. the linked source sentence-transformers/all-MiniLM-L12-v2 (2021).
74. Ozyurt, Y., Feuerriegel, S. & Zhang, C
Document-level in-context few-shot relation extraction via pre-trained language models. the linked source (2024).
75. Wang, W. et al
MINILM: deep self-attention distillation for task-agnostic compression of pre-trained transformers. In Proceedings of the 34th International Conference on Neural Information Processing Systems (2020).
76. Douze, M. et al 77. Li, S., Stenzel, L., Eickhoff, C. & Bahrainian, S
A. Enhancing retrieval-augmented generation: a study of best practices. In Proceedings of the 31st International Conference on Computational Linguistics, 6705–6717 (Association for Computational Linguistics, 2025).
78. Ginart, A., Kodali, N. & Emmons, J
Introducing LlamaRank: a state-of-the-art reranker for trusted AI. the linked source blog/llamarank/ (2024).
79. Grattafiori, A. et al 80. Feuerriegel, S. et al
Using natural language processing to analyse text data in behavioural science. Nat. Rev. Psychol. 4, 96–111 (2025).
81. Lin, Z
How to write effective prompts for large language models. Nat. Hum. Behav. 8, 611–615 (2024).
82. Giray, L
Prompt engineering with ChatGPT: a guide for academic writers. Ann. Biomed. Eng. 51, 2629–2633 (2023).
83. Meta AI
Llama-3.1-70B-Instruct. the linked source (2024).
84. Shrestha, Y
R., von Krogh, G. & Feuerriegel, S. Building open-source AI. Nat. Comput. Sci. 3, 908–911 (2023).
85. Together AI
Together AI the linked source (2025).
86. Feuerriegel, S. et al
A reporting checklist for large language models in behavioural science. Nature Human Behaviour (2026).
87. Grecco, H
E. Pint: makes units easy. the linked source pint (2024).
88. Board of Governors of the Federal Reserve System
G.5/H.10 for-eign exchange rates. the linked source datadownload/Choose.aspx?rel=H10 (2025).
89. Koo, T
K. & Li, M. Y. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J. Chiropr. Med. 15, 155–163 (2016).
90. Forster, K. et al
Assessing corporate sustainability with large lan-guage models: Evidence from Europe. GitHub repository archived at Zenodo. the linked source (2026).
World Health Organization. Timeline: WHO’s COVID-19 response. 91. the linked source! (2022).
Acknowledgements.
We thank Sandra Denk for her excellent research assistance and Tobias Schimanski for valuable comments on the draft manuscript.
K.F. implemented the ML pipeline. K.F., L.K., and S.F. designed the ML framework. V.W. and L.K. collected the datasets. K.F., V.W., and L.K. performed the analysis. K.F., M.A.M., T.S., and S.F. prepared the first draft of the manuscript. All authors contributed to conceptualization, manuscript writing, and approved the manuscript.
Funding
L.K., V.W., M.A.M., and T.S. disclose support for the research of this work from the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) – project ID 403041268 – TRR 266. S.F. discloses support from the Swiss National Science Foundation (SNSF), grants 186932 and 197485. Open Access funding enabled and organized by Projekt DEAL.
Competing interests
T.S. serves as a paid supervisory board member of Deloitte Germany. Deloitte Germany, like other consulting and assurance firms, may have clients potentially affected by ESG reporting regulation. This position is unrelated to the design, execution, or analysis of the present study, and all interpretations and conclusions are solely those of the authors. All other authors declare no competing interests.
Additional information
Supplementary information The online version contains supplementary material available at the linked source.
Peer review information Nature Communications thanks the anon-ymous reviewer(s) for their contribution to the peer review of this work. A peer review file is available.
Reprints and permissions information is available at the linked source
Publisher’s note Springer Nature remains neutral with regard to jur-isdictional claims in published maps and institutional affiliations.
© The Author(s) 2026
Author contributions