CogDeBias: An LLM-Based Multilingual Framework for Cognitive Bias Detection and Mitigation in Corporate Decision-Making Texts
1 More Paper · Full Reading

About this paper
A full audio edition of this paper.
Authors: Y. Shen, W. Yang, Y. Shen
Publication date: 2026
Read the paper: https://doi.org/10.1109/access.2026.3717981
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “CogDeBias: An LLM-Based Multilingual Framework for Cognitive Bias Detection and Mitigation in Corporate Decision-Making Texts,” by Y. Shen, W. Yang, and Y. Shen. Published in 2026.
Abstract.
Cognitive biases embedded in corporate strategic communications pose significant risks to investment decisions and governance quality, yet existing detection approaches rely on manual analysis that lacks scalability. This study presents CogDeBias, a multilingual framework integrating large language models (LLMs) and machine learning (ML) for automated detection and mitigation of cognitive biases in enterprise decision-making texts. A bilingual annotated corpus of 1,200 corporate annual reports (600 English, 600 Chinese) was constructed, encompassing 30,416 sentences, of which 8,082 are bias-positive (26.6%) and carry approximately 9,230 bias-category instances across six categories: confirmation bias, sunk cost fallacy, overconfidence, anchoring effect, bandwagon effect, and recency bias.
The hybrid architecture combines XLM-RoBERTa multilingual encoding, Llama-3-70B prompt-based classification, and XGBoost ensemble learning, achieving a weighted F1-score of 0.82 on the test set, surpassing rule-based (0.59), BERT-based (0.72), zero-shot GPT-4 (0.76), and few-shot Llama-3 (0.785) baselines, with the 3.5-percentage-point margin over the strongest baseline statistically significant (p = 0.008, McNemar’s test). On a fully expert-annotated subset labelled without any model assistance, the framework retained a weighted F1 of 0.81, indicating that the reported performance is not an artefact of model-to-model label agreement. Cross-lingual evaluation revealed English F1 of 0.84 versus Chinese F1 of 0.80, with zero-shot transfer from English to Chinese yielding F1 of 0.68, improving to 0.78 with minimal target-language fine-tuning.
The Mixtral-8x7B-powered mitigation generator produced actionable suggestions rated 4.1/5 by 30 financial analysts (Fleiss’ kappa = 0.71). These suggestions are decision-support outputs for human review, not automated corrections to regulated filings. Processing efficiency of 0.42 seconds for core model inference (3.0 seconds end-to-end per document) supports batch-oriented use as decision support in investor due diligence, corporate audit, and regulatory monitoring workflows, subject to domain-specific calibration. This work establishes a computational paradigm for decision science applications, demonstrating that linguistic patterns reliably surface systematic reasoning distortions across languages and corporate contexts.
Introduction.
Decision-making failure within corporations has significant economic implications. Recent history within the business
The associate editor coordinating the review of this manuscript and approving it for publication was Wai-Keung Fung.
world has highlighted examples of how cognitive biases within management discourse contributed to decision-making failure, ultimately costing companies billions of dollars in shareholder equity. These biases, or systematic deviations from rational decision-making, are often not directly observable within financial data analysis, occurring at the semantic level within annual reports, earnings announcements, and other strategic decision-making dis-course. This is in contrast to quantitatively derived metrics that indicate decision-making failure within companies, such as financial distress indicators. The increasing acknowledg-ment of cognitive factors within financial markets serves to further emphasize the need for an automated detection of cognitive biases.
Despite growing recognition of cognitive factors in finan-cial markets, three persistent gaps limit existing approaches: (i) limited scalability of manual annotation, (ii) linguistic fragmentation between English-dominant tools and bilingual corporate contexts, and (iii) methodological fragmentation across single-bias studies. A systematic review of these gaps, organized by research stream, is presented in Section II.
This research introduces three main innovations. First, this study establishes a six-category cognitive bias taxon-omy (confirmation bias, sunk cost fallacy, overconfidence, anchoring effect, bandwagon effect, and recency bias) operationalized through corpus-validated linguistic patterns, enabling automated detection at weighted F1 = 0.82 on a 240-document held-out test set—outperforming the strongest baseline (Few-shot Llama-3, F1 = 0.785) by 3.5 percent-age points (p = 0.008, McNemar’s test). The taxonomy adapts foundational heuristics-and-biases theory and recent medical-domain bias classifications to the corporate disclosure context, addressing a relatively underdeveloped area in enterprise decision-making bias research, building on prior investor-level bias modeling.
Second, this study introduces what is, to our knowledge, among the first end-to-end frameworks to jointly address cognitive-bias detection and mitigation in bilingual (English–Chinese) corporate disclosures (Table 1), integrating XLM-RoBERTa multi-lingual encoding, hybrid LLM-ML bias classification, and Mixtral-8x7B-powered mitigation generation for enterprise disclosure analysis. By moving beyond passive annotation to actionable rewriting suggestions, the framework delivers mitigation outputs rated 4.1/5 by 30 financial analysts (Fleiss’ kappa = 0.71, with 86% of suggestions judged highly rele-vant), and achieves end-to-end processing of 3.0 seconds per document (0.42-second core model inference), supporting scalable batch processing to assist investor due diligence and regulatory monitoring workflows.
This responds to the call of practitioners for useful tools, rather than purely proof-of-concept systems, in decision-making bias research. Third, this study constructs a sentence-level annotated bilin-gual (English–Chinese) cognitive-bias corpus for corporate disclosures (Table 1), comprising 1,200 documents and 8,082 bias-positive sentences (≈9,230 category-level bias instances) across six categories. Fine-grained, multi-domain annotated benchmarks with cross-platform transfer have proven valuable in adjacent text-classification tasks, a design principle this corpus follows. While benchmark datasets exist for adjacent financial NLP tasks such as opinion mining and question answering, no comparable bilingual annotated resource is available for cognitive-bias detection in corporate contexts.
The corpus enables systematic cross-lingual evaluation, demonstrating English F1 = 0.84 and Chinese F1 = 0.80, with zero-shot English-to-Chinese trans-fer at F1 = 0.68 and rapid recovery to F1 = 0.78 using only 100 target-language documents. As this few-shot adaptation effect does not survive Benjamini–Hochberg correction (p = 0.041), it is reported as suggestive rather than confirmatory.
Objectives of the research revolve around the development of an automated, scalable, and multilingual bias detection system that is deployable in a real-world setting. This includes the development of classification standards and annotation methods that have been approved by domain experts, the creation of a hybrid detection system that combines the strengths of neural and traditional ML tech-niques, as well as the empirical validation of the accuracy of the detection system and the utility of the mitigation system through human evaluation. The rest of the paper is organized as follows: Section II reviews related work across cognitive-bias detection, financial-text NLP, and multilingual NLP, and identifies the gaps this study addresses.
Section III describes data acquisition, bias taxonomy, system architecture, and evaluation; Section IV presents the results in terms of detection metrics, cross-lingual analysis, and mitigation quality assessment; Section V discusses the key findings, advantages, and limitations; and Section VI concludes the paper by highlighting the contributions and prospects.
II. RELATED WORK.
There are currently three major limitations associated with existing research: First, scalability is a major concern, and individual analysts examining corporate disclosures have severe throughput limits, making it difficult to scale up sample sizes and generalize findings on a larger scale. The study on corporate foresight revealed that detecting biases and fallacies within corporate planning documents demands specialized human expertise. Such expertise does not scale to thousands of annual reports and disclosures. The gap addressed here is that such manual, expert-dependent analysis does not scale to large bilingual corpora; CogDe-Bias replaces per-analyst throughput limits with automated sentence-level detection.
Second, linguistic fragmentation remains a major concern, and sentiment analysis tools have advanced for English-language financial disclosures, while corporate settings, especially Chinese A-share markets and Hong Kong markets, which use a mix of languages, are still not addressed. The detection of biases across languages, especially for Chinese and Hong Kong markets, remains underexplored, given that gaps have already been identified for multilingual NLP systems for non-English languages. The gap addressed here is the absence of a bias-detection system and annotated resource spanning English and Chinese corporate disclosures, which the present bilingual corpus and cross-lingual evaluation directly fill.
Third, methodological fragmentation is another concern, and prior research on AI and its application in behavioral finance have addressed individual biases, while a comprehensive framework remains absent, and applications remain at an early stage of practical deployment. Notably, the recent global review by Şeker et al. catalogues how cognitive biases have been operationalized across heuristic-based, machine-learning, and LLM-based approaches in investor decision-making, demonstrating that this research stream is well-established at the investor level but identifying a persistent gap in detection at the corporate-disclosure level. The gap addressed here is the lack of a unified framework that covers multiple bias categories and couples detection with actionable mitigation, rather than treating single biases in isolation—the gap the present work fills.
Positioned against the three literature streams, CogDeBias’s contribution is the specific combination none of them provides. Deep-learning and transformer financial-NLP work (e.g., lexicon-to-transformer sentiment, risk-disclosure topic model-ing ) targets sentiment or topic rather than reasoning-level cognitive bias, and is English-only. Cognitive-bias work in LLMs,,, probes biases inside models on medical or general benchmarks, not bias expressed in corporate disclosures, and stops at detection without mitigation. Multilingual NLP reviews, document the English–non-English gap but provide no bias-detection system or annotated resource.
CogDeBias is distinguished by the combination of four features: (i) operating on corporate disclosures, (ii) covering English and Chinese, (iii) integrating hybrid LLM–ML detection with generation-based mitigation, and (iv) supplying a sentence-level bilin-gual bias benchmark. No prior work in Table 1 combines all four; it is this combination, rather than any single element, that is novel.
Significant gaps in technical, cognitive, and behavioral capabilities of language-based systems have long been identified. Large language models (LLMs) hold tremendous promise in bridging these gaps. Transformer-based architec-tures, which are trained on multilingual data, have shown unprecedented capabilities in cross-lingual transfer learning. For example, the XLM-RoBERTa model has shown robust semantic understanding in 100 different languages using unsupervised learning of representations. In addition to technical capabilities, empirical research demonstrates that large language models are subject to human-like biases in their decision-making.
For example, in one of the first studies, it was shown that pre-ChatGPT models were subject to human-like framing effects, anchoring effects, and System 1 thinking, as observed in cognitive reflection and semantic illusion tests, which are subject to human irrationality. However, reinforcement learning from human feedback significantly reduces these effects. This demonstrates that, even as alignment training decreases the likelihood of human-like biases in decision-making, it also decreases the ability of the model to detect bias in input data. New methodologies in medical domain bias assessment are now being developed and are of significant importance to corporate applications. For example, in one recent study, datasets were created to assess six different decision-making biases in medical domains, including confirmation bias, availability heuristic, and overconfidence.
The results showed that state-of-the-art proprietary models are robust to bias-inducing prompts, whereas open-source models showed significant performance degradation. This framework of using adversarial prompts to detect latent patterns of bias can be applied to the present study. A second study, which surveyed the phenomenon of cognitive biases in large lan-guage models, validates template-based approaches to corporate text analysis. Table 1 summarizes ten representative works across these three streams and positions CogDeBias against the identified gaps.
III. MATERIALS AND METHODS.
A. DATASET CONSTRUCTION AND PREPROCESSING
The dataset was collected from four major sources, which cover various regulatory environments. English-language documents, totaling 600, were collected from the SEC EDGAR system. The authors specifically filtered the database to include the annual reports filed by publicly traded U.S. corporations, denoted as 10-K reports, from January 2023 to August 2025. Chinese-language documents, also totaling 600, were collected from the Shanghai Stock Exchange disclosure platform. The authors filtered the database to include annual reports filed by A-share listed companies during the same period. Supporting materials were collected from the Hong Kong Stock Exchange and OpenCorporates database, which included dual-language documents and disclosures filed by smaller enterprises. The final dataset included 1,200 corporate documents, with an equal number of English and Chinese documents.
The study focused on recent events and strategic narratives. These were shaped by the post-pandemic recovery and by various technology disruptions, both of which are documented to enhance managerial cognitive biases. A text mining study of financial disclosures has found that the narrative section of financial reports reveals behavioral patterns that are distinct from numerical statements.
To ensure the inclusion of the documents, the following criteria were satisfied: Firstly, the texts were required to have substantive management discussion sections, such as ‘‘Management Discussion and Analysis’’ for U.S. filings or the equivalent sections for Chinese reports. These sections carry forward-looking strategic discourse, where cognitive biases are most likely to occur. Secondly, the completeness thresholds for the documents were set such that those with less than 95% structural integrity, based on the absence of critical sections or corruption in the encoding, were excluded from the dataset. Lastly, tables, figures, and chart captions were excluded, as they are not part of the running text. The remaining text was filtered for bias-indicative linguistic patterns.
This step also addressed common NLP preprocessing problems, such as character-encoding errors in cross-platform aggregation, which are especially prevalent in multilingual and low-resource text.
Raw documents were subjected to systematic transforma-tion through a series of four stages. In the text cleaning stage, HTML residues, XML tags from regulatory templates, and non-standard Unicode characters were removed, which are known to interfere with tokenization. In the next stage of sentence segmentation, the text was divided into discrete analytical units, while language classification was performed to partition the text into English and Chinese categories using the langdetect library. Finally, text vectorization was done to transform the text into a compatible format for the neural network model, which was achieved through the sentence-transformers library. Through the text preprocessing stage, 30,416 annotated sentences were obtained, as indicated in the dataset composition and partitioning statistics as presented below in Table 2.
The composition of the dataset is evenly distributed across the training, validation, and test sets, while the bias levels were consistent at 26.6% throughout the dataset.
Each document was annotated with filing timestamp and industry sector (Global Industry Classification Standard (GICS)). Industry-specific stress events (e.g., semiconductor export controls in Oct 2023/July 2024, identified through Bloomberg/Reuters archives) were catalogued. Documents filed within 90 days of stress events were labeled for post-hoc temporal analysis (Section V-A). Baseline comparison group (n = 780, 65%) comprised documents from non-disruption periods.
Documents were sampled to reflect sectoral diversity: English corpus included Technology (22%), Financials (18%), Healthcare (15%), Consumer (12%), Industrials (10%), and other sectors (23%); Chinese corpus comprised Manufacturing (28%), IT (19%), Financials (14%), Materials (11%), Consumer (10%), and others (18%). Industry classi-fications followed GICS standards for consistency.
Market capitalization distribution: Large-cap >$10B (35%, n = 420), Mid-cap $2-10B (45%, n = 540), Small-cap
<$2B (20%, n = 240). The mid-cap oversampling (45% vs ∼30% typical market weight) reflects institutional research coverage patterns and enhances generalizability across firm sizes, though limits direct applicability to cap-weighted portfolio analysis.
B. COGNITIVE BIAS TAXONOMY AND GROUND TRUTH LABELING
The cognitive bias taxonomy was developed based on the foundational research by Kahneman and Tversky on ’heuristics and biases in judgment under uncertainty’. Six cognitive biases were identified based on their prevalence within corporate-level strategic discourse and relevance to the behavioral finance literature on investor decision-making processes. Cognitive biases are systematic deviations from rational economic decision-making processes that often permeate management discourse. Confirmation bias occurs when managers selectively use supporting evidence for predetermined corporate strategies while overlooking disconfirming evidence. Sunk cost fallacy occurs when managers justify continued investments in failing projects based on prior invested capital rather than expected return on investment.
Overconfidence bias, prevalent in equity markets and linked to suboptimal trading and portfolio performance, manifests as management discourse that uses absolute language in discussing market outlook or competitive advantage. Anchoring effect occurs through con-tinued reliance on initial targets or price anchors regardless of changing market conditions. Bandwagon effect occurs through the adoption of industry trends characterized as ‘‘best practices’’ without independent strategic analysis. Recency bias occurs through the use of short-term metrics as a basis for long-term planning decisions.
The theoretical origins of the six categories can be traced as follows: anchoring effect and recency bias derive directly from the original heuristics-and-biases framework; confirmation bias and overconfidence bias appear both in and in the recent medical-domain bias evaluation taxonomy; sunk cost fallacy and bandwagon effect were adapted from the broader behavioral-economics literature and operationalized specifically for the corporate-disclosure context in this work. The six categories were selected by two criteria: prevalence in management discussion sections (jointly about 86 percent of expert-flagged instances) and grounding in the established taxonomies cited above. Biases common in other settings but sparse in disclosure text were deliberately excluded; for example, status-quo bias, central to habitual decision-making, rarely surfaces in forward-looking strategic narratives.
This non-exhaustiveness is revisited in Section V-C.
Operational definitions specific to the context of corporate disclosures were derived through an iterative process involv-ing domain experts. Confirmation bias in annual reports tends to be characterized by the asymmetric presentation of data, where market research is prominently featured while risk factors are relegated to the fine print. Sunk cost fallacy tends to be exemplified by the use of the following words or phrases: ‘‘given our substantial R&D investment to date’’ or ‘‘considering resources already committed.’’ Overconfidence bias tends to be exemplified by the use of words such as ‘‘will achieve,’’ ‘‘guaranteed to succeed,’’ or claims to superiority without supporting quantitative evidence. Anchoring bias tends to be exemplified by references to the initial guidance or competitor levels established over the course of the preceding periods.
Bandwagon bias tends to be exemplified by the use of words such as ‘‘as leaders across our sector recognize’’ or ‘‘following proven models’’ to justify decisions. Recency bias tends to be exemplified by the use of words such as ‘‘based on last quarter’s momentum’’ to justify decisions over long periods. The prevalence and characteristic manifestations of the aforementioned biases are shown in Table 3, where it is evident that confirmation bias is the most dominant bias (28.3% of the training set), while recency bias is the least prevalent (6.7%), which may be due to the subtlety of its linguistic manifestations requiring temporal reasoning to detect.
The annotation process followed a two-stage hybrid approach. The first stage involved three business psychol-ogy experts with doctoral training in the field. All three experts independently annotated the same 300 randomly selected documents (150 English and 150 Chinese) using a standardized codebook, providing a complete three-rater intersection on which Fleiss’ kappa was computed. Eval-uations followed a single-blind protocol—analysts were unaware of each output’s source — with presentation order randomized per analyst.
The codebook contains, for each of the six bias categories: (i) a working definition; (ii) 6– 10 lexical and syntactic indicators (e.g., absolutist modal verbs, investment-anchor phrases) with positive and negative examples drawn from a 60-document pilot pool; (iii) decision rules for boundary cases; (iv) a five-point severity rubric; and (v) language-specific notes covering English and Chi-nese rhetorical conventions. The de-identified codebook is available from the corresponding author upon reasonable academic request, consistent with the Data Availability statement. The experts annotated the documents at the sentence level and provided one or more bias categories. Each expert also rated the bias severity using a five-point Likert scale.
Because three annotators participated, inter-annotator agreement was measured using Fleiss’ kappa, which is designed for three or more raters, and yielded κ = 0.74 (substantial agreement). For reference, the three pairwise Cohen’s kappa values were 0.78, 0.76, and 0.74 (mean = 0.76); these are reported only as supplementary information and not used as the primary aggregate metric. Disagreements were resolved by majority voting across the three annotators. The remaining 900 documents were annotated using a two-step LLM-assisted approach. GPT-4 (gpt-4-0613) first generated candidate sentence-level bias labels using six-shot prompts.
The prompt for each input sentence concatenated: (i) a system instruction describing the six-bias taxonomy and the JSON output schema ({biastype, span, confidence}); (ii) six in-context examples — one per bias category — randomly sampled from the stage-one expert-annotated pool, each pairing a positive sentence with its expert label and a one-line rationale; and (iii) the target sentence to be classified. Examples were re-sampled per document to reduce ordering bias, and decoding used temperature = 0 with a 256-token output cap. Two trained graduate research assistants (each having completed a 4-hour calibration session against the stage-one expert annotations and reaching ≥90% agreement on a 50-sentence calibration set before commencing verification) independently reviewed each GPT-4 candidate label, classifying it as accept, modify (with corrected bias category or span), or reject.
Disagreements between the two assistants were resolved in two tiers: cases where both assistants agreed were finalized directly; cases of inter-assistant disagreement, or any case flagged by either assistant as ambiguous, were escalated to the three stage-one doctoral experts, whose majority vote constituted the final label. Roughly 12% of candidate labels were escalated under this protocol. Of the candidate labels processed, 76.4% were accepted as-is, 15.1% had their bias category modified, 6.7% had their span boundaries adjusted, and 1.8% were rejected as false positives; verifiers additionally added bias instances missed by GPT-4, primarily for recency bias and Chinese-language sunk-cost cases.
Cost savings of approximately 60% were realized relative to full expert annotation, and a post-hoc re-evaluation of 100 randomly sampled stage-two documents against independent expert relabeling yielded Cohen’s kappa = 0.74, comparable to stage-one expert agree-ment (Fleiss’ kappa = 0.74; pairwise Cohen’s mean = 0.76).
The annotated corpus was split into training (720 doc-uments, 60%), validation (240 documents, 20%), and test (240 documents, 20%) sets using a stratified sampling approach. This split is an appropriate balance between the need for a large training set to cover linguistic diversity and the need for a sufficient number of documents in the other two sets for performance evaluation without overfitting. Document-level stratification used language (English/Chinese), industry sector (GICS), and bias-positive flag as joint stratification keys; per-category bias proportions across the three splits were 28.2/28.5/28.1% (confirma-tion), 18.7/18.5/18.9% (sunk cost), 22.1/22.3/21.8% (over-confidence), 12.4/12.2/12.6% (anchoring), 11.8/11.9/11.7% (bandwagon), and 6.7/6.6/6.9% (recency), confirming that proportions were preserved within ±0.4 percentage points.
Because the split was performed at the document level over a multi-year window (2023 to 2025), we additionally ran a company-identifier leakage check. Of the firms in the corpus, roughly 90 contributed documents to more than one split; the 24 test-set documents (10 percent of the test set) whose issuing firm also appeared in the training set were removed and the model re-evaluated. Test weighted F1 changed only marginally, from 0.82 to 0.815 (a drop of 0.5 percentage points), indicating that cross-split firm overlap does not materially inflate the reported performance.
C. COGDEBIAS FRAMEWORK ARCHITECTURE
Before describing the architecture, we formalize the detection task. The cognitive bias detection problem is formulated as a ‘‘multi-label sentence-level classification’’task. Let S = {s1, s2..., sM} denote the sentence collection across the corpus (M = 30,416). Each sentence s is mapped via the multilingual encoder into a 768-dimensional representation x ∈ R768, and is associated with a binary label vector y = [y1..., y6]T ∈ {0,1}6 indicating the presence of each of the six bias categories: confirmation bias, sunk cost fallacy, overconfidence, anchoring effect, bandwagon effect, and recency bias. A sentence may carry zero, one, or multiple bias labels simultaneously, reflecting the empirical finding that 14.2% of bias-positive sentences exhibit more than one bias type. The detection task seeks a hypothesis h:R768 →6 producing the per-category bias probability for each sentence.
In the realized system this hypothesis is instantiated as the weighted fusion of the embedding-based ML branch and the text-based LLM branch (Section III-C, Eq. ); the encoder representation therefore enters detection through the ML branch rather than as a standalone classifier. Encoder fine-tuning minimizes a combined objective L = LBCE + αLSupCon (α = 0.3), where LBCE is the binary cross-entropy averaged over the six categories and LSupCon is the supervised contrastive term encouraging category-consistent embedding clustering. With l2-normalized embeddings zi and temperature τ = 0.07, it is defined as