AI-Assisted Value Investing: A Human-in-the-Loop Framework for Prompt-Guided Financial Analysis and Decision Support
1 More Paper · Full Reading

About this paper
A full audio edition of this paper.
Authors: A. Caridi, M. Giovannini, L. Ricciardi Celsi
Publication date: 2026
Read the paper: https://doi.org/10.3390/electronics15061155
The authors and publisher do not sponsor or endorse this recording.
Source license: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/).
This audio adaptation adds an introduction and omits references and other narration distractions.
Transcript
You’re listening to “AI-Assisted Value Investing: A Human-in-the-Loop Framework for Prompt-Guided Financial Analysis and Decision Support,” by A. Caridi, M. Giovannini, and L. Ricciardi Celsi. Published in 2026.
Abstract.
Value investing remains grounded in intrinsic value estimation, margin-of-safety reasoning, and disciplined fundamental analysis, but its practical execution is increasingly constrained by the scale, heterogeneity, and velocity of modern financial information. Recent advances in artificial intelligence (AI), particularly large language models and automated information-extraction systems, create new opportunities to accelerate financial analysis; however, their outputs remain probabilistic, context-dependent, and potentially error-prone, making gov-ernance and verification essential. This article proposes an AI-assisted value investing framework that integrates automated extraction, valuation modeling, explainability, and human-in-the-loop (HITL) supervision into a unified decision-support architecture.
The framework is organized into three layers: (i) a data layer for traceable extraction and normalization of structured and unstructured financial information; (ii) a modeling layer for automated key performance indicator (KPI) computation, forecasting support, and discounted cash flow (DCF) valuation; and (iii) an explainability and governance layer for traceability, verification, model-risk control, and analyst oversight. A central contribution of the paper is the operational characterization of prompt literacy as a determinant of analytical reliability, showing that structured, context-aware prompts materially affect ex-traction correctness, usability, and verification effort.
The framework is evaluated through a case study using Rivanna AI on three large U.S. beverage firms—namely, The Coca-Cola Company, PepsiCo, and Keurig Dr Pepper—selected as a controlled, information-rich set-ting for comparative analysis. The results indicate that the proposed workflow can reduce end-to-end analysis time from approximately 25–40 h in a traditional manual process to approximately 8–12 h in an AI-assisted setting, including citation/source verification, unit and period reconciliation, and review of key valuation assumptions. Rather than elimi-nating analyst effort, AI shifts it from manual information processing toward verification, adjudication, and interpretation. Overall, the findings position AI not as an autonomous decision-maker, but as a governed reasoning accelerator whose effectiveness depends on structured human guidance, traceability, and disciplined validation.
In value investing, a discipline traditionally grounded in labor-intensive fundamental analysis and disciplined intrinsic value estimation, AI introduces the potential to scale analytical coverage and accelerate evidence synthesis. However, AI systems in financial contexts are probabilistic, context-sensitive, and inherently dependent on human interaction, raising critical questions about reliability, governance, and operational integration. This article proposes a structured framework for AI-driven value investing that preserves the foundational principles of intrinsic value, margin of safety, and economic reasoning, while redesigning the analytical workflow through automation, explainability, and human-in-the-loop (HITL) supervision. The proposed architecture integrates three layers: (i) an AI-enabled data layer for traceable
Academic Editors: Fernando De la Prieta Pintado, Valentina Emilia Balas, Xin Geng and George A. Papakostas
extraction and normalization of structured and unstructured financial information; (ii) a modeling and valuation layer combining automated KPI computation, machine learning forecasting, and discounted cash flow (DCF) valuation; and (iii) an explainability and governance layer ensuring traceability, verification, and model risk control. A central contribution of this work is the operational characterization of prompt literacy, namely the ability to formulate structured, context-aware requests to AI systems, as a critical determi-nant of system reliability and analytical correctness. Through a focused case study using an AI-assisted analysis platform (Rivanna AI) on three U.S. beverage firms, we provide evidence that structured prompt formulation can improve extraction consistency, reduce verification overhead, and increase workflow efficiency in a human-supervised setting.
In this setting, analysis time decreased from a manual range of approximately 25–40 h to 8–12 h with AI assistance and HITL validation, while preserving traceability and decision accountability. The reported hour savings should be interpreted as conservative estimates from the initial deployment phase; additional efficiency gains are expected as operational maturity increases, driven by learning-economy effects. The findings position AI not as an autonomous decision-maker but as a probabilistic reasoning accelerator whose effective-ness depends on structured human guidance, verification discipline, and prompt-driven interaction. These results redefine the role of the financial analyst from manual data proces-sor to reasoning architect, responsible for designing, guiding, and validating AI-assisted analytical workflows.
1. Introduction.
Value investing, formalized by Graham and Dodd and further developed by subse-quent practitioners and scholars, is grounded in a fundamental principle: financial securities should be evaluated based on intrinsic economic value rather than market price alone. This approach requires integrating quantitative financial analysis, qualitative business assess-ment, and forward-looking valuation models, traditionally through labor-intensive manual workflows involving financial statement analysis, ratio computation, and discounted cash flow (DCF) modeling.
While this methodology remains conceptually robust, its practical execution is increas-ingly challenged by structural changes in the financial information environment. Modern analysts must process vast and heterogeneous datasets, including regulatory filings, earn-ings call transcripts, alternative data sources, and macroeconomic signals, under time constraints that limit analytical depth and coverage. These constraints create operational bottlenecks and introduce cognitive risks, including confirmation bias, anchoring, and selective information processing.
Recent advances in artificial intelligence (AI), particularly in natural language process-ing, automated information extraction, and generative reasoning models, offer the potential to fundamentally reshape financial analysis workflows. AI systems can rapidly extract structured data from unstructured sources, compute financial metrics at scale, synthesize qualitative information, and assist in valuation modeling. However, these systems are not deterministic computational tools. Their outputs are probabilistic and depend on context, prompt formulation, and human interaction.
This creates an important operational challenge. The reliability of AI outputs depends not only on model architecture and training data but also on how clearly and completely the analyst formulates the request. In practical terms, prompt quality directly affects extraction accuracy, logical consistency, and analytical usability. For this reason, prompt literacy becomes a relevant operational skill in AI-assisted financial analysis. Poorly specified prompts can produce incomplete or inconsistent outputs, while structured and context-rich prompts improve reliability and efficiency in the considered case study.
This observation motivates a conceptual shift in how AI should be understood within financial analysis. Rather than replacing human analysts, AI systems function as cogni-tive amplifiers that extend analytical capacity while remaining dependent on structured human guidance and verification. In this augmented intelligence paradigm, the ana-lyst’s role evolves from manual data processor to reasoning architect, designing analytical workflows, structuring prompts, validating outputs, and integrating results into coherent investment decisions.
This article proposes a reference framework for AI-driven value investing that inte-grates automated extraction, valuation modeling, explainability, and human-in-the-loop (HITL) governance into a unified operational architecture. The framework explicitly in-corporates prompt literacy as a core operational competence, recognizing that effective interaction with AI systems is a prerequisite for reliable analytical outcomes. This way, we address the engineering of reliable AI-enabled decision-support systems, which is a core problem lying at the intersection between modern electronic and computer-based systems, on the one hand, and finance theory, on the other hand.
We propose a modular, auditable architecture for integrating large language models into a real-world workflow, emphasizing system-level requirements such as traceability of numerical outputs to primary sources, schema-constrained extraction, cross-model validation, and human-in-the-loop governance gates to manage model risk and operational drift. By framing LLMs as components within a verifiable pipeline (with explicit failure modes, monitoring, and escalation rules), the work contributes practical design patterns for dependable AI deployment in software-intensive electronic systems where correctness, accountability, and reproducibility are critical.
We provide evidence of this framework through a case study using an AI-assisted financial analysis platform (Rivanna AI), applied to comparative valuation of major US beverage companies. The case study evaluates extraction correctness, valuation consistency, and productivity improvements relative to traditional manual workflows while analyzing the operational role of structured prompt formulation in improving system reliability.
The main contributions of this work are fourfold:
• a structured architecture for AI-augmented value investing integrating extraction, modeling, explainability, and governance;
• a formalization of human-in-the-loop supervision as an operational control mechanism ensuring traceability and decision accountability;
• an operational characterization of prompt literacy as a determinant of AI system reliability and analytical correctness;
• an empirical evaluation demonstrating significant productivity gains and reliability improvements in AI-assisted financial analysis.
By linking AI performance to structured human interaction and prompt design, this work clarifies how AI can be integrated more safely and effectively into financial decision-support systems. More broadly, it suggests that future competitive advantage in AI-augmented investing will depend not only on access to data or computational resources but also on the ability to interact with AI systems in a structured, verifiable, and economically grounded manner.
The remainder of this paper is organized as follows. Section 2 reviews the related literature on AI-assisted financial analysis, decision-support systems, explainable AI, and generative models and identifies the research gap addressed by this work. Section 3 revisits the fundamental principles of value investing and analyzes the operational and cognitive limitations of traditional analytical workflows. Section 4 introduces the proposed AI-augmented value investing architecture, describing its design objectives and the three-layer system structure integrating data extraction, valuation modeling, and explainability and governance. Section 5 details the data layer, focusing on automated financial information extraction, normalization, and integration into analyst workflows.
Section 6 presents the modeling layer, including automated KPI computation, intrinsic value estimation via dis-counted cash flow, and NLP-assisted qualitative analysis. Section 7 discusses explainability and governance mechanisms, including traceability, model risk management, and the role of explainable AI in ensuring analytical reliability. Section 8 formalizes the human-in-the-loop operating model and introduces prompt literacy as a core operational competence enabling reliable human–AI interaction. Section 9 presents the case study using the Rivanna AI platform, including experimental design, validation metrics, and empirical results on extraction correctness, valuation consistency, and productivity gains. Finally, Section 10 concludes the paper by summarizing the main contributions, discussing limitations, and outlining directions for future research.
To improve readability and terminological consistency, the principal acronyms used throughout the manuscript are consolidated here: Artificial Intelligence (AI), Capital Asset Pricing Model (CAPM), Chain-of-Thought (CoT), Discounted Cash Flow (DCF), Deci-sion Support System (DSS), Extract–Load–Transform (ELT), Equity Risk Premium (ERP), Extract–Transform–Load (ETL), Fiscal Year (FY), Generally Accepted Accounting Princi-ples (GAAP), Human-in-the-Loop (HITL), International Financial Reporting Standards (IFRS), Key Performance Indicator (KPI), Large Language Model (LLM), Local Interpretable Model-Agnostic Explanations (LIME), Management’s Discussion and Analysis (MD&A), Margin of Safety (MoS), Natural Language Processing (NLP), Retrieval-Augmented Gener-ation (RAG), Relative Error (RE), Return on Invested Capital (ROIC), U.S.
Securities and Exchange Commission (SEC), Shapley Additive Explanations (SHAP), Terminal Value (TV), User Interface (UI), Weighted Average Cost of Capital (WACC), and Explainable Artificial Intelligence (XAI). Company tickers are The Coca-Cola Company (KO), PepsiCo, Inc. (PEP), and Keurig Dr Pepper Inc. (KDP).
2. Related Works.
The integration of artificial intelligence (AI) into financial analysis has attracted increas-ing attention across multiple research domains, including financial forecasting, automated information extraction, decision-support systems, and explainable AI. This section situates the present work within the existing literature and highlights the specific contribution of our approach.
2.1. AI for Financial Analysis and Valuation.
The application of computational methods to financial analysis has evolved signifi-cantly over the past decades. Classical approaches relied primarily on statistical models for valuation and forecasting, including regression-based discounted cash flow (DCF) modeling and factor-based valuation methods. These frameworks remain foundational to value investing, which emphasizes intrinsic value estimation and margin-of-safety reasoning.
More recently, machine learning techniques have been applied to financial predic-tion and classification tasks, including bankruptcy prediction, earnings forecasting, and asset pricing. Deep learning methods have further expanded these capabilities, en-abling automated extraction and interpretation of financial information from large and heterogeneous datasets.
Natural language processing (NLP) has played a particularly important role in fi-nancial analysis. Transformer-based architectures such as BERT and related models have demonstrated strong performance in extracting structured financial information from unstructured text, including regulatory filings, earnings transcripts, and financial disclo-sures. Sentiment analysis and textual embeddings have been shown to improve predictive accuracy in financial forecasting and risk assessment tasks.
However, most existing studies focus on predictive modeling rather than supporting the full analytical workflow required in value investing, which involves data extraction, fi-nancial normalization, ratio computation, valuation modeling, and strategic interpretation.
2.2. Decision-Support Systems and Augmented Intelligence.
Decision-support systems (DSS) have long been used to assist complex analytical and strategic decision-making in finance and management. Classical DSS archi-tectures integrate data management, analytical models, and user interfaces to support structured reasoning.
With the emergence of AI, these systems have evolved into AI-augmented decision-support platforms capable of automating data extraction, synthesis, and analysis. Rather than replacing human decision-makers, these systems operate under the paradigm of augmented intelligence, where AI enhances human cognitive capacity while preserving human supervision and responsibility.
Human-in-the-loop (HITL) systems have emerged as a critical design paradigm for ensuring reliability and accountability in AI-assisted workflows. HITL architectures incorporate human verification and oversight at key stages of the analytical process, allow-ing analysts to validate extracted information, correct errors, and ensure consistency with domain knowledge.
This paradigm is particularly relevant in financial analysis, where decisions require justification, traceability, and accountability.
2.3. Explainable AI and Model Governance in Finance.
Explainability has become a fundamental requirement for AI systems operating in high-stakes domains such as finance. Explainable AI (XAI) techniques enable users to interpret model outputs and understand the underlying reasoning processes.
Several methods have been proposed to improve explainability, including feature attribution methods such as SHAP (Shapley Additive Explanations) and LIME (Local Interpretable Model-Agnostic Explanations). These techniques provide insights into the relative importance of input variables and allow analysts to validate model behavior.
Explainability is particularly important in financial decision-making, where invest-ment decisions must be justified to stakeholders and regulatory bodies. Transparent AI systems improve trust, facilitate auditing, and reduce operational risk.
However, most existing explainability research focuses on predictive models rather than generative AI systems used for analytical synthesis and reasoning.
2.4. Generative AI and Large Language Models in Financial Workflows.
The emergence of large language models (LLMs) has significantly expanded the scope of AI-assisted analytical workflows. Transformer-based architectures introduced by Vaswani et al. and subsequent models such as GPT and BERT have enabled auto-mated document understanding, structured data extraction, and reasoning over complex textual information.
These models have demonstrated strong performance in financial document analysis, automated report generation, and information extraction tasks.
However, generative AI systems introduce new operational challenges, including hallucinations, sensitivity to prompt formulation, and probabilistic output variability. Unlike deterministic software systems, LLM outputs depend not only on input data but also on the structure and clarity of the prompt provided by the user.
Recent research in human–AI interaction suggests that structured prompts and task de-composition significantly improve system reliability and output quality. Nevertheless, this aspect has not been formally integrated into financial decision-support frameworks.
2.5. Research Gap and Contribution of This Work.
Despite significant progress in AI-assisted financial analysis, three important gaps remain.
• First, most existing work treats AI systems as autonomous analytical tools, rather than interactive systems whose reliability depends on structured human guidance.
• Second, the role of prompt formulation as a determinant of system reliability has not been formally characterized in financial analysis workflows.
• Third, there is limited empirical evidence linking human–AI interaction design to measurable improvements in analytical efficiency and reliability.
This work addresses these gaps by proposing a unified framework for AI-driven value investing that integrates automated extraction, valuation modeling, explainability, and human-in-the-loop governance, while explicitly incorporating prompt literacy as a critical operational competence.
Through an empirical case study using an AI-assisted financial analysis platform, we provide evidence that structured prompt formulation improves extraction correctness, reduces verification overhead, and enhances analytical efficiency.
This contribution advances the literature on augmented intelligence and AI-assisted financial decision support by framing financial analysis as a collaborative human–AI reasoning process rather than a fully automated pipeline.
3. Value Investing Principles and Workflow Limitations.
Value investing is grounded in the principle that financial securities represent owner-ship claims on real economic assets and should therefore be evaluated based on intrinsic value rather than short-term market fluctuations. Intrinsic value represents an estimate of the present value of future cash flows, adjusted for risk and capital structure. Because intrinsic value estimation is inherently uncertain, value investors rely on the concept of margin of safety, defined as the difference between estimated intrinsic value and market price, which provides protection against estimation errors and adverse developments.
A central conceptual framework supporting this approach is the distinction between price and value. Market prices reflect collective expectations, sentiment, and behavioral factors, whereas intrinsic value reflects underlying economic fundamentals such as prof-itability, reinvestment capacity, and competitive positioning. The role of the analyst is therefore to independently estimate intrinsic value and exploit deviations between price and value when sufficient margin of safety exists.
While these principles remain conceptually unchanged, their practical implementa-tion has become increasingly complex due to the rapid growth in financial data volume, heterogeneity of information sources, and increased reporting complexity.
Traditional Financial Analysis Workflow and Structural Bottlenecks
The traditional value investing workflow consists of five sequential stages: data ac-quisition, data normalization, financial ratio computation, qualitative business analysis, and valuation modeling. Analysts collect financial statements, regulatory filings, earnings transcripts, and market data, then normalize and integrate this information to compute performance indicators such as profitability, leverage, and valuation multiples. This infor-mation is subsequently used to build valuation models, typically based on discounted cash flow (DCF) analysis.
This workflow faces two fundamental limitations.
• First, operational constraints limit scalability. Data acquisition and normalization remain labor-intensive, particularly when integrating heterogeneous sources such as structured financial statements and unstructured textual disclosures. Differences in accounting standards, reporting conventions, and fiscal calendars further increase com-plexity.
• Second, cognitive constraints affect analytical reliability. Human analysts are subject to cognitive biases such as anchoring, confirmation bias, and selective attention, which may distort interpretation and decision-making.
These limitations create a trade-off between analytical depth and coverage, motivating the development of AI-assisted systems capable of automating repetitive tasks while preserving interpretability and analytical rigor.
4. AI-Augmented Value Investing System Architecture and.
Design Objectives
4.1. Design Objectives.
An AI-augmented financial analysis system must satisfy several functional and opera-tional requirements to be suitable for real-world decision support.
First, scalability is required to enable analysis of large numbers of firms without proportional increases in analyst effort. Second, evidence integrity must be preserved, ensuring that extracted information remains traceable to original sources. Third, model pluralism is necessary to integrate multiple analytical methods, including fundamental valuation models, statistical forecasting, and textual analysis.
Fourth, explainability and governance mechanisms must ensure that analytical out-puts remain interpretable, verifiable, and auditable. Finally, decision orientation must be maintained, ensuring that analytical outputs support actionable strategic decisions rather than purely predictive outputs.
These requirements motivate a modular system architecture combining automated extraction, valuation modeling, explainability, and human supervision.
4.2. Three-Layer System Architecture.
The proposed architecture consists of three functional layers, as shown in Figure 1: data extraction, analytical modeling, and explainability and governance. The end-to-end workflow is summarized in Algorithm 1, consistent with the three-layer architecture shown in Figure 1.
The data layer is responsible for extracting structured and unstructured financial information from heterogeneous sources, including financial statements, regulatory filings, earnings transcripts, and market data. Automated extraction techniques based on natural language processing and structured parsing enable scalable and consistent data acquisition.
The modeling layer transforms extracted data into economically meaningful analytical outputs. This includes automated computation of financial ratios, machine learning–based forecasting of key financial drivers, and intrinsic value estimation using discounted cash flow and relative valuation models.
The explainability and governance layer ensures that analytical outputs remain in-terpretable and verifiable. This layer provides traceability to original data sources, model interpretability mechanisms, and verification checkpoints enabling human oversight.
This layered architecture enables scalable automation while preserving analytical transparency and decision accountability.
Algorithm 1 Prompt-Driven AI Value Investing Workflow with HITL Gates
Input: Analyst goal, company C, period P, constraints (units/currency, required KPIs, valuation method), source set S.
Output: Decision-support package (sheets/memo), intrinsic value band, risk flags, full traceability log.
1. Specify task & constraints (UI): The analyst defines the objective, scope (C, P).
required outputs (KPIs/DCF/scenarios), and constraints (units, currency, disclosure priority).
2. Orchestrate workflow (Orchestrator): Load templates, route subtasks, and.
initialize logging/audit trail for prompts, sources, and outputs.
3. Layer A—Data & extraction (ETL/ELT): Retrieve sources S; extract numeric fields.
with schema-constrained prompts; map entities; normalize units/scale; align fiscal periods; attach citations.
4. G1—Data verification gate (HITL): Run consistency checks (units, missingness.
basic accounting ties). If failures occur, re-extract with tighter constraints or escalate to analyst correction; log decisions.
5. Layer B—Modeling & valuation: Compute KPIs; build valuation drivers; run.
DCF with scenarios and sensitivity; produce relative valuation checks; generate draft synthesis.
6. G2–G3—KPI & valuation gates (HITL): Validate KPI plausibility and valuation.
assumptions (bounds, driver sensitivity). Resolve discrepancies via source check/re-run; record adjudications.
7. Layer C—Explainability and governance: Generate traceability/lineage report.
uncertainty notes, XAI summaries, and drift/anomaly/bias risk flags; consolidate audit logs.
8. G4 and output delivery: Let the analyst validate the narrative against sources.
then allow them to export nonautonomous decision-support outputs (Excel-ready tables, memo, intrinsic value band, rationale, and risk flags).
5. Data Layer: Automated Financial Information Extraction.
5.1. Automated Extraction and Normalization.
The data layer is responsible for transforming heterogeneous financial information into structured, analysis-ready datasets. Modern AI techniques enable automated extraction of key financial variables from regulatory filings, earnings reports, and other disclosures. These methods support entity recognition, numerical extraction, and contextual interpreta-tion of financial information.
Normalization mechanisms align extracted data across reporting periods, account-ing conventions, and organizational structures. This process supports consistency and comparability across firms and time.
A critical requirement is traceability: each extracted data point must remain linked to its original source. This enables analysts to verify correctness and supports auditability of analytical outputs.
5.2. Integration with Analyst Workflows.
For practical adoption, automated extraction systems must integrate seamlessly with existing analyst workflows. Spreadsheet-based environments remain the dominant analyti-cal interface in finance. Automated systems can generate structured tables, standardized financial metrics, and valuation-ready datasets directly within these environments.
This integration reduces manual data manipulation and allows analysts to focus on higher-level analytical tasks, including interpretation, validation, and strategic reasoning.
6. Modeling Layer: Financial Analysis and Valuation.
6.1. Automated Financial Metrics and Consistency Validation.
Once financial data is structured, automated systems compute financial metrics includ-ing profitability ratios, leverage indicators, liquidity measures, and valuation multiples. Automated computation supports consistency and scalability.
Validation mechanisms are required to detect anomalies such as unit mismatches, missing values, and structural inconsistencies. These mechanisms improve reliability and reduce error propagation into valuation models.
6.2. Forecasting and Intrinsic Value Estimation.
Intrinsic value estimation requires forecasting future cash flows and discounting them using appropriate risk-adjusted discount rates. AI-based forecasting models can assist in estimating revenue growth, operating margins, and capital requirements.
These models should be treated as decision-support tools rather than autonomous predictors. Human analysts remain responsible for validating assumptions and ensuring consistency with economic fundamentals.
Intrinsic value estimation remains anchored in economic logic, with AI serving to improve scalability and analytical efficiency.
6.3. Natural Language Processing for Qualitative Analysis.
Qualitative information plays a critical role in financial analysis. NLP techniques enable automated analysis of earnings transcripts, management communication, and financial disclosures.
These methods extract sentiment indicators, detect linguistic uncertainty, and identify changes in managerial tone or strategic direction. These signals complement quantitative analysis and support more comprehensive evaluation of firm fundamentals.
6.4. Generative AI for Analytical Synthesis.
Generative AI systems assist analysts by synthesizing financial information, summa-rizing disclosures, and structuring analytical reports. These systems improve efficiency but introduce risks related to hallucinations and unverifiable outputs.
To mitigate these risks, generative outputs must remain traceable to source data and subject to human validation.
7. Explainability and Governance.
7.1. Explainability Requirements in Financial Decision Support.
Explainability is a fundamental requirement for AI-assisted financial decision support systems due to the economic, operational, and regulatory implications of investment decisions.
Unlike purely predictive applications, where performance metrics alone may suffice, financial analysis requires that conclusions be interpretable, verifiable, and economically justified.
Investment decisions often undergo multiple layers of review, including analyst val-idation, investment committee approval, audit processes, and regulatory scrutiny. In these contexts, analytical outputs must be accompanied by clear explanations that enable decision-makers to understand how conclusions were derived and which assumptions and data inputs influenced the result.
Explainability therefore serves several critical functions. First, it enables correctness verification, allowing analysts to confirm that extracted data, computed metrics, and valuation outputs are logically and numerically consistent. Second, it supports error detection by enabling the identification of incorrect assumptions, data inconsistencies, or model failures. Third, it improves trust and usability, allowing analysts to interpret AI-generated outputs in terms of economically meaningful variables such as profitability, growth, reinvestment, and capital structure.
From a regulatory and governance perspective, explainability is also essential for auditability and accountability. Financial decision processes must maintain documented reasoning chains linking conclusions to underlying evidence. AI-assisted systems must therefore provide transparent analytical pipelines that preserve traceability from raw data to final analytical outputs.
In this context, explainability is not merely a desirable feature but a structural require-ment for safe and reliable deployment of AI in financial analysis workflows.
7.2. Explainability Mechanisms.
Explainability in AI-assisted financial analysis can be implemented through a combi-nation of structural, analytical, and model-level mechanisms that ensure interpretability, traceability, and robustness.
A primary requirement is data lineage and traceability, ensuring that each extracted data point remains linked to its original source. This enables analysts to verify correctness by inspecting the originating document, such as a regulatory filing or earnings report. Traceability mechanisms typically include source references, document identifiers, and extraction metadata.
Feature attribution mechanisms provide insight into which input variables contribute most strongly to analytical outputs. In valuation models, these mechanisms allow analysts to identify key drivers such as revenue growth, operating margins, or discount rates that influence intrinsic value estimates. This improves interpretability and allows validation of economic plausibility.
Sensitivity analysis provides an additional layer of explainability by quantifying how analytical outputs change in response to variations in key assumptions. For example, intrinsic value estimates can be evaluated under alternative discount rate and growth assumptions, producing valuation ranges rather than single-point estimates. This ap-proach reflects the inherent uncertainty of financial forecasting and improves robustness of decision support.
Uncertainty quantification mechanisms further enhance interpretability by providing confidence intervals or scenario-based output ranges. Rather than presenting deterministic predictions, explainable systems communicate uncertainty explicitly, enabling analysts to assess risk and robustness.
In generative AI components, explainability also requires source-grounded genera-tion. Outputs such as summaries or analytical narratives must remain linked to source documents, preventing unverifiable or hallucinated conclusions. This supports the fact that narrative outputs remain evidence-based and suitable for decision support.
Collectively, these mechanisms ensure that AI-assisted financial analysis remains interpretable, verifiable, and aligned with fundamental economic reasoning.
7.3. Model Risk and Governance.
LLM-enabled decision-support for value investing introduces a layered risk surface across (i) source/provenance, (ii) extraction and normalization, (iii) KPI and valuation modeling, (iv) generative narrative, and (v) operations and lifecycle drift. We therefore treat governance as a system-engineering assurance layer that enforces traceability, measur-able verification gates, and documented accountability across the pipeline (Algorithm 1; Figure 1), consistent with AI risk-management and lifecycle monitoring practices.
Data and source risk. Failures arise from incomplete or stale disclosures, restate- ments, inconsistent accounting conventions, and mixing audited filings with lower-authority sources (decks/transcripts). The framework enforces a source-of-truth hierarchy (audited/regulatory first), timestamps all inputs, and applies complete-ness checks (coverage by year/statement/notes). When non-filing sources are used, items are explicitly labeled and excluded from “hard-number” computations unless independently validated. Documentation artifacts further improve auditability and reproducibility.
Extraction and normalization risk. Common failure modes include schema drift, unit/scale errors (thousands vs. millions), fiscal-period misalignment, and synonym collisions in line items. Mitigation combines (i) schema-constrained prompts (ta-ble/JSON), (ii) mandatory reporting of units/currency/period mapping, (iii) per-number citations, and (iv) deterministic checks (range/sign constraints and cross-statement reconciliation). Any missing citation or unit ambiguity is treated as a hard failure that triggers HITL Gate G1.
KPI and valuation model risk. Even with correct inputs, valuation uncertainty is dominated by assumptions (growth, margins, reinvestment, discount rate, terminal value). To avoid overconfident point estimates, the system produces intrinsic value as a band via scenario and sensitivity grids, enforces explicit parameter bounds (e.g., terminal growth constraints), and requires analyst justification for the highest-leverage drivers at HITL Gate G3. This reflects established valuation practice that emphasizes sensitivity to key drivers and conservative decision rules (margin of safety).
Generative and reasoning risk. LLMs may produce fluent but unsupported claims (“hallucinations”), particularly in qualitative synthesis (moat, strategy, competitive dynamics) or when extrapolating beyond evidence. The framework separates num-bers (must be cited and checked) from narratives (must be citation-grounded or labeled as hypotheses). Retrieval-grounded generation (retrieve-then-write) reduces unsupported statements, while structured prompting and consistency checks reduce numerical extraction errors.
Operational governance risk and verifier trust boundary. System behavior can drift with model updates, prompt/template changes, and data shifts. Moreover, verification components can become single points of failure if treated as infallible. We therefore treat the verifier as a fallible component: every run produces an audit trail (prompts, outputs, citations, adjudications, sign-offs), critical components are version-pinned, and monitoring tracks error rates, disagreement frequency, and regression failures over time. Disagreements escalate via a documented protocol (from source check, through re-prompt, to human adjudication), and final accountability remains with HITL sign-off at G2–G4.
Table 1 below sums up model risks, failure modes, detection signals, control, and ownership according to the five points discussed above.
7.4. Disagreement Resolution and Cross-Model Validation Protocol.
To improve robustness and reduce model-specific failure modes, the proposed frame-work includes a cross-model validation stage in which the same extraction task is indepen-dently executed by two large language models (e.g., GPT-4 and Claude). The objective is not to assume that agreement implies correctness but to use model diversity as a practical mechanism for detecting unstable outputs, unsupported values, and source inconsistencies before these propagate to KPI computation and valuation.
The protocol begins with parallel extraction. Given the same company, reporting period, and source set, both models receive an equivalent structured prompt requesting the extraction of predefined financial line items together with units, currency, fiscal-period mapping, and exact source citations. The outputs are then passed to a normalization layer, which harmonizes numerical scale (e.g., thousands vs. millions), currency denomina-tion, and fiscal-period labeling and maps synonymous labels to a common schema. This step is necessary because nominal disagreements may arise from formatting or reporting conventions rather than from true content divergence.
After normalization, the system applies a disagreement detection rule. A field is flagged if at least one of the following conditions holds: (i) the relative deviation between the two extracted values exceeds a predefined threshold; (ii) the cited source passages differ materially or one output lacks a citation; (iii) the reported units, currency, or fiscal-period mapping are inconsistent; or (iv) one model returns a value while the other returns a missing field or a qualitatively different interpretation. Flagged items are collected into a disagreement set and routed to an adjudication procedure.
The adjudication hierarchy follows three ordered steps. First, the system checks the primary filing or audited report as the source of truth. If the discrepancy can be resolved directly from the filing, the verified value is accepted and logged. Second, if the source text is ambiguous or the extraction failed because of formatting complexity, the task is re-run with a tighter prompt that imposes stricter schema constraints, explicit unit handling, and a narrower extraction scope. Third, if disagreement persists, the item is escalated to a human analyst, who makes the final determination and records the rationale for the decision. In this way, unresolved conflicts are not silently absorbed into downstream computations.
A key component of this protocol is the audit trail. For each disputed item, the system stores the original prompts, both model outputs, normalized values, cited passages, discrepancy type, re-prompt attempts, and the final adjudicated value. This provides full traceability for later review and supports reproducibility, accountability, and post hoc error analysis. The protocol therefore operationalizes governance not as a generic oversight principle but as a concrete mechanism for handling uncertainty, disagreement, and model fallibility in AI-assisted financial analysis.
It is important to note that, beyond error detection on quantitative fields, using mul-tiple LLMs enables what we term a scenario corridor for qualitative and strategic drivers. For numerical extraction, the goal of cross-model validation is convergence to a single source-grounded value after normalization and citation checks. In contrast, for qualita-tive judgments (e.g., competitive positioning, moat sustainability, management execution, regulatory exposure), the objective is not strict convergence to one “true” statement but the identification of a bounded set of plausible outcomes. Agreement across independent models on the direction and range of qualitative conclusions provides a more robust rep-resentation of uncertainty than relying on a single model’s narrative.
In the proposed governance protocol, this corridor is treated as an input to the analyst’s decision-making: convergent qualitative conclusions increase confidence, while divergent outputs trigger targeted evidence retrieval and human adjudication before final sign-off.
8. Human-in-the-Loop Operating Model.
8.1. Role of Human Supervision.
Human-in-the-loop (HITL) supervision represents a foundational component of AI-augmented financial analysis systems. Rather than fully automating analytical workflows, HITL architectures integrate human verification and oversight at critical stages of the analytical pipeline, ensuring correctness, interpretability, and accountability.
Financial analysis involves high-stakes decisions where small numerical or logical errors may propagate into significant valuation inaccuracies. Automated extraction, model-ing, and generative reasoning systems operate probabilistically and may introduce errors related to data extraction, model assumptions, or reasoning inconsistencies. HITL super-vision mitigates these risks by incorporating structured human validation at predefined verification gates.
Human supervision performs several essential functions. First, analysts validate ex-tracted data by checking unit consistency, accounting identity relationships, and source traceability. Second, analysts review model outputs to ensure economic plausibility and consistency with business fundamentals. Third, analysts verify generative outputs, ensur-ing that summaries, analytical narratives, and valuation conclusions remain grounded in verifiable evidence.
From a systems perspective, HITL supervision acts as a reliability layer. Its purpose is to detect and correct errors before they affect downstream analytical outputs. This approach enables scalable automation while preserving decision integrity and accountability.
HITL supervision does not eliminate efficiency gains. Instead, it shifts analyst effort away from repetitive tasks and toward validation, adjudication, and interpretation. This results in a system that combines the scalability of automation with the reliability of expert human judgment.
8.2. The Augmented Analyst: From Manual Processing to Strategic Oversight.
The integration of AI systems fundamentally transforms the role of financial analysts. In traditional workflows, analysts devote substantial time to manual data collection, nor-malization, spreadsheet manipulation, and mechanical ratio computation. These tasks are necessary but do not directly contribute to strategic insight or decision quality.
AI-assisted systems automate these operational tasks, allowing analysts to focus on higher-level cognitive functions. The role of the analyst evolves from manual processor to strategic supervisor of analytical workflows. Analysts now focus on defining analytical ob-jectives, validating model outputs, interpreting valuation results, and integrating analytical conclusions into coherent investment narratives.
This transformation introduces a shift from procedural execution to supervisory control. Analysts no longer perform every analytical operation manually but instead oversee automated processes, ensuring correctness, consistency, and economic validity.
The augmented analyst performs several key supervisory functions:
• defining analytical objectives and problem formulation;
• validating extracted data and resolving inconsistencies;
• evaluating model assumptions and forecasting plausibility;
• interpreting valuation outputs and sensitivity analysis;
• identifying anomalies and potential model failures;
• integrating quantitative and qualitative information into investment decisions.
This supervisory role preserves human responsibility and accountability while lever-aging the computational capabilities of AI systems.
Importantly, AI systems do not replace financial expertise but amplify it. Economic reasoning, domain knowledge, and judgment remain essential for interpreting model outputs and ensuring analytical correctness.
8.3. Prompt Literacy as a Core Operational Competence.
A critical skill emerging in AI-augmented financial analysis is prompt literacy, defined as the ability to formulate structured, precise, and context-aware instructions that guide AI systems toward reliable, verifiable, and decision-relevant outputs.
Unlike deterministic software, generative AI models respond probabilistically to the input they receive. The prompt therefore acts as an operational control signal. Its formulation influences extraction accuracy, analytical consistency, and interpretability.
Poorly specified prompts may lead to incomplete extraction, incorrect assumptions, logical inconsistencies, or hallucinated outputs. These failures increase verification over-head and reduce system reliability. Conversely, structured and precise prompts improve correctness, reproducibility, and analytical efficiency.
From an operational standpoint, effective prompts incorporate four essential components. First, explicit task definition clearly specifies the analytical objective, such as extracting specific financial variables, computing defined metrics, or performing structured valuation analysis. Precise task specification reduces ambiguity and improves extraction accuracy.
Second, contextual constraints provide necessary background information, including company identity, reporting period, accounting standards, and unit conventions. Context reduces uncertainty and ensures consistency with analytical requirements.
Third, structured output specification defines the required output format, such as tabular data, spreadsheet-compatible values, or predefined analytical structures. Structured outputs improve interpretability and facilitate downstream processing.
Fourth, verification orientation explicitly requires source traceability, citations, and uncertainty indicators. This ensures that outputs remain verifiable and suitable for financial decision support.
In addition to prompt structure, task decomposition plays an important role in improv-ing system reliability. Complex analytical tasks can be decomposed into smaller, sequential subtasks, allowing incremental validation and reducing error propagation. This approach improves robustness and enables systematic verification of intermediate outputs.
From a systems perspective, prompt formulation is an interaction-level control mecha-nism. Prompt literacy can therefore be viewed as an operational competence, similar in importance to model calibration or parameter tuning in traditional analytical systems.
This insight introduces an important conceptual shift: the reliability of AI-assisted financial analysis depends not only on model architecture and training data but also on the quality of human–AI interaction. The analyst actively shapes the reasoning trajectory of the AI system through structured prompt design.
In this sense, prompt literacy becomes a functional component of analytical rigor, ensuring that AI-assisted systems operate as reliable decision-support tools rather than autonomous decision-makers.
8.4. Closed-Loop Human–AI Interaction and Reliability Control.
The HITL operating model establishes a closed-loop interaction between human analysts and automated analytical systems. In this loop, AI systems generate structured analytical outputs, which are then validated, refined, and corrected by human supervisors. These corrections are incorporated into subsequent analytical iterations through prompt refinement and workflow adjustment.
This closed-loop structure improves system reliability through iterative error correc-tion and continuous validation. Human supervision ensures that system outputs remain aligned with economic reality and investment logic.
Over time, this interaction produces a stable and reliable analytical process in which AI systems provide scalable computational support while human analysts maintain supervisory control.
This hybrid human–AI architecture represents a practical and robust approach to integrating AI into financial decision-making, combining computational scalability with human judgment, accountability, and interpretability.
9. Case Study: Competitive Analysis with an AI Platform (Rivanna AI).
9.1. Use Case and Scope.
To validate the proposed AI-augmented value investing framework, we conducted a case study using Rivanna AI, an AI-assisted financial analysis platform, applied to a comparative analysis of three publicly traded companies in the U.S. beverage sector: The Coca-Cola Company (KO), PepsiCo Inc. (PEP), and Keurig Dr Pepper Inc. (KDP).
This sector provides a suitable evaluation environment due to its well-documented financial disclosures, comparable business models, and stable reporting practices. The focus on U.S.-listed firms supports high-quality and standardized primary data sources, including SEC filings (10-K, 10-Q, and 8-K), earnings transcripts, and financial statements, enabling rigorous evaluation of extraction accuracy, valuation consistency, and workflow efficiency.
The objective of the case study is to evaluate whether an AI-augmented, human-in-the-loop (HITL) analytical pipeline can improve the scalability and efficiency of financial analysis while preserving correctness, interpretability, and decision reliability.
9.2. Problem Formulation, Protocol, and Methods.
Fundamental financial analysis requires integrating heterogeneous information sources, including structured financial statements, unstructured regulatory disclosures, and market data. Analysts must extract and normalize financial variables, compute performance indicators, and construct valuation models such as discounted cash flow (DCF) analysis.
This process is time-intensive and prone to operational and cognitive failure modes. Operational challenges include unit inconsistencies, period misalignment, and data nor-malization complexity. Cognitive challenges include interpretation bias and inconsistency in analytical assumptions.
As a result, analysts face a structural trade-off between analytical depth and coverage. Increasing analytical coverage across multiple firms often reduces the time available for verification and interpretation, potentially degrading analytical quality.
While AI systems offer the potential to automate extraction and synthesis, financial workflows impose stringent requirements for traceability, correctness, and auditability. Even small numerical errors in extracted inputs may propagate into significant valuation deviations. Furthermore, generative AI outputs may introduce unsupported or unverifiable analytical claims if not properly constrained.
The key challenge is therefore not merely automating analysis but designing an AI-assisted analytical system that maintains correctness, transparency, and verifiability under real-world decision-making constraints.
Rivanna AI operates as an AI-augmented decision-support system rather than an autonomous decision-making agent. Its role is to automate operational components of the financial analysis pipeline while preserving human oversight and accountability.
The system performs several analytical functions, including automated extraction of financial variables from regulatory filings, structured generation of spreadsheet-compatible financial datasets, automated computation of financial metrics, and support for valuation modeling.
However, all outputs are subject to human verification through explicit HITL valida-tion gates. Analysts validate extracted financial values, confirm unit and period consistency, reconcile accounting identities, and review valuation assumptions before analytical outputs are used for decision-making.
An important operational observation emerging from the case study concerns the dependence of system performance on prompt formulation quality. When tasks were requested with insufficient context or as broad, open-ended analytical objectives, the system exhibited higher rates of inconsistency, incomplete extraction, or structurally incorrect outputs, increasing verification time and cognitive load for the analyst. Conversely, when the same analysis was decomposed into smaller, explicitly defined subtasks—with clear scope, period specification, and output format requirements—the reliability and consistency of the system improved significantly.
This behavior reflects a fundamental property of probabilistic generative systems: their outputs are conditioned not only on training data but also on the structure of the prompt itself. As a result, prompt formulation becomes an operational control mechanism that directly influences system correctness, efficiency, and usability. In this sense, prompt literacy is not merely a user convenience, but an integral component of the AI-supported analytical workflow.
The evaluation pipeline consisted of four sequential analytical tasks:
1. Extraction and normalization of fundamental financial variables from regulatory filings;.
2. Computation of key performance indicators, including profitability, leverage, and valuation multiples;.
3. Intrinsic value estimation using discounted cash flow (DCF) modeling;.
4. Generation of structured analytical summaries and investment rationale.
The intrinsic value was estimated using a standard DCF framework: n FCFt TV ∑ V = (1 + WACC)t + (1 + WACC)n t=1 where the terminal value is defined using the Gordon growth model:
FCFn+1 TV = WACC − g
DCF sensitivity analysis was performed across multiple discount rate and terminal growth assumptions to evaluate valuation robustness.
Prompts were designed to impose explicit extraction constraints, including fixed output schemas, unit and fiscal-period reporting, and source traceability requirements. To make this protocol more concrete, Box 1 reports an example of the structured prompt used in the case study for financial-variable extraction and numerical verification.
Task: Extract the following financial variables for [Company] for fiscal year [FY2024] from the provided primary filing.
Variables required: Revenue, Operating Income, Net Income, Total Assets, Total Liabilities, Shareholders’ Equity, Capital Expenditures, Depreciation and Amortization, Cash and Cash Equivalents, and Total Debt. Instructions:
1. Use only the provided primary filing as the source of truth.
Do not use prior knowledge or external information.
2. Return the results in a structured table with exactly the following columns:.
• Variable
• Reported value
• Unit (units/thousands/millions)
• Currency • Fiscal period
• Source citation
• Notes/flags
3. Preserve the exact reported numerical values as stated in the filing.
Do not round unless the source itself is rounded. 4. If a value is reported under a different but equivalent label, extract it and indicate the mapping in Notes/flags.
5. If a value is missing, ambiguous, or not directly stated, return “NOT FOUND” and explain.
briefly in Notes/flags.
6. Before returning the table, perform the following self-checks:.
verify that all values use consistent units and currency; # verify that fiscal-period labels are aligned; check whether the accounting identity Assets ≈ Liabilities + Equity is # approximately satisfied when the relevant fields are available; # flag any likely unit mismatch, missing value, inconsistency, or unsupported citation.
7. If the cited passage does not fully support the extracted value, return “REVIEW.
REQUIRED” in Notes/flags.
8. Do not provide free-form reasoning or narrative explanation.
Output format: Return only: (a) the structured table; and (b) a short bullet list titled Detected Issues, reporting any inconsistencies or flags.
This structured pipeline enables evaluation of both analytical correctness and operational efficiency.
The experimental evaluation addresses three research questions:
• RQ1: Extraction correctness—How accurately does the system extract financial vari-ables compared to manually verified ground truth?
• RQ2: Analytical consistency—Do AI-assisted KPI and valuation outputs remain con-sistent with manually verified analysis?
• RQ3: Productivity impact—How does AI assistance affect total analytical time and task allocation?
We report four extraction metrics.
1. Exact-match accuracy = correctly extracted fields/total evaluated fields.
3. Unit/scale error rate = fields requiring unit-scale correction/total fields.
4. Missingness = unavailable extracted fields/total requested fields.
For KPI consistency, mismatch rate is computed as inconsistent KPIs over total KPIs per family.
All percentages are reported with explicit denominators in the corresponding tables. In particular, extraction accuracy was evaluated using exact-match accuracy, relative error, unit consistency, and missing data rates.
Analytical consistency was evaluated by comparing computed KPIs and intrinsic valuation estimates with manually verified reference values.
Productivity impact was measured by comparing total workflow time between manual and AI-assisted analysis, including extraction, modeling, and validation tasks.
9.3. Results.
The experimental results provide evidence that AI-assisted analysis maintains high ex-traction accuracy, with an overall exact-match rate of 87.5% and negligible unit consistency errors. Relative extraction errors remained below 1% across evaluated financial variables, indicating high numerical reliability.
KPI computations and valuation outputs remained consistent with manually veri-fied reference results, confirming analytical correctness of AI-assisted workflows when combined with HITL validation.
Table 2 reports extraction correctness, through accuracy, relative error, scale error, missingness per company. Table 3 shows KPI consistency, via mismatch rate per KPI family (profitability, leverage, multiples). Table 4 reports the time study, that is, the time comparison and breakdown between a completely manual process and the Rivanna + HITL process (mean ± std). Finally, Table 5 reports the HITL ablation, showing the impact of verification gates on the critical-error rate.
Notes: Exact match is computed with a 1% relative-error tolerance. RE = relative error. Scale/unit error rate measures million/billion (or unit) mismatches. Missingness is the fraction of unavailable extracted fields among the selected benchmark fields.
Notes: Operating margin = EBIT/Revenues. Net margin = Net income/Revenues. P/E = Price/Diluted EPS. P/B = Price/(Equity per share). EV/EBITDA = (Market Cap + Total Debt − Cash)/EBITDA. FCF yield = (Operat-ing Cash Flow − Capex)/Market Cap. Values are rounded to two decimals and aligned to FY2024 inputs used in the case study.
Notes: Base intrinsic values come from the case-study DCF setup. Sensitivity bounds are computed with a standard DCF stress on discounting and terminal growth. Min intrinsic: higher discount rate and lower terminal growth (WACC +1%, g −0.5%). Max intrinsic: lower discount rate and higher terminal growth (WACC −1%, g +0.5%). Results are rounded to one decimal place. “Gap sign stable” indicates whether (Intrinsic–Market) keeps the same sign across min/base/max scenarios. “Ranking stable” indicates whether the relative attractiveness ranking remains unchanged across the sensitivity range.
Notes: The first two rows provide a consistent central decomposition aligned with the reported case-study envelope.
The most significant improvement was observed in operational efficiency. Total ana-lytical time was reduced from approximately 25–40 h using traditional manual workflows to approximately 8–12 h using AI-assisted analysis, corresponding to an efficiency improve-ment exceeding 65%.
Importantly, this efficiency gain did not eliminate human involvement but reallocated analyst effort from manual data processing to validation and interpretation tasks.
In particular, with respect to Table 5, the 32 h vs. 10 h values represent a central task-level decomposition, while the 25–40 h vs. 8–12 h values represent the empirical envelope observed across runs. Both are reported to distinguish structured baseline accounting from observed operational variability.
Overall effort reduction is computed as:
In the AI-assisted workflow, the reported verification gates comprise citation/source validation, unit and fiscal-period reconciliation, accounting-identity and consistency checks, and analyst review of the most influential valuation assumptions before approval of the final output.
The third row reports the empirical envelope already discussed in the manuscript (manual: 25–40 h; AI-assisted: 8–12 h).
9.4. Discussion and Operational Implications.
DCF analysis identified moderate valuation discrepancies across evaluated firms. Coca-Cola and Keurig Dr Pepper exhibited positive intrinsic value gaps relative to market prices, suggesting potential undervaluation under baseline assumptions. PepsiCo exhibited a negative intrinsic value gap, indicating limited margin of safety.
Sensitivity analysis confirmed that valuation conclusions remain dependent on dis-count rate and growth assumptions, reinforcing the importance of uncertainty-aware valuation interpretation.
These results provide evidence that AI-assisted analysis can produce economically meaningful valuation outputs suitable for decision support, provided that human valida-tion and interpretability mechanisms are maintained.
The primary operational benefit of AI-assisted analysis lies in the automation of data extraction, normalization, spreadsheet preparation, and preliminary KPI/valuation modeling, which substantially reduces the time traditionally spent on repetitive manual information processing. In the manual workflow, analysts must search across filings, locate relevant passages, transcribe numerical values, align fiscal periods and units, populate spreadsheets, and only then begin higher-level interpretation. In the AI-assisted workflow, much of this mechanical burden is compressed by prompt-driven extraction and structured output generation.
Importantly, the observed reduction in total analysis time, from approximately 25–40 h to 8–12 h, does not imply the removal of human effort. Rather, it reflects a shift in analyst effort from low-value procedural tasks toward higher-value supervisory tasks. In the proposed framework, human work is concentrated in the verification gates: checking source citations, reconciling units and fiscal periods, resolving discrepancies between extracted values and primary filings, reviewing accounting consistency, and adjudicating high-leverage valuation assumptions before final reporting. In this sense, AI changes the composition of cognitive work rather than eliminating it.
This distinction is important for interpreting productivity gains correctly. Verification is cognitively different from original manual research. Manual research is dominated by search, transcription, and spreadsheet construction; verification is dominated by focused review, exception handling, and judgment under traceability constraints. Because the AI system pre-structures the evidence and surfaces candidate values, the analyst can allocate more attention to economically meaningful questions—whether a number is supported by the filing, whether a discrepancy is material, and whether the valuation assumptions remain plausible—rather than to clerical processing. Accordingly, the efficiency gain should be understood not as “analysis without humans”, but as analysis with human effort reallocated from information processing to validation, adjudication, and interpretation.
From a systems perspective, this redistribution of effort is a central advantage of HITL architectures. Automation improves scalability, while structured verification preserves analytical reliability and decision accountability. The resulting workflow is therefore more suitable for repeatable financial decision support than either a purely manual process or an ungoverned autonomous AI pipeline.
9.5. Limitations, Threats to Validity, and Reproducibility Notes.
This study has four main limitations.
• First, valuation outputs remain assumption-sensitive (discount rates, growth, and terminal value settings).
• Second, the empirical evaluation is intentionally narrow (three firms, one sector), so external validity is limited.
• Third, performance depends on data availability and document quality.
• Fourth, prompt quality and operator expertise can influence outcomes, introducing interaction-related variance.
Accordingly, results should be interpreted as case-study evidence supporting the feasibility of AI-augmented, HITL-governed workflows, rather than universal performance claims.
The case study provides evidence that AI-assisted analytical architectures can improve efficiency and scalability of financial analysis while maintaining correctness and interpretability.
Rather than replacing human analysts, AI systems function as computational accel-erators within human-supervised analytical pipelines. The combination of automated extraction, explainable modeling, and human validation provides a robust framework for scalable and reliable financial decision support.
These findings support the viability of AI-augmented, human-in-the-loop architectures for real-world financial analysis applications.
An additional threat to external validity concerns the information richness of the empirical setting. The present case study is conducted on large, well-covered U.S. firms characterized by relatively standardized disclosures, high reporting frequency, and broad availability of structured and semi-structured financial information. These conditions are favorable for prompt-driven extraction, citation grounding, and downstream valuation consistency. By contrast, in small-cap firms, emerging-market issuers, or other information-scarce environments, the proposed framework may exhibit lower reliability due to thinner disclosures, more heterogeneous reporting conventions, lower-quality source formatting, and weaker representation of firm-specific patterns in the pretraining corpus of large language models.
Under such conditions, one may reasonably expect higher missingness, more frequent citation ambiguity, greater disagreement across model outputs, and a larger burden of Human-in-the-Loop adjudication. Therefore, the performance gains documented in this study should be interpreted as conditional on a relatively high-information regime, and future validation should assess the robustness of the framework across firms and markets with lower disclosure quality and reduced informational coverage.
A related qualification concerns the different disclosure depth and approach typically observed across large-cap, mid-cap, and small-cap firms. Smaller issuers often provide less extensive, less structured, and less frequent public disclosure than large firms, not necessarily because fewer strategic initiatives are undertaken, but because communica-tion resources, reporting practices, and regulatory environments may differ substantially across firms and jurisdictions. As a result, large firms are often more suitable for traceable extraction, reconciliation, and valuation support, particularly when disclosure require-ments are detailed and standardized. However, this limitation does not affect AI alone; it also constrains the work of the human analyst, whose judgment is equally bounded by the quality, granularity, and comparability of the available information.
In this sense, the proposed framework should be viewed as largely neutral with respect to disclosure scarcity: it does not remove differences in disclosure depth, but it provides a governed and traceable method for extracting, validating, and reasoning over whatever information is actually disclosed.
To improve reproducibility, we specify the operational protocol used in the case study.
• Scope and sources: three U.S. listed beverage firms (KO, PEP, KDP), FY2024-oriented comparative analysis based on publicly available filings and company disclosures used in the workflow.
• Task decomposition: each run follows four sequential tasks: extraction/normalization, KPI computation, DCF valuation, and memo synthesis.
• Prompt protocol: prompts are structured with explicit objectives, context constraints (firm, period, units), output schemas (tabular/spreadsheet-ready), and verification requests (source traceability and uncertainty flags).
• Ground-truth and adjudication: extracted values are compared against manually verified references; disagreements are resolved through analyst review and source-level reconciliation.
• Validation gates: HITL checks include unit/period consistency, accounting identity reconciliation, and DCF assumption review before final reporting.
This protocol is intended to make the experimental pipeline auditable and repeatable, while acknowledging that broader multi-sector validation remains future work.
10. Conclusions and Future Work.
This article presented a structured framework for AI-driven value investing that combines automated data extraction, scalable modeling, explainability, and human-in-the-loop governance to support strategic financial decision-making.
The results provide case-study evidence that AI-assisted workflows can materially improve analytical productivity (from a manual range of 25–40 h to 8–12 h in our setting) while maintaining accountability through traceability and HITL verification gates. It is important to note that the quantified savings reported in this study reflect the initial implementation stage; further productivity improvements are expected through iterative prompt refinement, HITL workflow standardization, and organizational learning-curve effects beyond the currently measured values.
A key conceptual contribution of this work is the identification of prompt literacy as a core operational competence of the augmented analyst. The case study shows that the clarity, structure, and contextual completeness of prompts directly influence extraction correctness, analytical consistency, and workflow efficiency. Structured, decomposed prompts improve reliability and reduce verification overhead, whereas vague or overly complex requests increase error rates and human correction effort.
These findings reinforce a broader conclusion: AI in financial analysis functions not as an autonomous decision-maker but as a probabilistic reasoning amplifier whose effectiveness depends on structured human guidance, verification discipline, and critical interpretation.
In the long run, competitive advantage in AI-augmented investing may depend less on access to raw data or computational power, and more on the ability to interact with AI systems effectively, namely formulating precise questions, validating outputs rigorously, and integrating AI-generated insights into coherent economic reasoning.
Future research directions include:
• formal quantification of prompt quality and its impact on extraction and valuation accuracy,
• integration of probabilistic uncertainty bounds into valuation outputs,
• extension of AI-augmented workflows across broader sectors and international markets, • development of standardized governance and verification protocols for AI-supported financial analysis.
A key next step is to validate the framework across multiple sectors, small- and mid-cap firms, and international regulatory environments, where disclosure quality, stan-dardization, and comparability may differ substantially. Such validation is important not because information scarcity is uniquely problematic for AI, but because it repre-sents a structural constraint on both human and AI-assisted analysis, which the proposed framework addresses through traceability and governance rather than by assuming richer information than the market actually discloses.
Ultimately, AI shifts the analyst’s role from information processor to reasoning archi-tect, namely designing, guiding, and validating analytical workflows that combine human judgment and machine intelligence.
analysis, A.C.; Investigation, A.C.; Writing—original draft, A.C. and L.R.C.; Writing—review & editing, L.R.C. All authors have read and agreed to the published version of the manuscript.
Funding: This research received no external funding.
Data Availability Statement: The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.
Conflicts of Interest: The authors declare no conflict of interest.
1. Graham, B.; Dodd, D.
Security Analysis; McGraw-Hill: New York, NY, USA, 1934.
2. Damodaran, A.
Investment Valuation: Tools and Techniques for Determining the Value of Any Asset; Wiley: New York, NY, USA, 2012.
Altman, E.I. Financial ratios, discriminant analysis and the prediction of corporate bankruptcy. J. Financ. 1968, 23, 589–609. 3. [CrossRef]
4. Gu, S.; Kelly, B.; Xiu, D.
Empirical asset pricing via machine learning. Rev. Financ. Stud. 2020, 33, 2223–2273. [CrossRef]
Heaton, J.; Polson, N.; Witte, J. Deep learning in finance. Annu. Rev. Financ. Econ. 2017, 11, 1–23. 5.
6. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. arXiv 2017. [CrossRef]
7. Devlin, J.; Chang, M.; Lee, K.; Toutanova, K.
BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019. [CrossRef]
8. Tetlock, P.
Giving content to investor sentiment. J. Financ. 2007, 62, 1139–1168. [CrossRef]
Loughran, T.; McDonald, B. When is a liability not a liability? Textual analysis, dictionaries, and financial reporting. J. Financ. 9. 2011, 66, 35–65. [CrossRef]
10. Power, D.J.
Decision Support Systems: Concepts and Resources; Business Expert Press: New York, NY, USA, 2002.
11. Shrestha, Y.R.; Ben-Menahem, S.M.; von Krogh, G.
Organizational decision-making structures in the age of artificial intelligence. Calif. Manag. Rev. 2019, 61, 66–83. [CrossRef]
12. Davenport, T.H.; Kirby, J.
Only Humans Need Apply; Harper Business: New York, NY, USA, 2016.
13. Amershi, S.; Weld, D.; Vorvoreanu, M.; Fourney, A.; Nushi, B.; Collisson, P.; Horvitz, E.; Iqbal, S.; Suh, J.; Teevan, J.; et al. Guidelines for human-AI interaction. In CHI Conference on Human Factors in Computing Systems; ACM: New York, NY, USA, 2019.
14. Doshi-Velez, F.; Kim, B.
Towards a rigorous science of interpretable machine learning. arXiv 2017, arXiv:1702.08608. [CrossRef] Lundberg, S.; Lee, S.I. A unified approach to interpreting model predictions. In NeurIPS; Curran Associates Inc.: Red Hook, NY, 15. USA, 2017.
16. Ribeiro, M.T.; Singh, S.; Guestrin, C.
Why should I trust you? Explaining the predictions of any classifier. In KDD; ACM: New York, NY, USA, 2016.
17. Molnar, C.
Interpretable Machine Learning; Lulu: Raleigh, NC, USA, 2020.
18. Araci, D.
FinBERT: Financial sentiment analysis with pre-trained language models. arXiv 2020, arXiv:1908.10063.
19. Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y.T.; Li, Y.; Lundberg, S.; et al. Sparks of artificial general intelligence: Early experiments with GPT-4. arXiv 2023, arXiv:2303.12712. [CrossRef]
20. Sahoo, P.; Singh, A.K.; Saha, S.; Jain, V.; Mondal, S.; Chadha, A.
A Systematic Survey of Prompt Engineering in Large Language Models. arXiv 2025, arXiv:2402.0792.
National Institute of Standards and Technology (NIST). Artificial Intelligence Risk Management Framework (AI RMF 1.0); NIST: 21. Gaithersburg, MD, USA, 2023.
ISO/IEC 23894:2023; Information Technology—Artificial Intelligence—Guidance on Risk Management. International Organization 22. for Standardization (ISO): Geneva, Switzerland; International Electrotechnical Commission (IEC): Geneva, Switzerland, 2023.
23. Gebru, T.; Morgenstern, J.; Vecchione, B.; Vaughan, J.W.; Wallach, H.; Daumé, H., III; Crawford, K. Datasheets for Datasets. In Communications of the ACM; ACM: New York, NY, USA, 2021; Volume 64, pp. 86–92.
Author Contributions: Conceptualization, M.G.; Methodology, A.C., M.G. and L.R.C.; Formal
24. Mitchell, M.; Wu, S.; Zaldivar, A.; Barnes, P.; Vasserman, L.; Hutchinson, B.; Spitzer, E.; Raji, I.D.; Gebru, T. Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT); ACM: New York, NY, USA, 2019; pp. 220–229.
25. Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-T.; Rocktäschel, T.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems (NeurIPS); Curran Associates Inc.: Red Hook, NY, USA, 2020.
26. Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.V.; Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Adv. Neural Inf. Process. Syst. 2022, 55, 1–38.
27. Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Zhou, D.
Self-Consistency Improves Chain-of-Thought Reasoning in Language Models. arXiv 2022, arXiv:2203.11171.
28. Liu, P.; Yuan, W.; Fu, J.; Jiang, Z.; Hayashi, H.; Neubig, G.
Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. In ACM Computing Surveys; ACM: New York, NY, USA, 2023; Volume 55, pp. 1–35.
29. Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.J.; Madotto, A.; Fung, P. Survey of Hallucination in Natural Language Generation. In ACM Computing Surveys; ACM: New York, NY, USA, 2023.
30. Rivanna AI.
Available online: the linked source (accessed on 16 February 2026).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.