Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
1 More Paper · Full Reading

About this paper
Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise ware- houses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an eval- uation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, mod- eled on the Oracle E-Business Suite schema. The simulator’s ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocat- ing courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solu- tion that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.
Authors: Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Published in: arXiv
Publication date: 2026-10-01
Read the paper: https://doi.org/10.48550/arXiv.2610.02122
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows,” by Gabriel Tomitsuka and colleagues. Published in arXiv on October 1, 2026.
A RGO -B ENCH: E VALUATING DATA AGENTS ON E NTERPRISE -S CALE W ORKFLOWS
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma TextQL
{gabriel,arman,emma,duke,joseph}@textql.com
A BSTRACT
Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise ware-houses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an eval-uation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives.
We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, mod-eled on the Oracle E-Business Suite schema. The simulator’s ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocat-ing courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solu-tion that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.
1 I NTRODUCTION
In recent years, data agents have advanced from answering simple text-to-SQL questions over small, well-documented schemas to performing long-horizon data science tasks, publishing durable company-wide assets such as dashboards, and acting on their findings, for instance, by adjusting promotions and banning fraudulent accounts. Rising scores on established data benchmarks could be read as a sign that enterprise data work is close to solved. The best Spider 2.0-Snow score has risen from 23.8% at release to 96.7%, and the top BIRD entry reaches 82.4% against a human 93.0%.1
However, current benchmarks primarily evaluate text-to-SQL performance and are not representa-tive of enterprise agentic data workflows in three important ways. First, they run on a patchwork of public data. Spider 2.0-Snow’s 547 questions sparsely cover a sprawling collection of 152 databases (Appendix H). Over a third of those databases are used for a single question only. Two thirds of the questions use BigQuery public datasets; another quarter use local sample databases. These public and local sources’ schemas are documented in tutorials and textbooks. Fur-thermore, there is a high degree of overlap among tables: at least two thirds of the tables are date-, geography-, or version-sharded copies of another table. In an enterprise warehouse, different mod-ules must agree on the same numbers, and one business event may touch ten or more tables.
Second, they are not end-to-end. Many of the most valuable enterprise data science tasks cannot be done in SQL. Fitting a demand forecast, training a fraud model, or optimizing a courier sched-ule requires statistical, machine learning, and optimization libraries. Additionally, findings lead to decisions, which are ultimately judged by their return on investment. For instance, a promotion with a high redemption rate can still lose money if most of those orders would have been placed anyway. A gold answer cannot distinguish between such decisions.
Third, they are not verifiable. On real data, a benchmark can only measure agreement with its annotators since writing the answer key requires solving the task. Even when the answer is in the data, annotators may miss it. In a recent audit, 62.8% of Spider 2.0-Snow’s released gold queries and 52.8% of BIRD Mini-Dev’s were found to be erroneous, most often because annotators misread the data or the schema, and correcting BIRD’s errors moved agents’ leaderboard standings by up to nine places. Furthermore, for many of the most valuable enterprise tasks, the answer may not be in the data at all. Fraud left undetected leaves no label, and the outcomes of a rejected decision are never observed.
Large enterprises keep the records that financial planning, forecasting, and fraud detection require in enterprise resource planning (ERP) systems such as Oracle E-Business Suite (EBS), SAP S/4HANA, and Oracle Fusion Cloud, extended with custom tables representing id-iosyncrasies and workflows unique to the business. Because the same systems hold the company’s ledgers, payroll, and customer records, access to them is heavily restricted. For this reason, ERP data has remained practically unexplored in benchmarks to date. When such data is released, it must be anonymized, which can break the relational structure that answers depend on. In an early release of BEAVER, the closest attempt to date, Chung et al. (2025) report primary keys that violate their uniqueness constraints and questions whose gold query is null.
We instead simulate a business and grade against the simulation’s ground truth (Figure 1). Argo-Bench models a food delivery platform in New York City in 2024, a three-sided marketplace whose economics are disclosed to the city every month. The simulator reflects a real minimum-pay increase in April 2024, and is calibrated to these disclosures and to public filings. None of its data is generated by a language model. We project this world into an
Oracle EBS warehouse that omits the simulator’s latent state, so tasks are harder to solve than to verify. The agent must reconstruct facts from the warehouse, while the grader reads them off the state. Agents work in a sandboxed Python environment with statistics, machine-learning, and optimization libraries (Appendix F), and file their decisions and results through a mission-control interface rather than returning a query. Because the grader knows the latent state, it can grade a decision by its consequences. For instance, a list of banned accounts is scored by the fraud losses it prevents, including losses from fraud that the platform never detected, net of the revenue lost from wrongly banned customers.
Simulated environments are an established practice in a broad range of domains, and synthetic warehouses have long been used to benchmark data systems. Even Spider 2.0 draws a fifth of its questions from synthetic, obfuscated, or textbook sample schemas (Appendix H). To our knowledge, no prior simulated bench-mark combines an enterprise-scale warehouse whose tables must agree with one another and a grader that scores the consequences of the agent’s actions against the simulator’s latent state.
We evaluate 14 frontier and open-weight models on Argo-Bench. The strongest, Claude Opus 5.5, solves 34.8% of tasks and averages 59.5 points. Nine of the fourteen average below 35. Models often analyze the wrong quantity or optimize the wrong objective.
We make the following contributions:
• A synthetic, large-scale public ERP dataset in Oracle EBS format for a 2024 New York City food delivery platform, comprising 235 mutually constraining tables, 81 million orders, 3.4 million active customers, and 7.5 billion rows, in which an order resolves into dispatch decisions, courier pay, merchant payouts, and balanced general-ledger journals.
• A grader that scores the consequences of an action rather than the correctness of a query, using the simulator’s latent state as ground truth, for instance, to score bans by the fraud losses they prevent and forecasts against held-out months, supporting tasks across fraud detection, forecasting, and financial planning.
• A benchmark of 210 such tasks, from publishing a dashboard data source to fitting forecasts and banning fraudulent accounts, with an evaluation of 14 frontier and open-weight models.
Argo-Bench is public. One world’s warehouse is released on Hugging Face under CC BY 4.0,2 and the tasks, reference agent, tool server, and sandboxes under the Apache License 2.0.3 The numbers and experiments in this paper use a second world from a private seed, with different customers, couriers, and answer keys. The leaderboard is scored only on this world, so a score cannot be earned by memorizing the released warehouse. A demo at the linked source readers to browse the orders and deliveries of the released world at a 10% scale.
2 B ENCHMARK C ONSTRUCTION 2.1 WORLD SIMULATION
We simulate a food delivery platform in New York City (NYC) in 2024, similar to DoorDash, Grub-hub, or Uber Eats. Public data on such platforms is aggregate. The city’s quarterly reports and the platforms’ own filings give totals, but order-level records cannot be released without exposing the platform’s customers, couriers, and margins. The world is therefore built from 34 public datasets and reports, each lending one mechanism (Appendix A). Uber and Lyft trips, for instance, give the time to drive between two zones at a given hour, and MenuStat the menus for restaurant chains. Donors play one of three roles. Identity donors are public records of real entities in the city, such as its 45,834 restaurants, 1.07 million addresses, and 260 taxi zones, and enter the world as they are, except that the released warehouse renames some restaurants (see the ethics statement).
Shape donors are measured elsewhere, on other people, in another city, or in another year, and lend the world only a distribution. Anchors are published totals that the world is calibrated to reproduce but never samples records from.
We chose food delivery in NYC because its economics are unusually well documented. Delivery apps must report their monthly orders, consumer spending, merchant fees, courier earnings, pro-ductivity, and hours worked to the NYC Department of Consumer and Worker Protection (DCWP), which publishes them quarterly, and the 10-K filings of DoorDash and Grubhub give the shape of a platform’s balance sheet (DoorDash, Inc., 2024b; 2025; Grubhub Inc., 2021). Following Walonoski et al. (2018), we calibrate the world to these anchors, sized as a dominant platform. The result has 81 million orders, about 55% of the 148 million deliveries that apps re-ported to the DCWP for 2024, and 3.4 million active customers, and it must balance incentives on three sides: quests and suggested pay for couriers, promotions and surge pricing for customers, and co-funded campaigns for merchants.
Its per-delivery economics stay within 5% of the DCWP’s figures in 13 of 16 quarterly comparisons (Figure 2a).
We selected 2024 because it contains a real extrinsic shock to these economics. NYC began enforc-ing a minimum pay rate of $17.96 per hour before tips for app-based restaurant delivery workers in December 2023 and raised it to $19.56 on April 1, 2024. We model the platforms’ response with a new courier scheduler that activates on that date, after which courier pay runs 6–9% above the anchor (Figure 2a). The levers our platform uses may differ from those of the real plat-forms, but the aggregate effect is the same in direction and, to within 9%, in size, and because the world absorbs the same shock, we can ask realistic forecasting questions about it (Appendix B).
Additionally, we insert fraud patterns that public evidence shows are major problems for delivery platforms. Couriers steal orders after pickup, spoof their GPS, grab offers with bots, or rent out their accounts (DoorDash, Inc., 2024a). On the customer side, rings of new accounts farm promotions (DoorDash, Inc., 2025; Incog-nia, 2025), and stolen cards fund account takeovers and bust-outs (DoorDash, Inc., 2025; Whittaker, 2018). Storefronts may be shells or collude with couriers or regular customers on refunds (DoorDash, Inc., 2023), and their payouts can be diverted to changed bank accounts. Each pattern is calibrated both to how separable real card fraud is and to how often honest customers share a device, an address, or a card, since the latter sets a detector’s precision (Ap-pendix A).
The simulator generates the world from the donors, calibrates it to the anchors, inserts the fraud, and projects the result to the EBS format (Section 2.2).
2.2 WAREHOUSE D ESIGN AND VALIDATION
The simulator’s last step projects the world into what an analyst actually sees: the analytics export of a greenfield Oracle E-Business Suite (EBS) 12.2 instance. We chose EBS because its data model is publicly documented, so our schema can be verified against a reference. We designed the schema with three ERP consultants who have 14 to 31 years of experience.
Every standard table and column in our warehouse exists in the data dictionary of Oracle’s EBS 12.2 Vision instance, the demo environment Oracle provides as a reference. Our tables, however, carry on average 52% of the columns of their Vision counterparts. EBS serves every industry, and analytics exports omit the columns a business does not use, here unused flexfields (a third of the omitted columns) and features such as shipping, inventory, foreign currency, and withholding tax. These columns would be empty in this business’s data, so no task loses information by their omission.
The warehouse does not contain data drift or inconsistencies, such as deprecated tables that overlap active ones or figures that fail to reconcile across tables. These inconsistencies sometimes accumu-late in real data warehouses over years of migrations and acquisitions. Though the consultants named this the most significant difference from their customers’ systems, we deliberately chose to keep this discrepancy. By doing so, we keep the ground truth unambiguous: if a legacy table were to disagree with an active one, the correct answer would depend on undocu-mented conventions. Mature warehouses compensate for their idiosyncrasies with semantic layers, data models, and institutional knowledge. While we could provide such a layer alongside a more realistic messy warehouse, this would shift the evaluation’s focus to testing an abil-ity to use a curated layer.
We instead aim to test an understanding of enterprise data organization: production workloads reuse only a few dozen combinations of hundreds of tables (van Renen et al., 2024), and an agent new to a warehouse must discover which ones matter by deciding what and how much of the warehouse to explore. A greenfield warehouse isolates this skill of understanding and exploring enterprise data organization.
Otherwise, the consultants found the schema realistic, with two further omissions. It records no foreign-currency transactions, since no donor dataset covers the currencies visitors pay in, and it has only 43 balance sheet accounts (Section 5).
2.3 TASK D ESIGN
Argo-Bench contains 210 tasks in five business areas (Figure 3a). Trust and safety tasks are en-forcement, where the agent finds fraud and abuse and acts on the accounts involved. FP&A tasks forecast unit economics and rebuild finance dashboards, marketplace tasks forecast courier supply and allocate budgets such as courier bonuses, accounting tasks report final values after the fact, and growth tasks measure, forecast, and publish the results of promotions and memberships.
Each task has four parts (Appendix G). The prompt states the problem as a stakeholder would. The scope sets the last month of 2024 visible to the agent. The expectations list the filings the grader requires, each with its action, keys, and grading mode. The answer key is frozen before any run and comes from SQL over the latent tables or from the simulator’s own labels, such as which couriers stole orders. Because the world is simulated, these labels are exact and need no anonymization.
Prompts cover a range of writing styles and levels of detail. Some reference the grading criteria or the exact output expected, while others are more subtle. This reflects how real users pose data questions: loosely, as high-level business questions, and in no set style or template. It also tests the skill of exploring data organization, since a less detailed prompt leaves the agent to discover which tables and conventions the question depends on. Several scenarios come in variants that differ only in such detail, and Appendix I compares them.
Unlike most data science and analytics benchmarks, Argo-Bench does not ask agents to return a query. Agents file actions to a mission-control interface through a Python library in their sandbox (Appendix E). This design has four advantages. First, it permits advanced data science tasks which require machine learning, operations research, and mathematical optimization libraries to complete. Second, it supports end-to-end workflows. Emitting the correct SQL is not enough in practice, as real tasks require taking actions, e.g., rebalancing courier incentives across zones and hours, banning a set of users who are likely committing fraud, or holding the payouts of a suspicious merchant. Third, filings are explicit declarations of intent.
When benchmarks compare SQL results, it is hard to determine whether a close number is the agent’s answer or an intermediate result, whereas a filing states the value, interval, or reason the agent commits to, which also makes partial credit well defined. Fourth, grading is objective, unlike the LLM-as-a-judge evaluation used in many related works.
Of the 210 tasks, 178 act on accounts, file a forecast, allocate a budget, or publish a dashboard data source, and 99 see only up to a cutoff month, as an analyst would at that date, so forecasts are graded on months the agent has not seen (Figure 3). Prompts average 159 words, 51 tasks require more than one filing, and the reference solution in Figure 1 joins six tables across three EBS modules to recover the minimum-pay rule and fit an interval. In size, Argo-Bench matches long-horizon agent benchmarks such as TheAgentCompany (175 tasks), τ -bench, KramaBench, and ELT-Bench.
Each expectation is scored from 0 to 100 (a forecast from −200) by one of nine grading modes (Appendix G). Forecasts are scored by their weighted interval score on a scale set by a reference forecast fixed before the outcome, so that filing one’s true median and interval is the best strategy, ban lists by the cost they save relative to banning nobody or everybody, budget allocations by the share of the attainable savings that the simulator realizes, and data sources and reported figures by their values. A task’s score is the mean of its expectations’ scores, weighted as the task specifies. A run that files nothing where the key expects action scores zero on every expectation (−200 on a forecast, the lowest a forecast can score).
We release the warehouse of one world on Hugging Face and keep a second, generated from a private seed, for official grading. We ran our experiments on BigQuery (941 GB uncompressed), and the released warehouse is a set of 1,219 Parquet files (76.5 GB), with a dataset card giving setup instructions for BigQuery, DuckDB, Snowflake, Trino, Delta Lake, and Iceberg. The two seeds share the simulator and its calibration, but every ID, customer, restaurant, and courier differs.4 Table 1 compares Argo-Bench with prior benchmarks.
3 E VALUATION
We run all experiments on Inspect AI 0.3.263, an open-source evaluation framework, on Google Kubernetes Engine. Each agent works in its own gVisor sandbox with 25 preinstalled Python libraries, an empty file system, and no out-bound internet access, and files to its own mission-control instance through a Python library (missioncontrol.py) that emulates a company’s internal tooling. A run ends after 500 model turns, to stop models that loop without progress, and Appendix F lists the other limits. We evalu-ate models available in September 2026 across price ranges (Table 2), calling open-weight models through Fireworks serverless endpoints and the others through their developers’ APIs, with default sampling settings for every model.
3.1 R ESULTS
Claude Opus 5.5 leads overall and in three of the five domains (Table 2), GPT-6 Astra leads in forecasting, and Claude Sonnet 5.5 leads in compliance. More reasoning effort helps the GPT-6 and Claude models at every step, with diminishing returns for Opus, which gains 17.0 points from low to medium effort, 5.9 from medium to high, and 4.0 from high to extra-high, while Claude Sonnet 5.5 gains 14.9 and 11.5 over the last two steps. Gemini 3.8 Flash gains 8.6 points from low to medium [4.9, 16.3] and Qwen 3.8 Max gains 5.2, and neither gains detectably afterward. Muse Spark 1.3 changes by less than 2 points past medium, and DeepSeek V4.1 Flash moves only at extra-high (Appendix I).
Gemini 3.8 Flash and Muse Spark 1.3 also average 197 and 250 model calls per task, compared to 82 for Opus, without scoring higher, and 57% of Muse’s warehouse spend goes to tasks on which it scores below 5 out of 100 (Appendix K).
Many tasks are prompt variants of one scenario, so we resample the 146 base scenarios when boot-strapping. The resulting 95% confidence intervals are about ±2 to 8 points on the score and up to
†Spider 2.0-Lite, as computed by Chen et al. (2024).
±8 points on the solved rate. Paired by task, Opus leads each of the next three models (GPT-6 Astra, Claude Sonnet 5.5, and GPT-6.1 Sol) by 7.7 to 10.0 points, and those three are not separated from one another. GPT-6.1 Sol leads its predecessor GPT-6 Sol by 12.8 points [6.8, 19.1], GPT-6 Sol leads Kimi K3 by 8.4 [0.5, 15.4], and the intervals of the next five overlap (Appendix I).
3.2 F INDINGS
Many failing runs use sound methods but read the wrong record, optimize the wrong objective, or measure the wrong quantity. Passing runs check definitions against a second source (Appendix I).
Wrong record. A marketing dashboard asks whether each discount offer paid off against the same push notification with the coupon left out for every thousand customers it was sent to, for members and non-members. Only Claude Opus 5.5, GPT-6 Astra, and GPT-6.1 Sol at extra-high effort publish the right table. What separates them is who counted as a member at the moment of targeting. A member whose card is declined keeps the benefits for seven days of grace and then loses them, but the contract stays on the books until it is canceled weeks later. Twenty-four of the failing runs take membership from the contract’s dates, and seven of them get every other figure right to within a few dollars. The three passing runs replay the billing history instead and check the result against the orders on which a member benefit was actually applied.
Opus reads the seven-day grace off those orders, and GPT-6.1 Sol starts from the contract dates, finds 1,057 orders that its rule calls members’ and that received no benefit, and starts over.
Wrong objective. One family of tasks asks where to cut $9.6 million from the annual budget for courier bonuses (quests). At extra-high effort and without hints, GPT-6 Astra finds the zones and hours where quests were randomly withheld, estimates supply responses with fixed effects and partial pooling, and cuts where quests buy the fewest courier-hours. The platform pays for quests to avoid surge pay; however, under the simulator’s response model, the plan loses $86,281 where a uniform cut would save $0.40 million, and it scores 0. From the same prompt, Claude Opus 5.5 at extra-high effort finds that quests substitute for surge pay, cuts where they save the least surge per bonus dollar, and saves $3.09 million of an attainable $3.12 million (score 99). Stating what quests are for and that a holdout exists raises the mean over all settings from 18.5 to 61.7.
Wrong quantity. Many failures come from measuring a quantity other than the one the prompt asks for. Of the 47 completed runs of a task that sizes the courier location service, 46 miss all 12 monthly counts despite a 1% tolerance because the warehouse keeps only hourly idle check-ins while the app sends them every half hour. GPT-6 Astra notices the gap, writes that its count does not establish how many reports the app sent, and files it anyway. Only Claude Sonnet 5.5 at extra-high effort adds the missing half-hourly check-ins back. Across the dashboard tasks, 66.9% of the data sources that pass their structural contract score zero on their values.
Overconfident forecasts. Across 4,553 forecast series from 3,346 runs on 72 tasks, nominal 80% intervals contain the realized value only 44.8% of the time. Apart from its floor, the grade is proper, so this overconfidence costs models points in expectation, although a per-series skill score clipped at zero would have rewarded a narrower interval in 72% of series (Appendix I). Grades also depend on the reference, which sets each series’ scale. Against the tighter reference of their harder variant, the June base-pay forecasts fall from a mean grade of 84.7 to 6.0, although their median absolute error is only 1.57%. Appendix I reports coverage and absolute error.
4 R ELATED WORK
Text-to-SQL and data science benchmarks. Text-to-SQL benchmarks have moved from databases with a handful of tables each to enterprise-scale schemas and private data warehouses, and data science bench-marks extend evaluation to multi-step analysis, data lakes, and data pipelines. These benchmarks compare an agent’s output to a gold answer, which audits have found to be frequently wrong, or to expert conclusions, often scored by an LLM judge. Argo-Bench instead grades the actions that an agent files against the simulator’s latent state.
Agent benchmarks in simulated environments. Simulated environments are the standard way to evaluate agents that act. AppWorld and τ -bench check the final state of the environment’s database. TheAgentCompany scores checkpoints in a simulated software company, and CRMArena-Pro shapes LLM-generated CRM records with latent variables. Vending-Bench and Business Arena score the net worth of a business that the agent runs. In these benchmarks, the state that determines success is either observable to the agent or changed by the agent inside a stylized game. In Argo-Bench, it is withheld and must be reconstructed from an enterprise warehouse. Simulators calibrated to public statistics have likewise supplied ground truth that real data lacks for patient records and money laundering.
Generated enterprises with hidden ground truth. AvalancheBench uses an LLM judge to score how much of a small latent e-commerce world an agent’s report recovers. The Era by Eon benchmark serves a generated company through a fleet of 66 simulated products, calibrated to operational and published statistics, and plants the records that answer each of its read-only questions, together with near misses. Two unrelated benchmarks named ERPBench evaluate decisions in a simulated manufacturer and computer-use tasks in a live ERP system. In contrast, the latent state of Argo-Bench is produced by the simulation itself rather than being planted. Its evidence is spread across an enterprise warehouse of 7.49 billion rows, and agents are graded on the consequences of the actions they file rather than on the answers they give.
5 L IMITATIONS & F UTURE WORK
The dataset still differs from the most complex ERP deployments in four ways. First, the dataset covers a single city, so it has no foreign currency and none of its representations as transaction, local, and reporting currency. Second, many complex warehouses combine several businesses, such as food delivery and grocery delivery, with shared accounts, such as driver payables and customer credits, mixing them in a single balance. Third, only one year is simulated. A longer history, such as 2017 to 2026, would add the market shocks of 2020 to 2022 to operations and forecasts. Fourth, only one ERP format is supported, and a natural extension would be to add SAP S/4HANA.
The simulator validates 23 distinct metrics, and we keep behaviors that public figures do not con-strain out of scope for tasks. The in-world membership program in particular rests on weakly grounded assumptions. More generally, a simulator encodes the assumptions of its authors, and a generator that shares the simplifying assumptions of the systems under evaluation can make tasks easier than their real counterparts. Calibration to aggregate targets also does not guarantee realistic tails. Argo-Bench therefore compares data agents and does not estimate their performance on a real company’s warehouse.
Each setting of the model and effort has one run per task, and all tasks share one simulated world, so our confidence intervals reflect the choice of tasks rather than run-to-run variation. Scores on some tasks are also sensitive to prompt wording, so Appendix I reports every hinted and unhinted pair, including one prompt we judged to be underspecified. Plan tasks are graded under the simulator’s frozen response model, and Appendix J lists open issues by task.
6 C ONCLUSION
We introduced Argo-Bench, which evaluates data agents on a simulated food delivery platform exported to an Oracle E-Business Suite warehouse of 235 tables and 7.49 billion rows. Because the simulator’s latent state is withheld from the warehouse, Argo-Bench grades the facts agents reconstruct and the actions they file, with every answer key computed from ground truth. The strongest of 14 models solves 34.8% of tasks and averages 59.5 points, and its most instructive failures are careful analyses of the wrong quantity, toward the wrong objective, or read from the wrong record. We release the public seed’s warehouse, tasks, reference solutions, and harness, and hope that Argo-Bench helps measure progress toward data agents that can understand, navigate, and act within real enterprise data environments.
AI USE STATEMENT
Large language models were used in three ways. First, as coding assistants for the simulator, graders, reference solutions, and figure and table code. Second, in the simulator’s data work: tuning its parameters, cleaning donor datasets, and reconciling restaurant records across sources (Section 2.1). No record in the simulated world is generated by a language model, and every value in the warehouse comes from the simulator, except the fictional names of the storefronts renamed in the released warehouse, which a language model drafted (see the ethics statement). Third, in writing: some task prompts were drafted by a language model and then curated and rewritten by the authors, and language models edited the text of the paper and checked its citations. The authors checked every claim, number, and citation, and take full responsibility for the content of the paper.
E THICS STATEMENT
Argo-Bench contains no data about real people. Consumer and courier names are drawn from pub-lic name-frequency tables and assigned to simulated people, and every order, shift, and payment is simulated. Restaurants are real New York City businesses. Their names and addresses come from Overture Maps places matched to the city’s inspection records, with some names updated to the busi-ness’s current listing, and some storefronts the simulator opens during the year, such as relaunches and virtual brands, take the name of a real business that was not trading at the time. All behavior attributed to them in the world, including the fraud patterns of Section 2.1, is simulated. The sim-ulator draws which storefronts play a fraud role, not from any record of a business’s conduct, so a label says nothing about the real business.
Even so, in the released warehouse every storefront that plays a fraud role in any scenario carries a fictional name instead of its real one, and fields derived from the name, such as its contact email, follow the new name. The released questions and agent transcripts use the same fictional names. A language model drafted the fictional names to read like real New York restaurant names, so that a renamed storefront does not stand out and point to the answers, and each was checked against the city’s inspection records and Overture Maps so that none is the name of a real restaurant. Addresses are unchanged, since the simulated geography depends on them. The donor datasets are used under their licenses.
Those released for research or non-commercial use, such as the Yelp Open Dataset and the Grubhub MDRP instances, serve only as shape donors, and none of their records enters the world or the warehouse (Appendix A). The fraud tasks reward detecting common schemes, not carrying them out. During development, one model escaped an insufficiently isolated sandbox and read the grader code (Appendix D). The final runs use isolated sandboxes without internet access, and we report the incident so that others building agent benchmarks can guard against it.
REPRODUCIBILITY STATEMENT
Two artifacts are released. The warehouse of one seed is released on Hugging Face under CC BY 4.0 at the linked source. The code that re-produces the paper’s runs is released under the Apache License 2.0 at the linked source TextQLLabs/Argo-Bench. It contains the 210 questions, the reference agent on Inspect AI 0.3.263, its warehouse and Python tool servers, the mission-control console through which the agent files, the three sandboxes in which runpython executes (a macOS Seatbelt profile, a Docker image with pinned libraries, and the network-less gVisor pod used for the paper’s runs), loaders for DuckDB and BigQuery, and the model and reasoning-effort configuration for every rung reported. The system prompt and tools are given in Appendices E and F, and Appendix L shows three refer-ence solutions in full.
The simulator and the graders are not released to prevent direct answer memorization. A run made with the released code exports a submission file recording every filing, the model and its settings, to-ken usage, and how each run ended, which we score on a best-effort basis. The paper’s runs queried a second world generated by the same simulator from a different seed (Section 2.3). The 15 ques-tions that name specific couriers, storefronts, or promotion codes draw them from the public world by the same selection rule. Re-running the released code therefore reproduces the paper’s procedure on a sibling world rather than its exact numbers. Proprietary models were accessed through their providers’ APIs, and all other models through Fireworks serverless endpoints with default settings in September 2026. Results from proprietary APIs may drift as providers update their models.
ACKNOWLEDGMENTS.
We thank Angela Peng, Alexander Baumstark, and Ben Van Sleen for their work on the design of the benchmark and on the evaluations, and JS Irick, David Dixon, and Scott Cairncross for build-ing the ERP, FP&A, and reporting components of the simulator. We also thank Mark Hay, Ben Mains, Matthew Abate, Sergi Domingo, and our colleagues at TextQL for their feedback and sup-port, the New York City agencies whose public reports and records the simulator is built on, and the maintainers of Inspect AI.
R EFERENCES https: Al Jazeera. Delivery driver pleads guilty to stealing $2.5m from DoorDash. //the linked source, 2025. 2025-05-14.
Erik Altman, Jovan Blanuša, Luc von Niederhäusern, Béni Egressy, Andreea Anghel, and Kubilay Atasu. Realistic synthetic financial transactions for anti-money laundering models. In Advances in Neural Information Processing Systems, volume 36, pp. 29851–29874, 2023.
Anthropic. Introducing the Model Context Protocol. the linked source model-context-protocol, 2024.
Axel Backlund and Lukas Petersson. Vending-Bench: A benchmark for long-term coherence of autonomous agents. arXiv preprint arXiv:2502.15840, 2025.
Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and Karthik Narasimhan. τ 2-Bench: Eval-uating conversational agents in a dual-control environment. arXiv preprint arXiv:2506.07982, 2025.
Kratika Bhagtani, Kusha Sridhar, Maziyar Baran Pouyan, Yuying Zhao, and Eugene Siow. ERP-Bench: A state-grounded evaluation paradigm for computer-use agents in enterprise software. arXiv preprint arXiv:2609.17885, 2026.
Nikos I Bosse, Sam Abbott, Anne Cori, Edwin van Leeuwen, Johannes Bracher, and Sebastian Funk. Scoring epidemiological forecasts on transformed scales. PLoS Computational Biology, 19: e1011393, 2023.
Johannes Bracher, Evan L Ray, Tilmann Gneiting, and Nicholas G Reich. Evaluating epidemic forecasts in an interval format. PLoS Computational Biology, 17:e1008618, 2021.
Lars Brehm, Armin Heinzl, and M Lynne Markus. Tailoring ERP systems: a spectrum of choices and their implications. In Proceedings of the 34th Annual Hawaii International Conference on System Sciences. IEEE, 2001.
Lizette Chapman and Kartikay Mehrotra. Instacart shoppers say they are battling order grabbing bots that cut their profits. the linked source, 2020. Bloomberg via Fortune, 2020-08-01.
Junqiao Chen, David Chun, Milesh Patel, Epson Chiang, and Jesse James. The validity of synthetic clinical data: a validation study of a leading synthetic data generator (Synthea) using clinical quality measures. BMC Medical Informatics and Decision Making, 19:44, 2019.
Peter Baile Chen, Devin Yang, Weiyue Li, Fabian Wenz, Yi Zhang, Nesime Tatbul, Michael Ca-farella, Ça ̆gatay Demiralp, and Michael Stonebraker. BEAVER: An enterprise benchmark for Text-to-SQL. arXiv preprint arXiv:2409.02038, 2024.
Yeounoh Chung, Gaurav T. Kakkar, Yu Gan, Brenton Milne, and Fatma Özcan. Is long context all you need? Leveraging LLM’s extended context for NL2SQL. Proceedings of the VLDB Endowment, 18:2735–2747, 2025.
Thomas H Davenport. Putting the enterprise into the enterprise system. Harvard Business Review, 76:121–131, 1998.
https: DoorDash, Inc. DoorDash further strengthens safeguards against account sharing. //about.doordash.com/en-us/news/doordash-further-strengthens-safeguards-against-account-sharing, 2024a. 2024-12-12.
Alex Egg, Martin Iglesias Goyanes, Friso Kingma, Andreu Mora, Leandro von Werra, and Thomas Wolf. DABstep: Data agent benchmark for multi-step reasoning. arXiv preprint arXiv:2506.23719, 2025.
Charles Elkan. The foundations of cost-sensitive learning. In Proceedings of the Seventeenth In-ternational Joint Conference on Artificial Intelligence (IJCAI), pp. 973–978. Morgan Kaufmann, 2001.
Andrea Gadotti, Luc Rocher, Florimond Houssiau, Ana-Maria Cre ̧tu, and Yves-Alexandre de Mon-tjoye. Anonymization: The imperfect science of using data while preserving privacy. Science Advances, 10:eadn7053, 2024. doi: 10.1126/sciadv.adn7053.
Ahmad Ghazal, Tilmann Rabl, Minqing Hu, Francois Raab, Meikel Poess, Alain Crolotte, and Hans-Arno Jacobsen. BigBench: Towards an industry standard benchmark for big data analytics. In Proceedings of the 2013 ACM SIGMOD International Conference on Management of Data, pp. 1197–1208, 2013.
Tilmann Gneiting. Quantiles as optimal point forecasts. International Journal of Forecasting, 27:197–207, 2011.
Tilmann Gneiting and Adrian E Raftery. Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102:359–378, 2007.
Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, and Or Itzahary. The Era by Eon benchmark: A generated enterprise estate with exact ground truth for benchmark-ing LLM agents. arXiv preprint arXiv:2609.09853, 2026.
Ken Gu, Ruoxi Shang, Ruien Jiang, Keying Kuang, Richard-John Lin, Donghe Lyu, Yue Mao, Youran Pan, Teng Wu, Jiaqian Yu, Yikun Zhang, Tianmai M. Zhang, Lanyi Zhu, Mike A Merrill, Jeffrey Heer, and Tim Althoff. BLADE: Benchmarking language model agents for data-driven science. In Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 13936– 13971, 2024.
https: Bernadette Heier. Uber Eats to remove thousands of duplicate virtual brands. //foodondemand.com/03302023/uber-eats-to-remove-thousands-of-duplicate-virtual-brands/, 2023. Food On Demand, 2023-03-30.
Kung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat, Divyansh Agarwal, Prafulla Kumar Choubey, Yixin Mao, Silvio Savarese, Caiming Xiong, and Chien-Sheng Wu. CRMArena-Pro: Holistic assessment of LLM agents across diverse business scenarios and interactions. Transac-tions on Machine Learning Research, 2026.
Rob J Hyndman and George Athanasopoulos. Forecasting: principles and practice. OTexts, 3rd edition, 2021.
Incognia. Incognia mobile app fraud insights report reveals food delivery apps are major target for location-based fraud. the linked source, 2022. 2022-08-30.
Incognia. Incognia’s gig economy fraud report shows refund abuse representing 48% of consumer the linked source in 2024. report-shows-refund-abuse-representing-48-percent-of-consumer-fraud-in-2024, 2025. 2025-02-26.
Tengjun Jin, Yuxuan Zhu, and Daniel Kang. ELT-Bench: An end-to-end benchmark for evaluating AI agents on ELT pipelines. Proceedings of the VLDB Endowment, 19:84–98, 2025.
Tengjun Jin, Yoojin Choi, Yuxuan Zhu, and Daniel Kang. Pervasive annotation errors break Text-to-SQL benchmarks and leaderboards. arXiv preprint arXiv:2601.08778, 2026.
Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao, Wenhao Yu, Kaixin Ma, Hongming Zhang, Xinya Du, and Dong Yu. DSBench: How far are data science agents from becoming data science experts? In International Conference on Learning Representations, volume 2025, pp. 32597–32649, 2025.
Sean Kandel, Andreas Paepcke, Joseph M Hellerstein, and Jeffrey Heer. Enterprise data analysis and visualization: An interview study. IEEE Transactions on Visualization and Computer Graphics, 18:2917–2926, 2012.
Darek Kłeczek, Fuheng Zhao, Alexander W. Lee, Julien Tissier, Paweł Liskowski, U ̆gur Çetintemel, and Anupam Datta. AvalancheBench: Evaluating enterprise data agents through latent world recovery. arXiv preprint arXiv:2605.24183, 2026.
Eugenie Lai, Gerardo Vitagliano, Ziyu Zhang, Om Chabra, Sivaprasad Sudhir, Anna Zeng, Anton Zabreyko, Chenning Li, Ferdi Kossmann, Jialin Ding, Jun Chen, Markos Markakis, Matthew Russo, Weiyang Wang, Ziniu Wu, Mike Cafarella, Lei Cao, Samuel Madden, and Tim Kraska. KramaBench: A benchmark for AI systems on data-to-insight pipelines over data lakes. In Inter-national Conference on Learning Representations, volume 2026, pp. 142883–142912, 2026.
Fangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao, Dongchan Shin, Hongjin Su, Zhaoqing Suo, Hongcheng Gao, Wenjing Hu, Pengcheng Yin, Victor Zhong, Caiming Xiong, Ruoxi Sun, Qian Liu, Sida Wang, and Tao Yu. Spider 2.0: Evaluating language models on real-world enterprise Text-to-SQL workflows. In International Conference on Learning Representations, volume 2025, pp. 28691–28735, 2025.
Fangyu Lei, Jinxiang Meng, Yiming Huang, Junjie Zhao, Yitong Zhang, Jianwen Luo, Xin Zou, Ruiyi Yang, Wenbo Shi, Yan Gao, Shizhu He, Jun Zhao, Zuo Wang, Qian Liu, Yang Wang, Ke Wang, and Kang Liu. DAComp: Benchmarking data agents across the full data intelligence lifecycle. In International Conference on Learning Representations, volume 2026, pp. 104463– 104501, 2026.
Viktor Leis, Andrey Gubichev, Atanas Mirchev, Peter Boncz, Alfons Kemper, and Thomas Neu-mann. How good are query optimizers, really? Proceedings of the VLDB Endowment, 9: 204–215, 2015.
Boyan Li, Yiran Peng, Yupeng Xie, Sirong Lu, Yizhang Zhu, Xing Mu, Xinyu Liu, and Yuyu Luo. DeepEye: A steerable self-driving data agent system. In Companion of the International Conference on Management of Data, pp. 74–77, 2026.
Jiacheng Li, Jingbo Shang, and Julian McAuley. UCTopic: Unsupervised contrastive learning for phrase representations and topic mining. In Proceedings of the 60th Annual Meeting of the Asso-ciation for Computational Linguistics (Volume 1: Long Papers), pp. 6159–6169, 2022.
Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, Xuanhe Zhou, Chenhao Ma, Guoliang Li, Kevin Chang, Fei Huang, Reynold Cheng, and Yongbin Li. Can LLM already serve as a database interface? A BIg bench for large-scale database grounded Text-to-SQLs. In Advances in Neural Information Processing Systems, volume 36, pp. 42330–42357, 2023.
Shu Liu, Soujanya Ponnapalli, Shreya Shankar, Sepanta Zeighami, Alan Zhu, Shubham Agarwal, Ruiqi Chen, Samion Suwito, Shuo Yuan, Ion Stoica, Matei Zaharia, Alvin Cheung, Natacha Crooks, Joseph E. Gonzalez, and Aditya G. Parameswaran. Supporting our AI overlords: Re-designing data systems to be agent-first. In Conference on Innovative Data Systems Research (CIDR), 2026.
Edgar Lopez-Rojas, Ahmad Elmir, and Stefan Axelsson. PaySim: A financial mobile money sim-ulator for fraud detection. In 28th European Modeling and Simulation Symposium (EMSS), pp. 249–255, 2016.
Khalil Maycock. Jacksonville restaurant loses thousands after DoorDash account hacked. the linked source, 2024. News4JAX, 2024-12-02.
Raghunath Othayoth Nambiar and Meikel Poess. The making of TPC-DS. In International Confer-ence on Very Large Data Bases (VLDB), pp. 1049–1058, 2006.
NYC Department of Consumer and Worker Protection (DCWP). Restaurant delivery app data: Quar-terly reports, Q1–Q4 2024. the linked source, 2024a. Accessed 2026-09.
NYC Department of Consumer and Worker Protection (DCWP). Mayor Adams announces first annual increase in minimum pay rate for app-based restaurant delivery workers. the linked source, 2024b.
Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, and Xiaoying Xing. Business Arena: Benchmarking LLM agents in a realistic marketplace. arXiv preprint arXiv:2608.08621, 2026.
Hasso Plattner. The impact of columnar in-memory databases on enterprise systems: implications of eliminating transaction-maintained aggregates. Proceedings of the VLDB Endowment, 7: 1722–1729, 2014.
Gaurav Sahu, Abhay Puri, Juan A. Rodriguez, Amirhossein Abaskohi, Mohammad Chegini, Alexandre Drouin, Perouz Taslakian, Valentina Zantedeschi, Alexandre Lacoste, David Vazquez, Nicolas Chapados, Christopher Pal, Sai Rajeswar, and Issam Laradji. InsightBench: Evaluating business analytics agents through multi-step insight generation. In International Conference on Learning Representations, volume 2025, pp. 4683–4715, 2025.
Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger. Defining and character-izing reward gaming. In Advances in Neural Information Processing Systems, volume 35, pp. 9460–9471, 2022.
Yuda Song, Hanlin Zhang, Carson Eisenach, Sham Kakade, Dean Foster, and Udaya Ghai. Mind the gap: Examining the self-improvement capabilities of large language models. In International Conference on Learning Representations, volume 2025, pp. 39894–39931, 2025.
Zhaoyan Sun, Jiayi Wang, Xinyang Zhao, Jiachi Wang, and Guoliang Li. Data agent: A holistic architecture for orchestrating data+AI ecosystems. arXiv preprint arXiv:2507.01599, 2025.
Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, and Niranjan Balasubramanian. AppWorld: A controllable world of apps and people for benchmarking interactive coding agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 16022– 16076, 2024.
UK AI Security Institute. Inspect AI: Framework for large language model evaluations. https: //github.com/UKGovernmentBEIS/inspectai, 2024.
Alexander van Renen, Dominik Horn, Pascal Pfeil, Kapil Vaidya, Wenjian Dong, Murali Narayanaswamy, Zhengchun Liu, Gaurav Saxena, Andreas Kipf, and Tim Kraska. Why TPC is not enough: An analysis of the Amazon Redshift fleet. Proceedings of the VLDB Endowment, 17:3694–3706, 2024.
Adrian Vogelsgesang, Michael Haubenschild, Jan Finis, Alfons Kemper, Viktor Leis, Tobias Mühlbauer, Thomas Neumann, and Manuel Then. Get real: How benchmarks fail to represent the real world. In Proceedings of the Workshop on Testing Database Systems, pp. 1–6, 2018.
Jason Walonoski, Mark Kramer, Joseph Nichols, Andre Quina, Chris Moesel, Dylan Hall, Carlton Duffett, Kudakwashe Dube, Thomas Gallagher, and Scott McLachlan. Synthea: An approach, method, and software mechanism for generating synthetic patients and the synthetic electronic health care record. Journal of the American Medical Informatics Association, 25:230–238, 2018.
https: Zack Whittaker. DoorDash customers say their accounts have been hacked. //techcrunch.com/2018/09/25/doordash-customers-say-their-accounts-have-been-hacked, 2018. TechCrunch, 2018-09-25.
Niklas Wretblad, Fredrik Riseby, Rahul Biswas, Amin Ahmadi, and Oskar Holmström. Under-standing the effects of noise in Text-to-SQL: An examination of the BIRD-Bench benchmark. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 356–369, 2024.
Frank F. Xu, Yufan Song, Boxuan Li, Yuxuan Tang, Kritanjali Jain, Mengxue Bao, Zora Wang, Xuhui Zhou, Zhitong Guo, Murong Cao, Mingyang Yang, Hao Yang Lu, Amaad Martin, Zhe Su, Leander Maben, Raj Mehta, Wayne Chi, Lawrence Jang, Yiqing Xie, Shuyan Zhou, and Graham Neubig. TheAgentCompany: Benchmarking LLM agents on consequential real world tasks. In Advances in Neural Information Processing Systems, volume 38, 2025a.
Muxi Xu, Kun Hu, Sudeep Das, and Bruce Wang. Causal machine learning for promotions: Industry evidence and applications. In KDD Workshop on Causal Inference and Machine Learning in Practice, 2025b.
An Yan, Zhankui He, Jiacheng Li, Tianyang Zhang, and Julian McAuley. Personalized showcases: Generating multi-modal explanations for recommendations. In Proceedings of the 46th Inter-national ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 2251–2255, 2023.
Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. τ -bench: A benchmark for tool-agent-user interaction in real-world domains. In International Conference on Learning Rep-resentations, volume 2025, pp. 9965–10017, 2025.
Yelp Inc. Yelp Open Dataset. the linked source, 2026. Accessed 2026-09.
Ethan G Young, Pengfei Zhu, Tyler Caraza-Harter, Andrea C Arpaci-Dusseau, and Remzi H Arpaci-Dusseau. The true cost of containing: A gVisor case study. In 11th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 19), 2019.
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and Text-to-SQL task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 3911–3921, 2018.
Xinran Zhang, Pengrui Lu, Lyumanshan Ye, and Pengfei Liu. ERPBench: Evaluating LLM agents for enterprise decision-making across competitive market ecologies. arXiv preprint arXiv:2609.04667, 2026.
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph Gonzalez, and Ion Stoica. Judg-ing LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems, volume 36, pp. 46595–46623, 2023.
A D ONOR DATASETS
Table 3 lists 34 public datasets and reports that were used to build the world (Section 2.1), catego-rized by the part of the world they support. Donor datasets released for research or non-commercial use, among them the Yelp Open Dataset (Yelp Inc., 2026) and the Google Local reviews, serve only as shape donors: none of their records reach the world or the released warehouse, only parameters fitted on them.
B QUESTIONS ENABLED BY THE MINIMUM - PAY CHANGE
On April 1, 2024, NYC raised the minimum pay rate for app-based restaurant delivery workers from $17.96 to $19.56 per hour, and the simulated platform responds on the same day. Its dispatcher starts restricting when couriers can go online, and its per-delivery pay card drops
(Section 2.1). Whenever a week’s pay falls short of the minimum, the platform tops it up with a true-up on the weekly courier pay invoice. The shock breaks the stationarity that simple forecasts rely on, and it changes how several standard figures must be computed. Table 4 groups the tasks that depend on it.
C WAREHOUSE TABLES
Table 5 lists every table in the warehouse with its column and row counts, grouped by the Oracle E-Business Suite module that owns it. The 159 standard tables fall under 15 Oracle products. The 76 custom extensions are grouped by the part of the business they record.
D R EWARD H ACKING
In an early round of runs, the Python sandbox did not isolate the agent from the evaluation host. In three runs of fraud-04, Gemini 3.8 Flash changed its root directory, located the grader for the task, and filed the answer it read there, scoring 0.99. We searched the traces of every run in that round for the same behavior and found it in no other model. The round was discarded. Every run reported in this paper executes Python in a network-less gVisor pod under Inspect’s Kubernetes sandbox provider on GKE, and each run’s BigQuery service account can read only the dataset scoped to its task. The relevant agent traces are included in the GitHub release.
Table 5, continued.
E S YSTEM PROMPT
You are an analyst for a food delivery platform operating in NYC. You have read-only access to a transformed export of the company's Oracle E-Business Suite warehouse.
You should explore the warehouse, investigate, and then perform appropriate actions based on your findings.
Warehouse
One Oracle EBS instance, order-to-cash, covering the calendar year 2024, following Oracle house conventions.
Timestamps are naive America/NewYork: compare "as of 2024-07-31 23:59:59 ET" directly against stored values, without converting.
Mission Control
File every finding from Python (runpython) through the Mission Control console. A number stated only in prose is not filed. Example: python from missioncontrol import MissionControl, Reason missioncontrol = MissionControl
a list can mix bare ids with per-item mappings, with the call's reason= overriding bare ones missioncontrol.bancustomers([{"id": 101, "reason": Reason.PROMOFARMING}, 102], reason=Reason.CARDTESTING,) # finish with a summary missioncontrol.summary.
Endpoints
help(MissionControl.<endpoint>) shows its docstring and arguments. Pass the complete set in one call, endpoints are plural. A rejected call files nothing and says which item is wrong.
Enforcement
- bancustomers, bancouriers, banmerchants: Close customer accounts, deactivate couriers, remove merchants.
- holdpayouts: Freeze a supplier's payouts pending review.
- discontinuepromocodes: Stop promo codes from being redeemed again.
- flagrings: Report sets of actors believed to be operating together
- endcampaigns: Stop promotional campaigns.
Payments and pay
- remediatepayments: Refund, recapture, void or write off payments on orders.
- reportpayperiods: State couriers' statutory minimum-pay position per pay period.
- issuepayadjustments: Issue pay corrections to couriers.
- electmethods: Record which minimum-pay method was elected per period, with both costs.
Accounting
- reportbalances: State account balances as of an instant.
- postjournalentries: Post balanced journal entries.
- fileadjustments: Record reconciling differences against an account.
- reportrollforwards: File an account roll-forward, one item per period.
- reportagings: Break balances into buckets, by age or any other dimension.
- flagorders: Flag orders for an accounting or control defect.
- fileschedules: File the tables a close produces.
Analysis and planning
- reportcampaignperformances: Record measured campaign economics.
- reportmetrics: Any other figure you were asked for.
- fileforecasts: Forecast figures that are not yet knowable, as a point and an interval.
- filepolicies: Propose operating policies as code.
Dashboards
- datasourcecontracts: The data sources this workspace's dashboard tiles are waiting for.
- publishdatasources: Publish dashboard data sources. This endpoint validates your data source against the tile's contract.
Session
- note: Attach free-text rationale or caveats that are not a decision.
- status: What has been filed so far.
- summary: Finish with this.
Money is compared to the cent. Round only at the end. Never modify the warehouse.
F AGENT TOOLS
The reference agent is given four tools, served by two MCP servers. A read-only warehouse server runs over the task’s database engine, and a Python server holds a persistent in-terpreter in the agent’s sandbox. Filings to Mission Control go through missioncontrol.py, which is importable from runpython (Appendix E).
Limits. Each sandbox runs under gVisor with 4 GiB of memory (12 GiB for the few runs repeated after the sandbox ran out of memory), one CPU, and 10 GiB of ephemeral disk. Each run is limited to 500 model turns and 240 minutes. A runpython cell times out after 900 seconds and returns at most 30,000 characters of output, and a runsql query may scan at most 20 GiB, with no limit on the number of queries, and saves at most 1,000,000 rows. A run that reaches a limit, exhausts the model’s context window, or crashes its sandbox is graded on whatever it submitted before stopping. Of the 9,869 graded runs, 56 (0.6%) reached the turn limit, 43 of them from Muse Spark 1.3 and 13 from Gemini 3.8 Flash, and 65 (0.7%) reached the context window, all from Qwen 3.8 Max.
G TASK DEFINITION AND GRADING
Task definition. A task consists of a prompt, a scope, a list of expectations, and an answer key. The prompt is the only text the agent receives besides the system prompt (Appendix E). The scope sets the last month of 2024 in the warehouse the agent queries. Each expectation names one filing the grader requires, with its action (for instance bancouriers or fileforecasts), the keys it is filed against, such as vendor IDs or months, and its grading mode. The answer key is computed once, before any run, and frozen. It comes either from truth queries, which are SQL over the latent tables, or from a local builder, which reads the simulator’s labels directly, for instance which couriers were made to steal orders. Listing 1 gives both definitions with their defaults.
Each expectation scores between 0 and 100, a forecast between −200 and 100, and Task score. the task score is their weighted mean. Mixing binary, graded, and forecast scales in one mean is a choice, so Table 2 also reports the solved rate, which does not depend on how the scales are mixed, and per-domain means, and it floors each model’s forecast tasks at 0 as a block so that a few very wrong forecasts cannot outweigh the rest. Of the 210 tasks, 159 have one scored expectation, 37 have two, and 14 have three. An expectation marked optional is graded and reported but left out of the mean. If a run files nothing while the key expects at least one action, every expectation scores 0, and a forecast expectation −200. A filing that carries many items, such as one bancustomers call with 300 accounts, is split into its items before grading, so it scores the same as 300 separate filings.
IDs and keys are compared as normalized text: case and surrounding whitespace do not matter, and 6057444.0 is 6057444. A key filed inside a longer string, such as payout period P17, also counts, but only on a token boundary, so P17 does not match P170.
Forecasts. A forecast files a point and an 80% interval for each series and is scored by its weighted interval score (WIS), with the median and one interval, on a scale set by a reference forecast. For 76 of the 98 forecast expectations the reference is the no-change forecast from the months the agent can see: its point r is the last visible value and its interval is r ± zs √ with s = ˆσ h, where ˆσ is the root mean square of the one-step changes and h the number of periods ahead. The other 22 cross a regime change, such as the minimum-pay rule, where no change is not a serious benchmark, and their reference is frozen into the answer key as a point r and a standard deviation s at the horizon. Most are a figure the prompt attributes to the business and asks the agent to check, computed by the method the prompt states from the data the agent can see.
Table 7 lists all 22 and how each was set, including seven whose point was set with the outcome in view.