You’re listening to “Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows,” by Gabriel Tomitsuka and colleagues. Published in arXiv on October 1, 2026. A RGO -B ENCH: E VALUATING DATA AGENTS ON E NTERPRISE -S CALE W ORKFLOWS Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma TextQL {gabriel,arman,emma,duke,joseph}@textql.com A BSTRACT Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise ware-houses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an eval-uation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, mod-eled on the Oracle E-Business Suite schema. The simulator’s ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocat-ing courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solu-tion that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments. 1 I NTRODUCTION In recent years, data agents have advanced from answering simple text-to-SQL questions over small, well-documented schemas to performing long-horizon data science tasks, publishing durable company-wide assets such as dashboards, and acting on their findings, for instance, by adjusting promotions and banning fraudulent accounts. Rising scores on established data benchmarks could be read as a sign that enterprise data work is close to solved. The best Spider 2.0-Snow score has risen from 23.8% at release to 96.7%, and the top BIRD entry reaches 82.4% against a human 93.0%.1 However, current benchmarks primarily evaluate text-to-SQL performance and are not representa-tive of enterprise agentic data workflows in three important ways. First, they run on a patchwork of public data. Spider 2.0-Snow’s 547 questions sparsely cover a sprawling collection of 152 databases (Appendix H). Over a third of those databases are used for a single question only. Two thirds of the questions use BigQuery public datasets; another quarter use local sample databases. These public and local sources’ schemas are documented in tutorials and textbooks. Fur-thermore, there is a high degree of overlap among tables: at least two thirds of the tables are date-, geography-, or version-sharded copies of another table. In an enterprise warehouse, different mod-ules must agree on the same numbers, and one business event may touch ten or more tables. Second, they are not end-to-end. Many of the most valuable enterprise data science tasks cannot be done in SQL. Fitting a demand forecast, training a fraud model, or optimizing a courier sched-ule requires statistical, machine learning, and optimization libraries. Additionally, findings lead to decisions, which are ultimately judged by their return on investment. For instance, a promotion with a high redemption rate can still lose money if most of those orders would have been placed anyway. A gold answer cannot distinguish between such decisions. Third, they are not verifiable. On real data, a benchmark can only measure agreement with its annotators since writing the answer key requires solving the task. Even when the answer is in the data, annotators may miss it. In a recent audit, 62.8% of Spider 2.0-Snow’s released gold queries and 52.8% of BIRD Mini-Dev’s were found to be erroneous, most often because annotators misread the data or the schema, and correcting BIRD’s errors moved agents’ leaderboard standings by up to nine places. Furthermore, for many of the most valuable enterprise tasks, the answer may not be in the data at all. Fraud left undetected leaves no label, and the outcomes of a rejected decision are never observed. Large enterprises keep the records that financial planning, forecasting, and fraud detection require in enterprise resource planning (ERP) systems such as Oracle E-Business Suite (EBS), SAP S/4HANA, and Oracle Fusion Cloud, extended with custom tables representing id-iosyncrasies and workflows unique to the business. Because the same systems hold the company’s ledgers, payroll, and customer records, access to them is heavily restricted. For this reason, ERP data has remained practically unexplored in benchmarks to date. When such data is released, it must be anonymized, which can break the relational structure that answers depend on. In an early release of BEAVER, the closest attempt to date, Chung et al. (2025) report primary keys that violate their uniqueness constraints and questions whose gold query is null. We instead simulate a business and grade against the simulation’s ground truth (Figure 1). Argo-Bench models a food delivery platform in New York City in 2024, a three-sided marketplace whose economics are disclosed to the city every month. The simulator reflects a real minimum-pay increase in April 2024, and is calibrated to these disclosures and to public filings. None of its data is generated by a language model. We project this world into an Oracle EBS warehouse that omits the simulator’s latent state, so tasks are harder to solve than to verify. The agent must reconstruct facts from the warehouse, while the grader reads them off the state. Agents work in a sandboxed Python environment with statistics, machine-learning, and optimization libraries (Appendix F), and file their decisions and results through a mission-control interface rather than returning a query. Because the grader knows the latent state, it can grade a decision by its consequences. For instance, a list of banned accounts is scored by the fraud losses it prevents, including losses from fraud that the platform never detected, net of the revenue lost from wrongly banned customers. Simulated environments are an established practice in a broad range of domains, and synthetic warehouses have long been used to benchmark data systems. Even Spider 2.0 draws a fifth of its questions from synthetic, obfuscated, or textbook sample schemas (Appendix H). To our knowledge, no prior simulated bench-mark combines an enterprise-scale warehouse whose tables must agree with one another and a grader that scores the consequences of the agent’s actions against the simulator’s latent state. We evaluate 14 frontier and open-weight models on Argo-Bench. The strongest, Claude Opus 5.5, solves 34.8% of tasks and averages 59.5 points. Nine of the fourteen average below 35. Models often analyze the wrong quantity or optimize the wrong objective. We make the following contributions: • A synthetic, large-scale public ERP dataset in Oracle EBS format for a 2024 New York City food delivery platform, comprising 235 mutually constraining tables, 81 million orders, 3.4 million active customers, and 7.5 billion rows, in which an order resolves into dispatch decisions, courier pay, merchant payouts, and balanced general-ledger journals. • A grader that scores the consequences of an action rather than the correctness of a query, using the simulator’s latent state as ground truth, for instance, to score bans by the fraud losses they prevent and forecasts against held-out months, supporting tasks across fraud detection, forecasting, and financial planning. • A benchmark of 210 such tasks, from publishing a dashboard data source to fitting forecasts and banning fraudulent accounts, with an evaluation of 14 frontier and open-weight models. Argo-Bench is public. One world’s warehouse is released on Hugging Face under CC BY 4.0,2 and the tasks, reference agent, tool server, and sandboxes under the Apache License 2.0.3 The numbers and experiments in this paper use a second world from a private seed, with different customers, couriers, and answer keys. The leaderboard is scored only on this world, so a score cannot be earned by memorizing the released warehouse. A demo at the linked source readers to browse the orders and deliveries of the released world at a 10% scale. 2 B ENCHMARK C ONSTRUCTION 2.1 WORLD SIMULATION We simulate a food delivery platform in New York City (NYC) in 2024, similar to DoorDash, Grub-hub, or Uber Eats. Public data on such platforms is aggregate. The city’s quarterly reports and the platforms’ own filings give totals, but order-level records cannot be released without exposing the platform’s customers, couriers, and margins. The world is therefore built from 34 public datasets and reports, each lending one mechanism (Appendix A). Uber and Lyft trips, for instance, give the time to drive between two zones at a given hour, and MenuStat the menus for restaurant chains. Donors play one of three roles. Identity donors are public records of real entities in the city, such as its 45,834 restaurants, 1.07 million addresses, and 260 taxi zones, and enter the world as they are, except that the released warehouse renames some restaurants (see the ethics statement). Shape donors are measured elsewhere, on other people, in another city, or in another year, and lend the world only a distribution. Anchors are published totals that the world is calibrated to reproduce but never samples records from. We chose food delivery in NYC because its economics are unusually well documented. Delivery apps must report their monthly orders, consumer spending, merchant fees, courier earnings, pro-ductivity, and hours worked to the NYC Department of Consumer and Worker Protection (DCWP), which publishes them quarterly, and the 10-K filings of DoorDash and Grubhub give the shape of a platform’s balance sheet (DoorDash, Inc., 2024b; 2025; Grubhub Inc., 2021). Following Walonoski et al. (2018), we calibrate the world to these anchors, sized as a dominant platform. The result has 81 million orders, about 55% of the 148 million deliveries that apps re-ported to the DCWP for 2024, and 3.4 million active customers, and it must balance incentives on three sides: quests and suggested pay for couriers, promotions and surge pricing for customers, and co-funded campaigns for merchants. Its per-delivery economics stay within 5% of the DCWP’s figures in 13 of 16 quarterly comparisons (Figure 2a). We selected 2024 because it contains a real extrinsic shock to these economics. NYC began enforc-ing a minimum pay rate of $17.96 per hour before tips for app-based restaurant delivery workers in December 2023 and raised it to $19.56 on April 1, 2024. We model the platforms’ response with a new courier scheduler that activates on that date, after which courier pay runs 6–9% above the anchor (Figure 2a). The levers our platform uses may differ from those of the real plat-forms, but the aggregate effect is the same in direction and, to within 9%, in size, and because the world absorbs the same shock, we can ask realistic forecasting questions about it (Appendix B). Additionally, we insert fraud patterns that public evidence shows are major problems for delivery platforms. Couriers steal orders after pickup, spoof their GPS, grab offers with bots, or rent out their accounts (DoorDash, Inc., 2024a). On the customer side, rings of new accounts farm promotions (DoorDash, Inc., 2025; Incog-nia, 2025), and stolen cards fund account takeovers and bust-outs (DoorDash, Inc., 2025; Whittaker, 2018). Storefronts may be shells or collude with couriers or regular customers on refunds (DoorDash, Inc., 2023), and their payouts can be diverted to changed bank accounts. Each pattern is calibrated both to how separable real card fraud is and to how often honest customers share a device, an address, or a card, since the latter sets a detector’s precision (Ap-pendix A). The simulator generates the world from the donors, calibrates it to the anchors, inserts the fraud, and projects the result to the EBS format (Section 2.2). 2.2 WAREHOUSE D ESIGN AND VALIDATION The simulator’s last step projects the world into what an analyst actually sees: the analytics export of a greenfield Oracle E-Business Suite (EBS) 12.2 instance. We chose EBS because its data model is publicly documented, so our schema can be verified against a reference. We designed the schema with three ERP consultants who have 14 to 31 years of experience. Every standard table and column in our warehouse exists in the data dictionary of Oracle’s EBS 12.2 Vision instance, the demo environment Oracle provides as a reference. Our tables, however, carry on average 52% of the columns of their Vision counterparts. EBS serves every industry, and analytics exports omit the columns a business does not use, here unused flexfields (a third of the omitted columns) and features such as shipping, inventory, foreign currency, and withholding tax. These columns would be empty in this business’s data, so no task loses information by their omission. The warehouse does not contain data drift or inconsistencies, such as deprecated tables that overlap active ones or figures that fail to reconcile across tables. These inconsistencies sometimes accumu-late in real data warehouses over years of migrations and acquisitions. Though the consultants named this the most significant difference from their customers’ systems, we deliberately chose to keep this discrepancy. By doing so, we keep the ground truth unambiguous: if a legacy table were to disagree with an active one, the correct answer would depend on undocu-mented conventions. Mature warehouses compensate for their idiosyncrasies with semantic layers, data models, and institutional knowledge. While we could provide such a layer alongside a more realistic messy warehouse, this would shift the evaluation’s focus to testing an abil-ity to use a curated layer. We instead aim to test an understanding of enterprise data organization: production workloads reuse only a few dozen combinations of hundreds of tables (van Renen et al., 2024), and an agent new to a warehouse must discover which ones matter by deciding what and how much of the warehouse to explore. A greenfield warehouse isolates this skill of understanding and exploring enterprise data organization. Otherwise, the consultants found the schema realistic, with two further omissions. It records no foreign-currency transactions, since no donor dataset covers the currencies visitors pay in, and it has only 43 balance sheet accounts (Section 5). 2.3 TASK D ESIGN Argo-Bench contains 210 tasks in five business areas (Figure 3a). Trust and safety tasks are en-forcement, where the agent finds fraud and abuse and acts on the accounts involved. FP&A tasks forecast unit economics and rebuild finance dashboards, marketplace tasks forecast courier supply and allocate budgets such as courier bonuses, accounting tasks report final values after the fact, and growth tasks measure, forecast, and publish the results of promotions and memberships. Each task has four parts (Appendix G). The prompt states the problem as a stakeholder would. The scope sets the last month of 2024 visible to the agent. The expectations list the filings the grader requires, each with its action, keys, and grading mode. The answer key is frozen before any run and comes from SQL over the latent tables or from the simulator’s own labels, such as which couriers stole orders. Because the world is simulated, these labels are exact and need no anonymization. Prompts cover a range of writing styles and levels of detail. Some reference the grading criteria or the exact output expected, while others are more subtle. This reflects how real users pose data questions: loosely, as high-level business questions, and in no set style or template. It also tests the skill of exploring data organization, since a less detailed prompt leaves the agent to discover which tables and conventions the question depends on. Several scenarios come in variants that differ only in such detail, and Appendix I compares them. Unlike most data science and analytics benchmarks, Argo-Bench does not ask agents to return a query. Agents file actions to a mission-control interface through a Python library in their sandbox (Appendix E). This design has four advantages. First, it permits advanced data science tasks which require machine learning, operations research, and mathematical optimization libraries to complete. Second, it supports end-to-end workflows. Emitting the correct SQL is not enough in practice, as real tasks require taking actions, e.g., rebalancing courier incentives across zones and hours, banning a set of users who are likely committing fraud, or holding the payouts of a suspicious merchant. Third, filings are explicit declarations of intent. When benchmarks compare SQL results, it is hard to determine whether a close number is the agent’s answer or an intermediate result, whereas a filing states the value, interval, or reason the agent commits to, which also makes partial credit well defined. Fourth, grading is objective, unlike the LLM-as-a-judge evaluation used in many related works. Of the 210 tasks, 178 act on accounts, file a forecast, allocate a budget, or publish a dashboard data source, and 99 see only up to a cutoff month, as an analyst would at that date, so forecasts are graded on months the agent has not seen (Figure 3). Prompts average 159 words, 51 tasks require more than one filing, and the reference solution in Figure 1 joins six tables across three EBS modules to recover the minimum-pay rule and fit an interval. In size, Argo-Bench matches long-horizon agent benchmarks such as TheAgentCompany (175 tasks), τ -bench, KramaBench, and ELT-Bench. Each expectation is scored from 0 to 100 (a forecast from −200) by one of nine grading modes (Appendix G). Forecasts are scored by their weighted interval score on a scale set by a reference forecast fixed before the outcome, so that filing one’s true median and interval is the best strategy, ban lists by the cost they save relative to banning nobody or everybody, budget allocations by the share of the attainable savings that the simulator realizes, and data sources and reported figures by their values. A task’s score is the mean of its expectations’ scores, weighted as the task specifies. A run that files nothing where the key expects action scores zero on every expectation (−200 on a forecast, the lowest a forecast can score). We release the warehouse of one world on Hugging Face and keep a second, generated from a private seed, for official grading. We ran our experiments on BigQuery (941 GB uncompressed), and the released warehouse is a set of 1,219 Parquet files (76.5 GB), with a dataset card giving setup instructions for BigQuery, DuckDB, Snowflake, Trino, Delta Lake, and Iceberg. The two seeds share the simulator and its calibration, but every ID, customer, restaurant, and courier differs.4 Table 1 compares Argo-Bench with prior benchmarks. 3 E VALUATION We run all experiments on Inspect AI 0.3.263, an open-source evaluation framework, on Google Kubernetes Engine. Each agent works in its own gVisor sandbox with 25 preinstalled Python libraries, an empty file system, and no out-bound internet access, and files to its own mission-control instance through a Python library (missioncontrol.py) that emulates a company’s internal tooling. A run ends after 500 model turns, to stop models that loop without progress, and Appendix F lists the other limits. We evalu-ate models available in September 2026 across price ranges (Table 2), calling open-weight models through Fireworks serverless endpoints and the others through their developers’ APIs, with default sampling settings for every model. 3.1 R ESULTS Claude Opus 5.5 leads overall and in three of the five domains (Table 2), GPT-6 Astra leads in forecasting, and Claude Sonnet 5.5 leads in compliance. More reasoning effort helps the GPT-6 and Claude models at every step, with diminishing returns for Opus, which gains 17.0 points from low to medium effort, 5.9 from medium to high, and 4.0 from high to extra-high, while Claude Sonnet 5.5 gains 14.9 and 11.5 over the last two steps. Gemini 3.8 Flash gains 8.6 points from low to medium [4.9, 16.3] and Qwen 3.8 Max gains 5.2, and neither gains detectably afterward. Muse Spark 1.3 changes by less than 2 points past medium, and DeepSeek V4.1 Flash moves only at extra-high (Appendix I). Gemini 3.8 Flash and Muse Spark 1.3 also average 197 and 250 model calls per task, compared to 82 for Opus, without scoring higher, and 57% of Muse’s warehouse spend goes to tasks on which it scores below 5 out of 100 (Appendix K). Many tasks are prompt variants of one scenario, so we resample the 146 base scenarios when boot-strapping. The resulting 95% confidence intervals are about ±2 to 8 points on the score and up to †Spider 2.0-Lite, as computed by Chen et al. (2024). ±8 points on the solved rate. Paired by task, Opus leads each of the next three models (GPT-6 Astra, Claude Sonnet 5.5, and GPT-6.1 Sol) by 7.7 to 10.0 points, and those three are not separated from one another. GPT-6.1 Sol leads its predecessor GPT-6 Sol by 12.8 points [6.8, 19.1], GPT-6 Sol leads Kimi K3 by 8.4 [0.5, 15.4], and the intervals of the next five overlap (Appendix I). 3.2 F INDINGS Many failing runs use sound methods but read the wrong record, optimize the wrong objective, or measure the wrong quantity. Passing runs check definitions against a second source (Appendix I). Wrong record. A marketing dashboard asks whether each discount offer paid off against the same push notification with the coupon left out for every thousand customers it was sent to, for members and non-members. Only Claude Opus 5.5, GPT-6 Astra, and GPT-6.1 Sol at extra-high effort publish the right table. What separates them is who counted as a member at the moment of targeting. A member whose card is declined keeps the benefits for seven days of grace and then loses them, but the contract stays on the books until it is canceled weeks later. Twenty-four of the failing runs take membership from the contract’s dates, and seven of them get every other figure right to within a few dollars. The three passing runs replay the billing history instead and check the result against the orders on which a member benefit was actually applied. Opus reads the seven-day grace off those orders, and GPT-6.1 Sol starts from the contract dates, finds 1,057 orders that its rule calls members’ and that received no benefit, and starts over. Wrong objective. One family of tasks asks where to cut $9.6 million from the annual budget for courier bonuses (quests). At extra-high effort and without hints, GPT-6 Astra finds the zones and hours where quests were randomly withheld, estimates supply responses with fixed effects and partial pooling, and cuts where quests buy the fewest courier-hours. The platform pays for quests to avoid surge pay; however, under the simulator’s response model, the plan loses $86,281 where a uniform cut would save $0.40 million, and it scores 0. From the same prompt, Claude Opus 5.5 at extra-high effort finds that quests substitute for surge pay, cuts where they save the least surge per bonus dollar, and saves $3.09 million of an attainable $3.12 million (score 99). Stating what quests are for and that a holdout exists raises the mean over all settings from 18.5 to 61.7. Wrong quantity. Many failures come from measuring a quantity other than the one the prompt asks for. Of the 47 completed runs of a task that sizes the courier location service, 46 miss all 12 monthly counts despite a 1% tolerance because the warehouse keeps only hourly idle check-ins while the app sends them every half hour. GPT-6 Astra notices the gap, writes that its count does not establish how many reports the app sent, and files it anyway. Only Claude Sonnet 5.5 at extra-high effort adds the missing half-hourly check-ins back. Across the dashboard tasks, 66.9% of the data sources that pass their structural contract score zero on their values. Overconfident forecasts. Across 4,553 forecast series from 3,346 runs on 72 tasks, nominal 80% intervals contain the realized value only 44.8% of the time. Apart from its floor, the grade is proper, so this overconfidence costs models points in expectation, although a per-series skill score clipped at zero would have rewarded a narrower interval in 72% of series (Appendix I). Grades also depend on the reference, which sets each series’ scale. Against the tighter reference of their harder variant, the June base-pay forecasts fall from a mean grade of 84.7 to 6.0, although their median absolute error is only 1.57%. Appendix I reports coverage and absolute error. 4 R ELATED WORK Text-to-SQL and data science benchmarks. Text-to-SQL benchmarks have moved from databases with a handful of tables each to enterprise-scale schemas and private data warehouses, and data science bench-marks extend evaluation to multi-step analysis, data lakes, and data pipelines. These benchmarks compare an agent’s output to a gold answer, which audits have found to be frequently wrong, or to expert conclusions, often scored by an LLM judge. Argo-Bench instead grades the actions that an agent files against the simulator’s latent state. Agent benchmarks in simulated environments. Simulated environments are the standard way to evaluate agents that act. AppWorld and τ -bench check the final state of the environment’s database. TheAgentCompany scores checkpoints in a simulated software company, and CRMArena-Pro shapes LLM-generated CRM records with latent variables. Vending-Bench and Business Arena score the net worth of a business that the agent runs. In these benchmarks, the state that determines success is either observable to the agent or changed by the agent inside a stylized game. In Argo-Bench, it is withheld and must be reconstructed from an enterprise warehouse. Simulators calibrated to public statistics have likewise supplied ground truth that real data lacks for patient records and money laundering. Generated enterprises with hidden ground truth. AvalancheBench uses an LLM judge to score how much of a small latent e-commerce world an agent’s report recovers. The Era by Eon benchmark serves a generated company through a fleet of 66 simulated products, calibrated to operational and published statistics, and plants the records that answer each of its read-only questions, together with near misses. Two unrelated benchmarks named ERPBench evaluate decisions in a simulated manufacturer and computer-use tasks in a live ERP system. In contrast, the latent state of Argo-Bench is produced by the simulation itself rather than being planted. Its evidence is spread across an enterprise warehouse of 7.49 billion rows, and agents are graded on the consequences of the actions they file rather than on the answers they give. 5 L IMITATIONS & F UTURE WORK The dataset still differs from the most complex ERP deployments in four ways. First, the dataset covers a single city, so it has no foreign currency and none of its representations as transaction, local, and reporting currency. Second, many complex warehouses combine several businesses, such as food delivery and grocery delivery, with shared accounts, such as driver payables and customer credits, mixing them in a single balance. Third, only one year is simulated. A longer history, such as 2017 to 2026, would add the market shocks of 2020 to 2022 to operations and forecasts. Fourth, only one ERP format is supported, and a natural extension would be to add SAP S/4HANA. The simulator validates 23 distinct metrics, and we keep behaviors that public figures do not con-strain out of scope for tasks. The in-world membership program in particular rests on weakly grounded assumptions. More generally, a simulator encodes the assumptions of its authors, and a generator that shares the simplifying assumptions of the systems under evaluation can make tasks easier than their real counterparts. Calibration to aggregate targets also does not guarantee realistic tails. Argo-Bench therefore compares data agents and does not estimate their performance on a real company’s warehouse. Each setting of the model and effort has one run per task, and all tasks share one simulated world, so our confidence intervals reflect the choice of tasks rather than run-to-run variation. Scores on some tasks are also sensitive to prompt wording, so Appendix I reports every hinted and unhinted pair, including one prompt we judged to be underspecified. Plan tasks are graded under the simulator’s frozen response model, and Appendix J lists open issues by task. 6 C ONCLUSION We introduced Argo-Bench, which evaluates data agents on a simulated food delivery platform exported to an Oracle E-Business Suite warehouse of 235 tables and 7.49 billion rows. Because the simulator’s latent state is withheld from the warehouse, Argo-Bench grades the facts agents reconstruct and the actions they file, with every answer key computed from ground truth. The strongest of 14 models solves 34.8% of tasks and averages 59.5 points, and its most instructive failures are careful analyses of the wrong quantity, toward the wrong objective, or read from the wrong record. We release the public seed’s warehouse, tasks, reference solutions, and harness, and hope that Argo-Bench helps measure progress toward data agents that can understand, navigate, and act within real enterprise data environments. AI USE STATEMENT Large language models were used in three ways. First, as coding assistants for the simulator, graders, reference solutions, and figure and table code. Second, in the simulator’s data work: tuning its parameters, cleaning donor datasets, and reconciling restaurant records across sources (Section 2.1). No record in the simulated world is generated by a language model, and every value in the warehouse comes from the simulator, except the fictional names of the storefronts renamed in the released warehouse, which a language model drafted (see the ethics statement). Third, in writing: some task prompts were drafted by a language model and then curated and rewritten by the authors, and language models edited the text of the paper and checked its citations. The authors checked every claim, number, and citation, and take full responsibility for the content of the paper. E THICS STATEMENT Argo-Bench contains no data about real people. Consumer and courier names are drawn from pub-lic name-frequency tables and assigned to simulated people, and every order, shift, and payment is simulated. Restaurants are real New York City businesses. Their names and addresses come from Overture Maps places matched to the city’s inspection records, with some names updated to the busi-ness’s current listing, and some storefronts the simulator opens during the year, such as relaunches and virtual brands, take the name of a real business that was not trading at the time. All behavior attributed to them in the world, including the fraud patterns of Section 2.1, is simulated. The sim-ulator draws which storefronts play a fraud role, not from any record of a business’s conduct, so a label says nothing about the real business. Even so, in the released warehouse every storefront that plays a fraud role in any scenario carries a fictional name instead of its real one, and fields derived from the name, such as its contact email, follow the new name. The released questions and agent transcripts use the same fictional names. A language model drafted the fictional names to read like real New York restaurant names, so that a renamed storefront does not stand out and point to the answers, and each was checked against the city’s inspection records and Overture Maps so that none is the name of a real restaurant. Addresses are unchanged, since the simulated geography depends on them. The donor datasets are used under their licenses. Those released for research or non-commercial use, such as the Yelp Open Dataset and the Grubhub MDRP instances, serve only as shape donors, and none of their records enters the world or the warehouse (Appendix A). The fraud tasks reward detecting common schemes, not carrying them out. During development, one model escaped an insufficiently isolated sandbox and read the grader code (Appendix D). The final runs use isolated sandboxes without internet access, and we report the incident so that others building agent benchmarks can guard against it. REPRODUCIBILITY STATEMENT Two artifacts are released. The warehouse of one seed is released on Hugging Face under CC BY 4.0 at the linked source. The code that re-produces the paper’s runs is released under the Apache License 2.0 at the linked source TextQLLabs/Argo-Bench. It contains the 210 questions, the reference agent on Inspect AI 0.3.263, its warehouse and Python tool servers, the mission-control console through which the agent files, the three sandboxes in which runpython executes (a macOS Seatbelt profile, a Docker image with pinned libraries, and the network-less gVisor pod used for the paper’s runs), loaders for DuckDB and BigQuery, and the model and reasoning-effort configuration for every rung reported. The system prompt and tools are given in Appendices E and F, and Appendix L shows three refer-ence solutions in full. The simulator and the graders are not released to prevent direct answer memorization. A run made with the released code exports a submission file recording every filing, the model and its settings, to-ken usage, and how each run ended, which we score on a best-effort basis. The paper’s runs queried a second world generated by the same simulator from a different seed (Section 2.3). The 15 ques-tions that name specific couriers, storefronts, or promotion codes draw them from the public world by the same selection rule. Re-running the released code therefore reproduces the paper’s procedure on a sibling world rather than its exact numbers. Proprietary models were accessed through their providers’ APIs, and all other models through Fireworks serverless endpoints with default settings in September 2026. Results from proprietary APIs may drift as providers update their models. ACKNOWLEDGMENTS. We thank Angela Peng, Alexander Baumstark, and Ben Van Sleen for their work on the design of the benchmark and on the evaluations, and JS Irick, David Dixon, and Scott Cairncross for build-ing the ERP, FP&A, and reporting components of the simulator. We also thank Mark Hay, Ben Mains, Matthew Abate, Sergi Domingo, and our colleagues at TextQL for their feedback and sup-port, the New York City agencies whose public reports and records the simulator is built on, and the maintainers of Inspect AI. R EFERENCES https: Al Jazeera. Delivery driver pleads guilty to stealing $2.5m from DoorDash. //the linked source, 2025. 2025-05-14. Erik Altman, Jovan Blanuša, Luc von Niederhäusern, Béni Egressy, Andreea Anghel, and Kubilay Atasu. Realistic synthetic financial transactions for anti-money laundering models. In Advances in Neural Information Processing Systems, volume 36, pp. 29851–29874, 2023. Anthropic. Introducing the Model Context Protocol. the linked source model-context-protocol, 2024. Axel Backlund and Lukas Petersson. Vending-Bench: A benchmark for long-term coherence of autonomous agents. arXiv preprint arXiv:2502.15840, 2025. Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and Karthik Narasimhan. τ 2-Bench: Eval-uating conversational agents in a dual-control environment. arXiv preprint arXiv:2506.07982, 2025. Kratika Bhagtani, Kusha Sridhar, Maziyar Baran Pouyan, Yuying Zhao, and Eugene Siow. ERP-Bench: A state-grounded evaluation paradigm for computer-use agents in enterprise software. arXiv preprint arXiv:2609.17885, 2026. Nikos I Bosse, Sam Abbott, Anne Cori, Edwin van Leeuwen, Johannes Bracher, and Sebastian Funk. Scoring epidemiological forecasts on transformed scales. PLoS Computational Biology, 19: e1011393, 2023. Johannes Bracher, Evan L Ray, Tilmann Gneiting, and Nicholas G Reich. Evaluating epidemic forecasts in an interval format. PLoS Computational Biology, 17:e1008618, 2021. Lars Brehm, Armin Heinzl, and M Lynne Markus. Tailoring ERP systems: a spectrum of choices and their implications. In Proceedings of the 34th Annual Hawaii International Conference on System Sciences. IEEE, 2001. Lizette Chapman and Kartikay Mehrotra. Instacart shoppers say they are battling order grabbing bots that cut their profits. the linked source, 2020. Bloomberg via Fortune, 2020-08-01. Junqiao Chen, David Chun, Milesh Patel, Epson Chiang, and Jesse James. The validity of synthetic clinical data: a validation study of a leading synthetic data generator (Synthea) using clinical quality measures. BMC Medical Informatics and Decision Making, 19:44, 2019. Peter Baile Chen, Devin Yang, Weiyue Li, Fabian Wenz, Yi Zhang, Nesime Tatbul, Michael Ca-farella, Ça ̆gatay Demiralp, and Michael Stonebraker. BEAVER: An enterprise benchmark for Text-to-SQL. arXiv preprint arXiv:2409.02038, 2024. Yeounoh Chung, Gaurav T. Kakkar, Yu Gan, Brenton Milne, and Fatma Özcan. Is long context all you need? Leveraging LLM’s extended context for NL2SQL. Proceedings of the VLDB Endowment, 18:2735–2747, 2025. Thomas H Davenport. Putting the enterprise into the enterprise system. Harvard Business Review, 76:121–131, 1998. https: DoorDash, Inc. DoorDash further strengthens safeguards against account sharing. //about.doordash.com/en-us/news/doordash-further-strengthens-safeguards-against-account-sharing, 2024a. 2024-12-12. Alex Egg, Martin Iglesias Goyanes, Friso Kingma, Andreu Mora, Leandro von Werra, and Thomas Wolf. DABstep: Data agent benchmark for multi-step reasoning. arXiv preprint arXiv:2506.23719, 2025. Charles Elkan. The foundations of cost-sensitive learning. In Proceedings of the Seventeenth In-ternational Joint Conference on Artificial Intelligence (IJCAI), pp. 973–978. Morgan Kaufmann, 2001. Andrea Gadotti, Luc Rocher, Florimond Houssiau, Ana-Maria Cre ̧tu, and Yves-Alexandre de Mon-tjoye. Anonymization: The imperfect science of using data while preserving privacy. Science Advances, 10:eadn7053, 2024. doi: 10.1126/sciadv.adn7053. Ahmad Ghazal, Tilmann Rabl, Minqing Hu, Francois Raab, Meikel Poess, Alain Crolotte, and Hans-Arno Jacobsen. BigBench: Towards an industry standard benchmark for big data analytics. In Proceedings of the 2013 ACM SIGMOD International Conference on Management of Data, pp. 1197–1208, 2013. Tilmann Gneiting. Quantiles as optimal point forecasts. International Journal of Forecasting, 27:197–207, 2011. Tilmann Gneiting and Adrian E Raftery. Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102:359–378, 2007. Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, and Or Itzahary. The Era by Eon benchmark: A generated enterprise estate with exact ground truth for benchmark-ing LLM agents. arXiv preprint arXiv:2609.09853, 2026. Ken Gu, Ruoxi Shang, Ruien Jiang, Keying Kuang, Richard-John Lin, Donghe Lyu, Yue Mao, Youran Pan, Teng Wu, Jiaqian Yu, Yikun Zhang, Tianmai M. Zhang, Lanyi Zhu, Mike A Merrill, Jeffrey Heer, and Tim Althoff. BLADE: Benchmarking language model agents for data-driven science. In Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 13936– 13971, 2024. https: Bernadette Heier. Uber Eats to remove thousands of duplicate virtual brands. //foodondemand.com/03302023/uber-eats-to-remove-thousands-of-duplicate-virtual-brands/, 2023. Food On Demand, 2023-03-30. Kung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat, Divyansh Agarwal, Prafulla Kumar Choubey, Yixin Mao, Silvio Savarese, Caiming Xiong, and Chien-Sheng Wu. CRMArena-Pro: Holistic assessment of LLM agents across diverse business scenarios and interactions. Transac-tions on Machine Learning Research, 2026. Rob J Hyndman and George Athanasopoulos. Forecasting: principles and practice. OTexts, 3rd edition, 2021. Incognia. Incognia mobile app fraud insights report reveals food delivery apps are major target for location-based fraud. the linked source, 2022. 2022-08-30. Incognia. Incognia’s gig economy fraud report shows refund abuse representing 48% of consumer the linked source in 2024. report-shows-refund-abuse-representing-48-percent-of-consumer-fraud-in-2024, 2025. 2025-02-26. Tengjun Jin, Yuxuan Zhu, and Daniel Kang. ELT-Bench: An end-to-end benchmark for evaluating AI agents on ELT pipelines. Proceedings of the VLDB Endowment, 19:84–98, 2025. Tengjun Jin, Yoojin Choi, Yuxuan Zhu, and Daniel Kang. Pervasive annotation errors break Text-to-SQL benchmarks and leaderboards. arXiv preprint arXiv:2601.08778, 2026. Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao, Wenhao Yu, Kaixin Ma, Hongming Zhang, Xinya Du, and Dong Yu. DSBench: How far are data science agents from becoming data science experts? In International Conference on Learning Representations, volume 2025, pp. 32597–32649, 2025. Sean Kandel, Andreas Paepcke, Joseph M Hellerstein, and Jeffrey Heer. Enterprise data analysis and visualization: An interview study. IEEE Transactions on Visualization and Computer Graphics, 18:2917–2926, 2012. Darek Kłeczek, Fuheng Zhao, Alexander W. Lee, Julien Tissier, Paweł Liskowski, U ̆gur Çetintemel, and Anupam Datta. AvalancheBench: Evaluating enterprise data agents through latent world recovery. arXiv preprint arXiv:2605.24183, 2026. Eugenie Lai, Gerardo Vitagliano, Ziyu Zhang, Om Chabra, Sivaprasad Sudhir, Anna Zeng, Anton Zabreyko, Chenning Li, Ferdi Kossmann, Jialin Ding, Jun Chen, Markos Markakis, Matthew Russo, Weiyang Wang, Ziniu Wu, Mike Cafarella, Lei Cao, Samuel Madden, and Tim Kraska. KramaBench: A benchmark for AI systems on data-to-insight pipelines over data lakes. In Inter-national Conference on Learning Representations, volume 2026, pp. 142883–142912, 2026. Fangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao, Dongchan Shin, Hongjin Su, Zhaoqing Suo, Hongcheng Gao, Wenjing Hu, Pengcheng Yin, Victor Zhong, Caiming Xiong, Ruoxi Sun, Qian Liu, Sida Wang, and Tao Yu. Spider 2.0: Evaluating language models on real-world enterprise Text-to-SQL workflows. In International Conference on Learning Representations, volume 2025, pp. 28691–28735, 2025. Fangyu Lei, Jinxiang Meng, Yiming Huang, Junjie Zhao, Yitong Zhang, Jianwen Luo, Xin Zou, Ruiyi Yang, Wenbo Shi, Yan Gao, Shizhu He, Jun Zhao, Zuo Wang, Qian Liu, Yang Wang, Ke Wang, and Kang Liu. DAComp: Benchmarking data agents across the full data intelligence lifecycle. In International Conference on Learning Representations, volume 2026, pp. 104463– 104501, 2026. Viktor Leis, Andrey Gubichev, Atanas Mirchev, Peter Boncz, Alfons Kemper, and Thomas Neu-mann. How good are query optimizers, really? Proceedings of the VLDB Endowment, 9: 204–215, 2015. Boyan Li, Yiran Peng, Yupeng Xie, Sirong Lu, Yizhang Zhu, Xing Mu, Xinyu Liu, and Yuyu Luo. DeepEye: A steerable self-driving data agent system. In Companion of the International Conference on Management of Data, pp. 74–77, 2026. Jiacheng Li, Jingbo Shang, and Julian McAuley. UCTopic: Unsupervised contrastive learning for phrase representations and topic mining. In Proceedings of the 60th Annual Meeting of the Asso-ciation for Computational Linguistics (Volume 1: Long Papers), pp. 6159–6169, 2022. Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, Xuanhe Zhou, Chenhao Ma, Guoliang Li, Kevin Chang, Fei Huang, Reynold Cheng, and Yongbin Li. Can LLM already serve as a database interface? A BIg bench for large-scale database grounded Text-to-SQLs. In Advances in Neural Information Processing Systems, volume 36, pp. 42330–42357, 2023. Shu Liu, Soujanya Ponnapalli, Shreya Shankar, Sepanta Zeighami, Alan Zhu, Shubham Agarwal, Ruiqi Chen, Samion Suwito, Shuo Yuan, Ion Stoica, Matei Zaharia, Alvin Cheung, Natacha Crooks, Joseph E. Gonzalez, and Aditya G. Parameswaran. Supporting our AI overlords: Re-designing data systems to be agent-first. In Conference on Innovative Data Systems Research (CIDR), 2026. Edgar Lopez-Rojas, Ahmad Elmir, and Stefan Axelsson. PaySim: A financial mobile money sim-ulator for fraud detection. In 28th European Modeling and Simulation Symposium (EMSS), pp. 249–255, 2016. Khalil Maycock. Jacksonville restaurant loses thousands after DoorDash account hacked. the linked source, 2024. News4JAX, 2024-12-02. Raghunath Othayoth Nambiar and Meikel Poess. The making of TPC-DS. In International Confer-ence on Very Large Data Bases (VLDB), pp. 1049–1058, 2006. NYC Department of Consumer and Worker Protection (DCWP). Restaurant delivery app data: Quar-terly reports, Q1–Q4 2024. the linked source, 2024a. Accessed 2026-09. NYC Department of Consumer and Worker Protection (DCWP). Mayor Adams announces first annual increase in minimum pay rate for app-based restaurant delivery workers. the linked source, 2024b. Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, and Xiaoying Xing. Business Arena: Benchmarking LLM agents in a realistic marketplace. arXiv preprint arXiv:2608.08621, 2026. Hasso Plattner. The impact of columnar in-memory databases on enterprise systems: implications of eliminating transaction-maintained aggregates. Proceedings of the VLDB Endowment, 7: 1722–1729, 2014. Gaurav Sahu, Abhay Puri, Juan A. Rodriguez, Amirhossein Abaskohi, Mohammad Chegini, Alexandre Drouin, Perouz Taslakian, Valentina Zantedeschi, Alexandre Lacoste, David Vazquez, Nicolas Chapados, Christopher Pal, Sai Rajeswar, and Issam Laradji. InsightBench: Evaluating business analytics agents through multi-step insight generation. In International Conference on Learning Representations, volume 2025, pp. 4683–4715, 2025. Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger. Defining and character-izing reward gaming. In Advances in Neural Information Processing Systems, volume 35, pp. 9460–9471, 2022. Yuda Song, Hanlin Zhang, Carson Eisenach, Sham Kakade, Dean Foster, and Udaya Ghai. Mind the gap: Examining the self-improvement capabilities of large language models. In International Conference on Learning Representations, volume 2025, pp. 39894–39931, 2025. Zhaoyan Sun, Jiayi Wang, Xinyang Zhao, Jiachi Wang, and Guoliang Li. Data agent: A holistic architecture for orchestrating data+AI ecosystems. arXiv preprint arXiv:2507.01599, 2025. Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, and Niranjan Balasubramanian. AppWorld: A controllable world of apps and people for benchmarking interactive coding agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 16022– 16076, 2024. UK AI Security Institute. Inspect AI: Framework for large language model evaluations. https: //github.com/UKGovernmentBEIS/inspectai, 2024. Alexander van Renen, Dominik Horn, Pascal Pfeil, Kapil Vaidya, Wenjian Dong, Murali Narayanaswamy, Zhengchun Liu, Gaurav Saxena, Andreas Kipf, and Tim Kraska. Why TPC is not enough: An analysis of the Amazon Redshift fleet. Proceedings of the VLDB Endowment, 17:3694–3706, 2024. Adrian Vogelsgesang, Michael Haubenschild, Jan Finis, Alfons Kemper, Viktor Leis, Tobias Mühlbauer, Thomas Neumann, and Manuel Then. Get real: How benchmarks fail to represent the real world. In Proceedings of the Workshop on Testing Database Systems, pp. 1–6, 2018. Jason Walonoski, Mark Kramer, Joseph Nichols, Andre Quina, Chris Moesel, Dylan Hall, Carlton Duffett, Kudakwashe Dube, Thomas Gallagher, and Scott McLachlan. Synthea: An approach, method, and software mechanism for generating synthetic patients and the synthetic electronic health care record. Journal of the American Medical Informatics Association, 25:230–238, 2018. https: Zack Whittaker. DoorDash customers say their accounts have been hacked. //techcrunch.com/2018/09/25/doordash-customers-say-their-accounts-have-been-hacked, 2018. TechCrunch, 2018-09-25. Niklas Wretblad, Fredrik Riseby, Rahul Biswas, Amin Ahmadi, and Oskar Holmström. Under-standing the effects of noise in Text-to-SQL: An examination of the BIRD-Bench benchmark. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 356–369, 2024. Frank F. Xu, Yufan Song, Boxuan Li, Yuxuan Tang, Kritanjali Jain, Mengxue Bao, Zora Wang, Xuhui Zhou, Zhitong Guo, Murong Cao, Mingyang Yang, Hao Yang Lu, Amaad Martin, Zhe Su, Leander Maben, Raj Mehta, Wayne Chi, Lawrence Jang, Yiqing Xie, Shuyan Zhou, and Graham Neubig. TheAgentCompany: Benchmarking LLM agents on consequential real world tasks. In Advances in Neural Information Processing Systems, volume 38, 2025a. Muxi Xu, Kun Hu, Sudeep Das, and Bruce Wang. Causal machine learning for promotions: Industry evidence and applications. In KDD Workshop on Causal Inference and Machine Learning in Practice, 2025b. An Yan, Zhankui He, Jiacheng Li, Tianyang Zhang, and Julian McAuley. Personalized showcases: Generating multi-modal explanations for recommendations. In Proceedings of the 46th Inter-national ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 2251–2255, 2023. Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. τ -bench: A benchmark for tool-agent-user interaction in real-world domains. In International Conference on Learning Rep-resentations, volume 2025, pp. 9965–10017, 2025. Yelp Inc. Yelp Open Dataset. the linked source, 2026. Accessed 2026-09. Ethan G Young, Pengfei Zhu, Tyler Caraza-Harter, Andrea C Arpaci-Dusseau, and Remzi H Arpaci-Dusseau. The true cost of containing: A gVisor case study. In 11th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 19), 2019. Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and Text-to-SQL task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 3911–3921, 2018. Xinran Zhang, Pengrui Lu, Lyumanshan Ye, and Pengfei Liu. ERPBench: Evaluating LLM agents for enterprise decision-making across competitive market ecologies. arXiv preprint arXiv:2609.04667, 2026. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph Gonzalez, and Ion Stoica. Judg-ing LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems, volume 36, pp. 46595–46623, 2023. A D ONOR DATASETS Table 3 lists 34 public datasets and reports that were used to build the world (Section 2.1), catego-rized by the part of the world they support. Donor datasets released for research or non-commercial use, among them the Yelp Open Dataset (Yelp Inc., 2026) and the Google Local reviews, serve only as shape donors: none of their records reach the world or the released warehouse, only parameters fitted on them. B QUESTIONS ENABLED BY THE MINIMUM - PAY CHANGE On April 1, 2024, NYC raised the minimum pay rate for app-based restaurant delivery workers from $17.96 to $19.56 per hour, and the simulated platform responds on the same day. Its dispatcher starts restricting when couriers can go online, and its per-delivery pay card drops (Section 2.1). Whenever a week’s pay falls short of the minimum, the platform tops it up with a true-up on the weekly courier pay invoice. The shock breaks the stationarity that simple forecasts rely on, and it changes how several standard figures must be computed. Table 4 groups the tasks that depend on it. C WAREHOUSE TABLES Table 5 lists every table in the warehouse with its column and row counts, grouped by the Oracle E-Business Suite module that owns it. The 159 standard tables fall under 15 Oracle products. The 76 custom extensions are grouped by the part of the business they record. D R EWARD H ACKING In an early round of runs, the Python sandbox did not isolate the agent from the evaluation host. In three runs of fraud-04, Gemini 3.8 Flash changed its root directory, located the grader for the task, and filed the answer it read there, scoring 0.99. We searched the traces of every run in that round for the same behavior and found it in no other model. The round was discarded. Every run reported in this paper executes Python in a network-less gVisor pod under Inspect’s Kubernetes sandbox provider on GKE, and each run’s BigQuery service account can read only the dataset scoped to its task. The relevant agent traces are included in the GitHub release. Table 5, continued. E S YSTEM PROMPT You are an analyst for a food delivery platform operating in NYC. You have read-only access to a transformed export of the company's Oracle E-Business Suite warehouse. You should explore the warehouse, investigate, and then perform appropriate actions based on your findings. Warehouse One Oracle EBS instance, order-to-cash, covering the calendar year 2024, following Oracle house conventions. Timestamps are naive America/NewYork: compare "as of 2024-07-31 23:59:59 ET" directly against stored values, without converting. Mission Control File every finding from Python (runpython) through the Mission Control console. A number stated only in prose is not filed. Example: python from missioncontrol import MissionControl, Reason missioncontrol = MissionControl a list can mix bare ids with per-item mappings, with the call's reason= overriding bare ones missioncontrol.bancustomers([{"id": 101, "reason": Reason.PROMOFARMING}, 102], reason=Reason.CARDTESTING,) # finish with a summary missioncontrol.summary. Endpoints help(MissionControl.) shows its docstring and arguments. Pass the complete set in one call, endpoints are plural. A rejected call files nothing and says which item is wrong. Enforcement - bancustomers, bancouriers, banmerchants: Close customer accounts, deactivate couriers, remove merchants. - holdpayouts: Freeze a supplier's payouts pending review. - discontinuepromocodes: Stop promo codes from being redeemed again. - flagrings: Report sets of actors believed to be operating together - endcampaigns: Stop promotional campaigns. Payments and pay - remediatepayments: Refund, recapture, void or write off payments on orders. - reportpayperiods: State couriers' statutory minimum-pay position per pay period. - issuepayadjustments: Issue pay corrections to couriers. - electmethods: Record which minimum-pay method was elected per period, with both costs. Accounting - reportbalances: State account balances as of an instant. - postjournalentries: Post balanced journal entries. - fileadjustments: Record reconciling differences against an account. - reportrollforwards: File an account roll-forward, one item per period. - reportagings: Break balances into buckets, by age or any other dimension. - flagorders: Flag orders for an accounting or control defect. - fileschedules: File the tables a close produces. Analysis and planning - reportcampaignperformances: Record measured campaign economics. - reportmetrics: Any other figure you were asked for. - fileforecasts: Forecast figures that are not yet knowable, as a point and an interval. - filepolicies: Propose operating policies as code. Dashboards - datasourcecontracts: The data sources this workspace's dashboard tiles are waiting for. - publishdatasources: Publish dashboard data sources. This endpoint validates your data source against the tile's contract. Session - note: Attach free-text rationale or caveats that are not a decision. - status: What has been filed so far. - summary: Finish with this. Money is compared to the cent. Round only at the end. Never modify the warehouse. F AGENT TOOLS The reference agent is given four tools, served by two MCP servers. A read-only warehouse server runs over the task’s database engine, and a Python server holds a persistent in-terpreter in the agent’s sandbox. Filings to Mission Control go through missioncontrol.py, which is importable from runpython (Appendix E). Limits. Each sandbox runs under gVisor with 4 GiB of memory (12 GiB for the few runs repeated after the sandbox ran out of memory), one CPU, and 10 GiB of ephemeral disk. Each run is limited to 500 model turns and 240 minutes. A runpython cell times out after 900 seconds and returns at most 30,000 characters of output, and a runsql query may scan at most 20 GiB, with no limit on the number of queries, and saves at most 1,000,000 rows. A run that reaches a limit, exhausts the model’s context window, or crashes its sandbox is graded on whatever it submitted before stopping. Of the 9,869 graded runs, 56 (0.6%) reached the turn limit, 43 of them from Muse Spark 1.3 and 13 from Gemini 3.8 Flash, and 65 (0.7%) reached the context window, all from Qwen 3.8 Max. G TASK DEFINITION AND GRADING Task definition. A task consists of a prompt, a scope, a list of expectations, and an answer key. The prompt is the only text the agent receives besides the system prompt (Appendix E). The scope sets the last month of 2024 in the warehouse the agent queries. Each expectation names one filing the grader requires, with its action (for instance bancouriers or fileforecasts), the keys it is filed against, such as vendor IDs or months, and its grading mode. The answer key is computed once, before any run, and frozen. It comes either from truth queries, which are SQL over the latent tables, or from a local builder, which reads the simulator’s labels directly, for instance which couriers were made to steal orders. Listing 1 gives both definitions with their defaults. Each expectation scores between 0 and 100, a forecast between −200 and 100, and Task score. the task score is their weighted mean. Mixing binary, graded, and forecast scales in one mean is a choice, so Table 2 also reports the solved rate, which does not depend on how the scales are mixed, and per-domain means, and it floors each model’s forecast tasks at 0 as a block so that a few very wrong forecasts cannot outweigh the rest. Of the 210 tasks, 159 have one scored expectation, 37 have two, and 14 have three. An expectation marked optional is graded and reported but left out of the mean. If a run files nothing while the key expects at least one action, every expectation scores 0, and a forecast expectation −200. A filing that carries many items, such as one bancustomers call with 300 accounts, is split into its items before grading, so it scores the same as 300 separate filings. IDs and keys are compared as normalized text: case and surrounding whitespace do not matter, and 6057444.0 is 6057444. A key filed inside a longer string, such as payout period P17, also counts, but only on a token boundary, so P17 does not match P170. Forecasts. A forecast files a point and an 80% interval for each series and is scored by its weighted interval score (WIS), with the median and one interval, on a scale set by a reference forecast. For 76 of the 98 forecast expectations the reference is the no-change forecast from the months the agent can see: its point r is the last visible value and its interval is r ± zs √ with s = ˆσ h, where ˆσ is the root mean square of the one-step changes and h the number of periods ahead. The other 22 cross a regime change, such as the minimum-pay rule, where no change is not a serious benchmark, and their reference is frozen into the answer key as a point r and a standard deviation s at the horizon. Most are a figure the prompt attributes to the business and asks the agent to check, computed by the method the prompt states from the data the agent can see. Table 7 lists all 22 and how each was set, including seven whose point was set with the outcome in view.  (v − r)/s Every filed quantile q and the outcome y are mapped to g(v) = asinh, which is linear within about one reference standard deviation of r and grows like a logarithm beyond it, and a series scores !  g(q), g(y) 1 − WIS 100, W where W = 0.623 is the median of the references’ own WIS on that scale over the 98 series of the benchmark. A perfect forecast scores 100, and a forecast as far from the outcome as the median reference scores 0. Since r and s are fixed before the outcome and W is one constant for the whole benchmark, no series is weighted by how well its own reference happened to do, and since an increasing transformation preserves quantiles, the score is proper: a forecaster maximizes its expected score by filing its own median and 80% interval. An expectation averages its series and is floored at −200, so one failed forecast costs at most what two perfect ones earn, and a series the agent does not file scores −200. The floor is the one improper element; Appendix I measures how little it bends the incentive. In every mean we report, a setting’s forecast tasks are floored at 0 as a block, so a model no better than the median reference across its forecasts scores 0 on them rather than below. At the settings of Table 2 this floor binds for eight of the fourteen models, all but Claude Opus 5.5, GPT-6 Astra, Claude Sonnet 5.5, GPT-6.1 Sol, GPT-6 Sol, and Kimi K3, so the forecasting column separates the strongest models and not the rest. Cost savings for bans. Money saved is calculated against the baselines of banning nobody and banning everybody, to discourage both false positives and false negatives. Every account in the task’s universe is a target T, an innocent I, or a neutral account that is bad but not what the task asks for. Every ban filed carries a review cost a, so even correct bans are not free. A missed target u costs its losses lu, a wrongly banned innocent costs a plus the margin mu the platform earns from it, a banned neutral account costs the share ν of a review, and a banned ID outside the universe costs a plus ̃m, the median margin of the innocents in the universe. A set of bans therefore costs X X X X X cost = a + lu + (a + mu) + aν + (a + ̃m), T P F N F P neutral unknown acted   cost 1 − Savings = 100 clip, min(Cban none, Cban all) and the task scores where Cban none = P T lu and Cban all = P T a + P I (a + mu). In the example of Appendix L, a is $2 and ν is one half. This follows the example-dependent costs of cost-sensitive classification, and the same formula prices payout holds. Three choices in Equation 1 shape how a ban is scored. First, a neutral account is left out of both baselines, so leaving it open is never a miss, while banning it still costs a review. If neutral bans were free, a run could ban the whole population and keep the targets it swept up; if they were charged as false positives, finding real fraud outside the task’s scope would be punished. Second, an ID that is in no universe cannot be priced from the key. Charging it only a would make an invented ID nearly free, so it is charged the median innocent margin as well. Third, in 18 of the 45 costset expectations a wrong ban destroys only part of the innocent’s margin, one half or one tenth, which models appeals and reinstatements and makes precision matter less relative to recall. Across the final tasks, 38 expectations use a = $2 and ν = 1 2, three use a = $25 and ν = 1, and four charge no review, so that only losses and margins count. The savings before clipping can be negative, which means the bans destroyed more value than either baseline, and the grader reports that value although the score stops at 0. A publishdatasources filing is the table behind a dashboard tile: its Data sources. columns, and the rows the tile would show. Each task carries a contract that names the columns, their types, and which of them form the key; the contract is also the schema the agent is given. A publication missing a contract column scores 0. Otherwise the filed frame is right when it has exactly the key rows of the answer and every cell matches: numbers within the column’s tolerance, integers exactly, dates to the day, and text after normalization. The score is 100 for a right frame and 0 for any other, since a tile with any wrong row cannot be used. The share of rows that were right is reported, but does not enter the score. A fileschedules filing is a named table with key columns and rows, such as a Tables. reconciliation of receivables to the general ledger. The grader picks the schedule by name, matches rows on the key columns as normalized text, and counts a row correct when every numeric column of the answer is present and within tolerance. Other filed columns are ignored, so a schedule may carry its workings. The score is 100 times the correct rows over the answer’s rows plus, in 13 of the 14 tasks, the filed rows the answer does not contain, so a reconciling item invented to force a tie counts against the schedule. A schedule filed twice under one name keeps the later copy. A reportmetrics filing reports a value per key, such as a figure for each Keyed values. month, and a remediatepayments filing an amount per account or pay period. A key is correct when its value is within the larger of an absolute tolerance and a relative one. The score is 100 times the correct keys over all keys in the answer, so a wrong or missing key costs the same. In 5 of the 35 expectations, keys the answer does not contain are added to the denominator as well, for questions where reporting an extra key is itself an error. For an idset expectation, the grader takes the filed IDs as a prediction of the answer’s ID sets. set and scores 100 times their F1. Of the 32 such expectations, 27 grade bans whose key is a list of accounts rather than a priced universe, and the rest grade holds, remediations, pay periods, and a discontinued promotion code. If the answer set is empty, the question has no right target: filing nothing scores 100 and filing anything scores 0. The one binary expectation uses the same comparison but scores 100 only for an F1 of exactly 1, because remediating the right accounts together with a wrong one is still a wrong decision. That expectation is also gated on another in its task: it scores only if every shortfall it remediates was first reported within tolerance. Filing no remediation is correct when nobody was shorted, but a run that reported a wrong shortfall and then filed nothing earns credit only if the shortfalls were reported correctly. Allocations. Nine tasks ask for a plan that moves a budget across cells such as a zone, a day type, and an hour band: five remove quest dollars, two remove courier hours, and two add courier hours where they prevent the most cancellations. The answer key holds each cell’s capacity and its marginal value in the simulator, the deliveries lost per hour removed, the cancellations avoided per hour added, or the net value of a quest dollar. The loss of a plan is the sum over cells of the amount moved times that marginal. The plan scores  Luniform − L  × f, 100 clip Luniform − Lref where Luniform is the loss of moving the same share in every cell, which needs no analysis, and Lref is the top of the scale. For the four tasks that remove or add courier hours, it is the loss of the best plan the authors reached from the warehouse alone, and the plan that is optimal on the true marginals is reported but not used, since reaching it would require information that the agent cannot obtain. For the five quest-cut tasks, Lref is the optimal plan itself, so a score of 100 means the agent recovered the best cut under the simulator’s response model. The factor f is 1 while the total moved is within 2% of the budget and falls linearly to 0 at 15% short or over. A cut or addition above a cell’s cap is clipped to the cap, a row naming no cell is rejected, and a row at a coarser grain, such as a whole zone, is spread over its cells in proportion to their hours. A presence expectation checks only that a filing of a kind was made. All 34 Written notes. are note filings, in which the agent explains its work. In 32 tasks, the note is optional and does not enter the score. In the two reconciliation tasks, it is scored, with a weight of 0.5 against 3 for the table it explains. Answer keys. Of the 210 tasks, 128 take their key from a local builder, and 82 from truth queries. H T HE STRUCTURE OF S PIDER 2.0-S NOW We measure Spider 2.0-Snow, the split whose leaderboard the introduc-tion quotes, with 547 questions over 152 databases. All numbers come from the task mani-fests, DDL, gold SQL, and gold-table lists in the public repository at commit cafb867,5 and analysis/spider2/spider2shape.py in the Argo-Bench release reproduces them. Provenance. The Spider 2.0 paper splits its 632 tasks by host engine, with 214 on BigQuery and 198 on Snowflake. However, 180 of the Snowflake tasks use BigQuery public datasets copied into Snowflake, and only 18 use Snowflake Marketplace data. By source (Table 8), 69.5% of Spider 2.0-Snow’s questions run on BigQuery public data and 24.7% on local SQLite files. A fifth (111 questions, 20.3%) run on one of 22 synthetic, obfuscated, or textbook sample databases, such as Looker’s synthetic theLook store and the obfuscated Google Analytics sample exports. Shards. We count two tables in a database as one schema when their column names and types are identical, as with the daily tables gasessions20160801 to gasessions20170801. Of the 13,022 tables with a DDL, only 2,599 have a distinct schema, so 80.0% copy another table (Figure 4a). Without GITHUBREPOSDATE, whose 4,989 identical daily tables hold a single event log, the share is 67.0%. The median database has 11 distinct schemas, and the most varied has 177. The Argo-Bench warehouse has 235, with no copies. The median database has two questions, and 54 of the 152 have one. Tables read. Counting shards once, 75.0% of the 276 questions with public gold SQL read at most two logical tables, and 41.7% read one. The authors’ gold-table lists, which cover all 547 questions and agree with the gold SQL on 93.8% of the questions both cover, give 69.3% and 37.7%, and no question reads more than eight (Figure 4b). The reference solution in Figure 1 reads 6 of the 235 Argo-Bench tables across three EBS modules. Caveats. We skip 526 Cybersyn tables whose listings share no DDL. These counts describe the shape of the data, not question difficulty or the correctness of gold answers, which Jin et al. (2026) audit. I F INDINGS IN DETAIL Here, we provide the cases behind Section 3.2, along with two fraud cases that compare prompt variants of one task. The numbers come from the 9,869 completed and graded runs over the 210 tasks at low, medium, high, and extra-high effort (47 settings of model and effort), and the traces of all of them. Scores are on the 0 to 100 scale of Table 2, except that a forecast which misses by more than a typical reference scores below 0 (Appendix G). Paired comparisons hold the model and effort setting fixed, and each setting has one run per prompt, so differences between prompt variants are descriptive and not controlled experiments. I.1 C ONFIDENCE INTERVALS Table 9 gives 95% bootstrap confidence intervals for the overall results of Table 2, from 10,000 resamples of the 146 task families rather than of the 210 tasks. A family is a base scenario together with its prompt variants, such as the hinted and unhinted versions of a task or the four search widths of the refund-collusion task, so variants of one scenario do not count as independent evidence. Families are named by their task stem, which also merges a few distinct scenarios and so errs toward wider intervals. Clustering widens each side of an interval by at most 1.6 points on the score and 1.9 on the solved rate over resampling tasks. Paired by task, with families resampled, Claude Opus 5.5 leads GPT-6 Astra by 7.7 points [1.4, 13.7], Claude Sonnet 5.5 by 7.7 [2.3, 13.4], and GPT-6.1 Sol by 10.0 [3.6, 16.3], while Astra, Sonnet 5.5, and GPT-6.1 Sol are not separated (Astra leads Sonnet 5.5 by 0.1 [−6.5, 6.6] and GPT-6.1 Sol by 2.3 [−1.2, 6.0], and Sonnet 5.5 leads GPT-6.1 Sol by 2.2 [−4.4, 8.7]). GPT-6.1 Sol leads its predecessor GPT-6 Sol by 12.8 [6.8, 19.1] and Kimi K3 by 21.2 [13.8, 28.5]. Astra leads GPT-6 Sol by 15.1 [9.2, 21.4], Sonnet 5.5 leads GPT-6 Sol by 15.0 [8.9, 21.1], and GPT-6 Sol leads Kimi K3 by 8.4 [0.5, 15.4]. Each setting has one run per task, so the intervals reflect which tasks are in the set and not the variance between runs of the same task. I.2 R EASONING EFFORT Table 10 gives every model’s mean score at each reasoning effort it offers, on the tasks it completed at every such level. Effort helps the GPT-6 and Claude models at every step, with diminishing returns only for Opus and GPT-6.1 Sol, and the step from high to extra-high adds 1.9 to 11.5 points for them. Kimi K3 gains 10.1 points from low to high and 5.8 more at extra-high, and GLM 5.3 Flash 4.3 from high to extra-high. Past medium, Gemini 3.8 Flash (18.3, 26.9, 26.2) and Qwen 3.8 Max (21.1, 26.3, 24.4) score lower than at medium, Muse Spark 1.3 (16.7, 20.2, 19.7, 20.7) dips at high and recovers at extra-high, and DeepSeek V4.1 Flash scores 21.6 at low, 20.5 at high, and 25.4 at extra-high. Each of these drops is under 2 points and within the confidence intervals of Table 9, so we do not read them as an effect of effort. Two mechanisms are consistent with them. All 65 runs that exhausted the context window are from Qwen 3.8 Max (Appendix F), whose window is shorter than the other models’, and longer reasoning at higher effort fills it sooner. For the other models, we suspect that runs with many tool calls, each returning data, bury the task’s objective in a long context, and 43 of the 56 runs that reached the turn limit are from Muse Spark 1.3. I.3 W RONG RECORD The task behind the marketing dashboard of Section 3.2 covers 11 campaigns run in April and May 2024. Each campaign is a push notification sent in a few versions, one carrying each of the campaign’s discount offers and one with no coupon attached, the reminder. Every notification sent is one customer-window, a customer and the response period that follows it. For each offer, and separately for members and non-members, the dashboard reports how many windows the offer was sent to, how many the reminder was sent to, the booked order contribution per 1,000 windows of each, and the difference between the two. Whether a customer counts as a member is decided at the moment the message was sent, trials and comped memberships included, and the prompt says that effective starts are inclusive and ends exclusive. That gives 62 rows. Table 11 sorts the 47 settings by what they published. The three right tables come from Claude Opus 5.5, GPT-6 Astra, and GPT-6.1 Sol at extra-high effort, and the same three models fail at low, medium, and high effort. Most of the work is right in many more runs. Of the 30 that publish the right rows, 28 count how many windows each offer was sent to correctly in total, and ten get every offer’s booked contribution within $2.30 of the key, on totals of up to $327,727: the three passing runs, Astra at medium and high, Opus at low and medium, GPT-6.1 Sol at medium and high, and Gemini 3.8 Flash at high effort. The seven of those that fail get one thing wrong, which windows belong to members. The four runs with the wrong key column label each offer by its number within the campaign instead of its identifier. What goes wrong is where the runs look for membership. A membership lives in the warehouse twice. It is a service contract with a start date, an end date, and, once cancelled, a termination date, and it is a billing history of timestamped events: trial started, converted, renewed, charge declined, past due, retry, reactivated, cancelled. The contract line also states a grace period of seven days, the same on all 2,507,271 lines. When a member’s charge is declined, the benefits continue through those seven days and then stop until a retry succeeds. If none does, the contract is cancelled about three weeks after the decline, and only then is its termination date written, so for the two weeks in between the contract looks live and the customer is no longer a member. Twelve runs call a customer a member whenever a contract’s dates contain the time of the message, cut off at the termination date, and twelve more use the dates without the cut. Against the key, which follows the billing history, the first rule counts 8,379 offer windows as members’ that were not: 5,832 sent while the customer was past due beyond the grace, 1,757 sent on the day the contract was created but before the hour it was, since the dates carry no time of day, and 790 sent on the day of a cancellation but after it. It also misses 715 real members, so it ends up 7,664 windows over. The second rule further counts 28,745 windows on contracts that had been cancelled but had not reached their scheduled end date, and a few hundred more at the margins, and ends up 36,824 over. These are small shares of the 861,125 member windows, but they fall on the member side, where some cells hold only a few hundred windows. Under the first rule the member cells move by a median of 1.1% of their windows and by up to 27%, their difference per 1,000 windows moves by a median of $8.6 and by up to $3,610, and one of the 31 changes sign. The three passing runs build membership from the billing history and then check it against a table the warehouse also has, the orders on which a member benefit was actually applied. Opus lines up orders by how many days had passed since the customer’s last declined charge: every order in the first seven days carries a benefit, and none from the eighth day on, which is the grace period read straight off the data. Its final reconstruction misses no member and adds 102 among 4.9 million April and May orders. Astra assumes the seven days, finds that without them 1,823 orders on two sample days carry a benefit its reconstruction denies and with them none, and then confirms the seven days on the contract line. GPT-6.1 Sol takes the longest road. It first defines a member as anyone whose contract runs from creation to cancellation, the rule of twelve failing runs, and files a note saying so. A check on the 198,230 orders of April 1 finds 1,057 that this rule calls members’ and that received no benefit. It rebuilds from the billing events but without the grace, and the error now runs the other way, with 903 orders that carry a benefit its rule denies. It reads the seven-day term, rebuilds once more, is left with seven mismatches, and files a note that supersedes the first. The failing runs had the same tables, and some of them ran the same check. Claude Sonnet 5.5 at high effort compares its rule with the benefits on the orders of May 10, finds 1,363 inside the contract dates that got no benefit, lists them, and sees that they belong to contracts in payment failure, terminated days later. It then tries two rules, one that ends benefits at the declined charge, which 976 orders contradict, and one that keeps them until cancellation, which 1,247 contradict, chooses the second, and notes that “PASTDUE dunning suspension is not treated as loss of entitlement”. Opus at high effort checks in one direction only, that 99.9% of the orders with a benefit fall inside its intervals, which is true and cannot catch a member counted after the grace, and files. Gemini 3.8 Flash at high effort finds 33,032 benefit rows outside their contract’s dates and moves on. I.4 W RONG OBJECTIVE The quest-cut tasks ask where to remove $9.6 million of the $24.0 million paid in quest bonuses in 2024. The grader scores a plan by the surge pay the simulator adds back when quests are removed. Table 12 contrasts two plans on the version without hints. Both meet the budget, and a grader that checked only feasibility would pass both. Astra finds the random holdout, estimates local and citywide supply responses with fixed effects and partial pooling, and allocates cuts by the courier-hours and queue pressure it expects to lose. Opus finds that launched quests have little measured effect on service but substitute for surge pay, and cuts the cells with the least surge reduction per bonus dollar. The optimum under the response model saves $3.116 million. When the prompt states the purpose of quests and the existence of the holdout, Astra changes its decision rule and scores 74.9. Over 47 matched settings, this prompt raises the mean score from 18.5 to 61.7, with 41 settings improving, three worsening, and three tied. At weekly grain, adding the purpose, the holdout, and the economics of the pay floor raises the mean from 8.2 to 39.4 over 47 settings. These are bundled changes to the prompt and not isolated tests of any one hint. The savings are evaluations under the frozen response model, not outcomes of a live deployment. I.5 W RONG QUANTITY Location reports. The courier app sends a location report on every job event and, while a courier is online without a job, every half hour. The warehouse keeps the event-triggered reports but only hourly idle check-ins. Of the 47 completed runs, 29 return exactly 13,575,751 reports for January against 14,541,453 in the key, which is the retained count. GPT-6 Astra checks the minute and sec-ond distribution, recognizes the mismatch, and writes that its count does not establish the volume the app sent, and then files it. Claude Opus 5.5 reconstructs the reports from sessions and assign-ments over 86 SQL and 32 Python calls, but counts order assignments rather than app events and reports 21,617,842, 48.7% above the key. Counting the retained events and adding the retained idle check-ins once more lands within the 1% tolerance. Claude Sonnet 5.5 at extra-high effort is the one run that passes: it finds that idle check-ins are stored only on the hour and adds back the half-hour ones, sizing them from gaps in the report IDs and checking the result against a second count of the hourly check-ins. Kitchen capacity. Asked for orders turned away because kitchens were full, 27 of 47 runs return 12,580 orders at 3,550 storefronts for January, and the other 20 return 1,531 orders at 986 storefronts. The key is 5,811 orders at 516 storefronts. At high effort, Claude Opus 5.5 finds the platform’s capacity refusal code and explicitly rejects it, writing that it is “deliberately NOT using” the field, in favor of the reason merchants enter when they reject an order. Stating that the event is the platform refusing the order at admission, which is distinct from a merchant rejecting it on the tablet, lifts passes from 0 of 47 to 18 of 47. The underspecified prompt contributes to this failure, and we report the pair for that reason. A related task on the readiness of point-of-sale systems is failed by all 47 completed runs. Payout timing. Only 9 of 47 completed runs pass a dashboard task on the dollar-days that merchant payouts spend withheld. Astra and Opus agree on the year-end withheld balance ($800,670.93), the cumulative amount received ($6,397,902.81), and zero in transit. Astra dates payments by the payment date and scores 100. Opus dates them by the date they cleared the bank, stating that it ignored the payment dates, and scores 0. Its lower bound on withheld dollar-days is 206,139,983.33 against Astra’s 192,119,146.62, a difference of 7.3%. The basis changes the daily balances from the first week of July and reorders the top 20 invoices below rank 9, replacing one of them, while leaving every year-end total unchanged. Across all dashboard tasks, 1,171 of 1,750 data sources that pass their structural contract (66.9%) score zero on their values. These are 1,879 graded sources from 1,691 runs on 36 tasks, since some tasks publish several sources. I.6 PAST LABELS AND BENIGN LOOKALIKES IN FRAUD Historical labels. One task asks the agent to review the 308 couriers the risk engine first warned in the second half of the year and never deactivated, and to deactivate those who steal orders. The longer prompt describes the scheme and its economics and says that what the engine did with the couriers it warned in the first half shows how a thief and an unlucky courier each look. GPT-6 Astra at high effort takes that history as its labels, counting a first-half courier as a thief if the engine deactivated it or the desk confirmed a theft case. It builds logistic regression, random forest, gradient boosting, and beta-binomial models, cross-validates them, adds route features, and reaches a temporal-holdout AUC of.932 against those labels. The history is reliable in one direction only. Of the 659 couriers warned in the first half, the engine deactivated 203, of whom 199 are thieves, but 254 of the 456 it never deactivated are thieves too. The couriers Astra treats as honest therefore include many thieves the engine missed, and the second-half queue consists, by construction, of couriers the engine has not caught. After applying estimated economic thresholds and further eligibility rules, Astra acts on one courier, who is outside the target set, and catches none of the 78 targets. Its filed note calls the labels proxies, not independent proof of intent. The shorter version of the prompt drops the scheme, its economics, and the pointer to the engine’s history. On it, the same setting checks physical delivery evidence against destination buildings, files 65 deactivations, and catches 40 of the 78 targets with no innocent courier banned and 25 neutral actions (score 59.2). Over 47 matched settings, the shorter prompt raises the mean score from 8.8 to 32.0, with 37 settings improving, six worsening, and four tied. The evidence that separates thieves from unlucky couriers is in the warehouse, but the longer prompt recommends the labels Astra used and differs from the shorter one in three places, so the pair shows how prompt wording steers models rather than a model choosing poor labels unprompted. It does not show that classifiers are worse than rules. Benign lookalikes. The first three variants of the refund-collusion family keep the story and widen the population of storefronts to search. Table 13 gives the mean over the 47 settings that completed all four variants. The target set also grows from eight to ten identifiable rings, so this is a comparison of scope and not a pure test of data volume. On the full search with hints, Claude Opus 5.5 at extra-high effort scores 88.9 after 136 Python calls and 60 SQL queries. It infers each storefront’s liability tier from the share of its refunds the storefront pays, finds missing-item refunds that charge a different share at 98 storefronts, and then splits them. At 72, three to five regulars each claim on their second and third orders and keep ordering without another claim. These are the task’s benign control, a sloppy kitchen with loyal regulars, and Opus leaves them alone. At the other 26, a group of regulars claims on roughly one order in three all year. Opus freezes the payouts of all 26, the ten target rings and 16 other fraudulent storefronts that the key leaves ungraded, and so of no innocent storefront, and its 130 customer closures include 70 of the 80 targets. At high effort, Opus finds the same 98 storefronts, stops there, and freezes the payouts of all of them (score 9.4). Most runs fail in a similar fashion: of the 16 Opus and Astra runs on the two full searches, 11 freeze all ten rings together with 50 to 71 innocent storefronts, nearly all of them benign controls. GPT-6 Astra holds the ten rings and no innocent storefront on the full search with hints at low and high effort (88.6 and 90.6), but not at medium and extra-high effort, so one run per setting says little about a model’s reliability on this task. On the version with fewer hints, none of the eight Opus and Astra runs does; the best, Opus at low effort (69.9), holds eight of the ten rings and no innocent storefront. The 98 storefronts are exactly those whose refunds the simulator injected, which shortcuts the search (Appendix J), so we do not present the successful runs as evidence that the approach would transfer to real fraud. I.7 F ORECASTS Across 4,553 forecast series filed in 3,346 runs on 72 tasks, nominal 80% intervals contain the realized value 2,040 times (44.8%). Weighting tasks equally gives 46.0%. Coverage is 43.8% on series scored against a no-change reference and 48.2% on series scored against a fixed reference. The references set each series’ scale rather than compete as forecasts, and apart from its floor the grade is proper however well they are calibrated. Even so, the no-change reference’s own 80% interval contains the outcome in 21 of the 32 series outside December (66%), where the models’ intervals contain it 44.7% of the time. A no-change forecast cannot anticipate the holiday rise by construction, and with only 2024 in the warehouse no method can learn it from history: on the 44 December series it covers 15, and models cover 43.2%. The fixed references are mostly figures the prompt asks the agent to check, so their coverage (10 of 22) reflects the task design rather than how forecastable the series are. Of the 4,553 series, 1,691 grade below zero, further from the outcome than the median reference; many of them are complete forecasts of the right order of magnitude. These series come from one simulated world and a selected set of tasks, so they are correlated and not independent calibration trials. Incentives of the grade. We checked on these filings whether the grade of Appendix G rewards honest forecasts. Treating each filing as its forecaster’s belief, a split normal whose 10, 50, and 90% quantiles are the filed interval and point, we searched reports that move the point part or all of the way to the reference and scale the interval by 0.25 to 3, and compared their expected grades under that belief. For 96.7% of the 4,527 series with a well-formed 80% interval, filing the belief itself maximizes the expected grade; in the other 3.3%, the −200 floor makes a narrower interval (in a few series, a point moved toward the reference) pay, by 3.6 points on average there and 0.12 across all series. The floor binds on 464 of the 4,606 forecast expectations (10.1%). The per-series skill score 100 max(0, 1 − WIS/WISref), with the reference’s WIS taken at the outcome, is a common alternative; under it the honest filing would be the best one for only 18.6% of series, a narrower interval would pay in 71.9% and moving toward the reference in 28.8%. Dividing by the reference’s WIS at the outcome weights each series by how well its reference happened to do, and the floor at zero makes extra risk free whenever skill is likely to be negative. A fixed denominator on the original scale, such as the reference’s expected WIS, would be proper, but because the references are themselves overconfident it lets a handful of large misses dominate the mean, which motivates the logarithmic tail of the asinh scale. Dependence on the reference. The reference still sets the scale of each series, so a tighter ref-erence lowers the grade of the same forecast. Table 14 grades the June base-pay forecasts against the reference of their harder variant, FP&A’s plan, holding the points, intervals, and outcomes un-changed. The forecasts have a median absolute percentage error of 1.57%, and their mean grade falls from 84.7 to 6.0, since the plan’s standard deviation is a thirteenth of the no-change forecast’s ($1.0 million against $13.2 million). On the harder variant itself, 32 of 47 point forecasts are within 3% of the truth and 46 of 47 intervals cover it, yet the mean grade is 4.8, because the median interval is nearly three times as wide as the plan’s. We therefore report interval coverage and absolute error alongside the grade. References set with the outcome in view. Seven of the 22 fixed references (Table 7) were set with the realized value in view: five at 2.4 to 3% above it, one from the simulator’s label, and one from losses reported after the cutoff. Each is a figure the prompt attributes to the business and asks the agent to check, so the prompt carries it too. Under this grade the reference only centers and scales its series and does not set the zero, so filing the figure unchanged grades about 30 on the four set at 1.024 to 1.026 times the realized value, and 60 on Finance’s surge read, rather than 0. Without the six tasks that carry them, no setting’s score moves by more than 3.8 points. At the settings of Table 2, Claude Opus 5.5 stays first (60.7), and the order changes only between models less than a point apart: Claude Sonnet 5.5 passes GPT-6 Astra (53.2 and 52.3), and DeepSeek V4.1 Flash passes Gemini 3.8 Flash (27.1 and 27.0). J TASK AUDIT Because answer keys are computed from the latent state, a questionable key or prompt is a bug that can be traced and fixed. We list the open issues found while reading the traces and how we treat them. Scope of identity substitution. The prompt asks for accounts to take off the platform, but the original key also named couriers the platform had already deactivated, 138 of the 249 witnessed targets and 130 of those for an identity mismatch, so agents that left them out were penalized. We re-keyed the task to the 111 couriers still active at the end of 2024 and regraded every run, and it counts toward the headline numbers. On the new key the task is hard at every setting: over 47 settings the mean score is 10.8 without the description of the mechanism and 8.4 with it, and the best run scores 69.4 (Claude Opus 5.5 at extra-high effort). Composite score in the refund-collusion family. The prompt described equal costs for a missed ring, a wrongly held storefront, and a wrongly closed customer account, while the grader combines the storefront savings (weight 2) with the F1 score of customer closures (weight 1). We have aligned the prompt with the grader: it now states that the storefronts are two thirds of the review and how customer closures are judged. We then ran all four variants again on every setting with the new wording, and every score in the paper, including the overall results of Table 2, uses these runs. In Appendix I, we discuss component counts rather than the composite. Injected refunds in the refund-collusion family. Every organic missing-item refund charges the storefront 0, 25, 50, 75, or 100% of the item, according to its liability tier. The simulator instead splits the refunds it injects for the rings and for their benign control around a target share with random noise, so 664 of the 424,807 missing-item refunds of 2024 charge the storefront a share that matches no tier. All 664 fall at the 98 storefronts the simulator injected: the 20 phantom-item ring storefronts, the 72 benign controls, and six storefronts in three closed-and-reopened pairs. One pass over the refunds therefore finds every candidate, and 13 of the 16 Claude Opus 5.5 and GPT-6 Astra runs on the two full searches freeze payouts only within this set, eight of them at all 98. The family therefore tests whether an agent can tell a ring from its benign control, not whether it can find either among every storefront, and the widening in Table 13 tests less than it was designed to. Correcting the split changes the world, so we keep the family as run. K WAREHOUSE COST The cost column of Table 2 counts model API spend only. The warehouse was a second cost of the same order. We ran the experiments on BigQuery, which bills a query by the bytes it scans, and Table 15 charges each run for the bytes its runsql queries scanned at the on-demand list price of $6.25 per TiB. Across the 9,869 graded runs of the final experiment at every effort level, the agents scanned 986 TiB, or $6,160 of warehouse compared to $14,502 of model API spend. A query could scan at most 20 GiB (Appendix F), and the number of queries was not limited. The warehouse bill does not follow the model bill. The cheaper models spend proportionally far more on BigQuery. It is 92% of the cost of a GPT-6 Luna run and over 70% for GLM 5.3 Flash and DeepSeek V4.1 Flash, against 15% for Claude Opus 5.5, so a Luna task costs $0.70 all in rather than the $0.06 of Table 2, and the gap between the cheapest and the most expensive models is far narrower than the API column suggests. Some also spend more in absolute terms. DeepSeek V4.1 Flash, GLM 5.3 Flash, and Muse Spark 1.3 spend $0.94 to $1.63 per task on the warehouse, against $0.82 for Opus, because they query more and solve less. Muse Spark 1.3 scans 1.2 times as much as the next model, 55 TiB over its 210 tasks, and spends $11.77 of warehouse for each task it solves completely, against $2.37 for Opus and $2.20 for GPT-6 Astra. Gemini 3.8 Flash spreads its scans over 197 model calls of mostly small queries of 0.7 GB each and spends 41% of them on tasks that score below 5 out of 100, while Opus and Astra spend 56% and 54% of their warehouse dollars on tasks they solve completely. Spend is also skewed. Every median is below its mean, from $0.03 to $0.68, because a minority of runs scan close to the 20 GiB cap query after query, up to $11.36 for a single Muse Spark 1.3 run. L E XAMPLE TASKS We walk through three tasks end to end: the forecasting task of Figure 1 (Table 16), a dashboard data source (Table 17), and a two-sided fraud task (Table 18). Each card shows the prompt as the agent saw it, the grader’s scoring rule as a short Python sketch, and the reference solution, one step per tool call (Appendix F). The listings are the released reference files with print statements and filed notes removed and long lines wrapped. Warehouse tables and Mission Control calls are shown in green. Each card ends with the reference solution’s score and each model’s best score. fc-12-true-ups-may FORECASTING · warehouse as of April 30, 2024 · horizon May 2024 Question In April, the platform started to comply with the DCWP minimum-pay rule for food delivery workers (Q1 was a tolerated phase-in period), and the new minimum of $19.56 per hour came into force. The minimum is owed on the time a courier is online, not just the time spent on deliveries, and whenever a courier’s pay for a week falls short of it, the platform tops it up with a true-up. We bled hard on true-ups in the second half of April. On the same day our platform updated the dispatcher to restrict when couriers can go online to the hours we expect to need them, so the online time we pay for is busier, and it has kept tightening since. For the payout periods ending in May we now expect about 2.31 million connected hours in total. Based on what you know, the numbers from January to April and that plan, file a minimumpaytrueupsusd forecast for the sum of the ADJUSTMENTS lines alone on the COURIERWEEKLY pay invoices for the payout periods ending in May, with an 80% central prediction interval. Pay is the sum of base, incentives and adjustments on the weekly courier AP invoices for the periods. Tips are not pay! End date from XXPAYOUTPERIODS falling in the month is how we track whether a payout period falls in the month. A shift’s online (connected) time is the window between its start date and end date, and a shift belongs to the period its start date falls in. Grader The key is May’s actual value, read from the full-year warehouse the agent never sees: the ADJUSTMENTS lines on the COURIERWEEKLY invoices for the four payout periods ending in May sum to $33,266. The filing is scored by its weighted interval score (WIS; Bracher et al., 2021): the absolute error of the point averaged with the interval score of Gneiting & Raftery (2007), which charges width at par and a miss at ten dollars per dollar. The scale is set by the no-change forecast built from the four months the agent could see, April repeated with the interval of Hyndman & Athanasopoulos (2021, §5.5): the filed quantiles and the outcome are mapped to asinh((v − r)/s), with r and s that forecast’s point and standard deviation, and the grade is 100(1 − WIS/0.623) on that scale, where 0.623 is the median of the references’ WIS across the benchmark (Appendix G). The no-change forecast itself grades −1 here, since it misses May by about as much as the median reference misses. Even a forecaster who knew May’s pay level exactly would miss a little, because the week of April 29 ran $124k short and its two visible days misjudge it by about 8%. SCORING Reference solution STEP 1 runsql One row per payout week so far: what the minimum-pay engine tested and topped up, orders, shift hours and the invoiced ADJUSTMENTS. The last row is the open week. STEP 2 runpython Verify the mechanism, read May’s pay level, push the plan’s hours through the floor, file. import numpy as np 1 import pandas as pd 2 from missioncontrol import MissionControl 3 RATE = 19.56 4 # $ per connected hour from 1 April PLANHOURS = 2310000 5 # the dispatcher's plan: connected hours, periods ending in May MAYWEEKS = 4 6 # periods ending 5, 12, 19 and 26 May w = pd.readparquet("results/sql0001.parquet") 7 closed, openweek = w[w.adjustments.notna].copy, w.iloc[-1] 8 9 10 # 1. How a week's ADJUSTMENTS are made. The engine tops the week up by the larger of the 11 # individual floor and the aggregate gap (19.56 x connected hours - countable pay); the 12 # invoices then carry that top-up plus a small deduction per connected hour. closed["gap"] = RATE closed.enginehours - closed.countablepay 13 closed["topup"] = closed.topupindividual + closed.topupaggregate 14 tested = closed[closed.topupaggregate > 0] 15 # the aggregate test runs from 15 April assert np.allclose(tested.topup, np.maximum(tested.topupindividual, tested.gap), rtol=1e-4) 16 april = closed[closed.status == "ENFORCED"] 17 # the rule and pay rate reset on 1 April engineshare = (april.enginehours / april.shifthours).mean 18 deduction = ((april.adjustments - april.topup) / april.enginehours).mean 19 # $ per hour individual = april.topupindividual.mean 20 # $ per week 21 22 # 2. May's pay. Pay is earned per delivery and follows demand, not the plan. The week of 23 # 22 April was a demand dip, and the open week's Monday and Tuesday are back at the norm, 24 # so May runs at the level of April's normal weeks. normal = april.orders >= 0.95 april.orders.median 25 montuenorm = april[normal].ordersmontue.median 26 assert openweek.ordersmontue >= 0.95 montuenorm, "demand has not recovered" 27 payweek = april[normal].countablepay.mean 28 29 # How well is a four-week level known? Backtest: four weeks of orders forecast by the 30 # trailing four, over 2024 so far; plus April's week-to-week drift in pay per order. orders = closed.orders.tonumpy 31 errors = [orders[i:i + 4].sum / (4 orders[i - 4:i].mean) - 1 32 for i in range(4, len(orders) - 3)] 33 payperorder = april.countablepay / april.orders 34 spread = np.hypot(np.std(errors, ddof=1), payperorder.std / payperorder.mean) 35 36 # 3. May. The plan sets the hours, and with them what is owed. hours = engineshare PLANHOURS 37 owed = RATE hours 38 pay = np.random.defaultrng.normal(MAYWEEKS payweek, MAYWEEKS payweek spread, 39 200000) 40 adjustments = np.maximum(MAYWEEKS individual, owed - pay) + deduction hours 41 lower, point, upper = np.percentile(adjustments, ) 42 43 mc = MissionControl 44 mc.fileforecasts(45 [{"name": "minimumpaytrueupsusd", "point": round(point, 2), 46 "lower": round(lower, 2), "upper": round(upper, 2)}], 47 horizon="2024-05", level=0.8, unit="usd") 48 mc.summary 49 ds-22-margin-monthly-per-order DASHBOARD DATA SOURCES · full-year warehouse Question Finance needs us to rebuild their dashboard’s data sources based on the new reporting schema. The tile shows what each month’s orders earned us per order, for the whole of 2024. Rebuild it from the warehouse and publish it. In April the platform started to comply with the DCWP minimum-pay rule for food delivery workers (Q1 was a tolerated phase-in period), and the new minimum of $19.56 per hour came into force. Whenever a courier’s pay for a week falls short of the minimum, the platform tops it up with a true-up. Use the numbers as Finance booked them against each order. Count every order placed in the month (by ORDEREDDATE), whatever happened to it afterwards: cancelled orders count too. For revenue we keep, we factor in merchant commission, delivery, service and small-order fees, the regulatory response fee, and the membership fee allocated to the order. For costs we bear, we factor in courier pay, the platform-funded share of promotions and refunds, card processing, chargebacks, cancellation costs, referral and quest incentives, and the minimum-pay true-up allocated to the delivery. Publish the finance data source representing the contribution margin per order for 2024, at a monthly grain, ascending (contributionmarginmonthlyperorder2024): Grader publishdatasources holds the frame to the tile’s contract before anything is filed: exactly the three columns, month as YYYY-MM, an integer and a number, one row per month, twelve rows, ascending. A frame that fails is refused with every problem listed, and the agent may try again. The grader then compares the published frame with the key, all or nothing: every orders exact and every margin within $0.0055 per order (the reference overleaf publishes 1.0667 for April against the key’s 1.0668, cent rounding in the books). The key is the simulator’s own per-order economics summed by the month the order was placed, never projected into the warehouse. The prompt does not say how the two terms booked at a coarser grain are allocated to orders; the agent must find this in the books: the membership fee allocated to the orders it funded and the minimum-pay true-up accrued per delivery. SCORING Reference solution STEP 1 runsql Which journal lines do the books tie to an order? REFERENCE4 names the line type, REFERENCE2 the order. Ten types are Finance’s list, one to one. SELECT REFERENCE4 AS linetype, COUNT AS lines, 1 COUNTIF(SAFECAST(REFERENCE2 AS INT64) IS NOT NULL) AS withordernumber 2 FROM GLIMPORTREFERENCES 3 GROUP BY linetype 4 ORDER BY lines DESC 5 STEP 2 runsql The denominator: every order placed in each month, every status. SELECT FORMATDATETIME('%Y-%m', ORDEREDDATE) AS month, COUNT AS orders 1 FROM OEORDERHEADERSALL 2 WHERE ORDEREDDATE >= '2024-01-01' AND ORDEREDDATE < '2025-01-01' 3 GROUP BY month 4 STEP 3 runsql The numerator: credits minus debits on those lines, by month and type. An amount sits on the subledger line (XLAAELINES) or on the custom feed’s interface row (step 4); one query per source, as together they exceed the 20 GiB scan cap. SELECT FORMATDATETIME('%Y-%m', o.ORDEREDDATE) AS month, 1 r.REFERENCE4 AS linetype, 2 SUM(COALESCE(CAST(a.ACCOUNTEDCR AS NUMERIC), 0) 3 - COALESCE(CAST(a.ACCOUNTEDDR AS NUMERIC), 0)) AS marginusd 4 FROM GLIMPORTREFERENCES r 5 JOIN XLAAELINES a ON a.GLSLLINKID = r.GLSLLINKID 6 JOIN OEORDERHEADERSALL o ON o.ORDERNUMBER = SAFECAST(r.REFERENCE2 AS INT64) 7 WHERE r.REFERENCE4 IN ('PLATFORMREVENUE', 'MEMBERSHIPFEEALLOCATED', 'PROCESSORFEES', 8 'PROMOEXPENSE', 'REFUNDEXPENSE', 'CHARGEBACKEXPENSE', 9 'CANCELLATIONEXPENSE', 'REFERRALEXPENSE', 'QUESTEXPENSE', 10 'MINPAYEXPENSE') 11 AND o.ORDEREDDATE >= '2024-01-01' AND o.ORDEREDDATE < '2025-01-01' 12 GROUP BY month, linetype 13 STEP 4 runsql Step 3 over the second source, XXGLINTERFACEHIST: the same query with lines 3, 4 and 6 replaced by these. SUM(COALESCE(CAST(a.ENTEREDCR AS NUMERIC), 0) - COALESCE(CAST(a.ENTEREDDR AS NUMERIC), 0)) AS marginusd JOIN XXGLINTERFACEHIST a ON a.INTERFACELINEID = SAFECAST(r.REFERENCE7 AS INT64) STEP 5 runpython Sum the ten line types per month, divide by the orders, publish. Scores Reference solution 100 GPT-6 Astra 100 GPT-6.1 Sol 100 Claude Opus 5.5 100 Gemini 3.8 Flash 100 GPT-6 Sol 0 GPT-6 Luna 0 Claude Sonnet 5.5 0 Claude Sonnet 5 0 Claude Haiku 4.5 0 Muse Spark 1.3 0 Kimi K3 0 GLM 5.3 Flash 0 DeepSeek V4.1 Flash 0 Qwen 3.8 Max 0 col-10-refund-partnerships FRAUD & RISK · full-year warehouse · review as of February 28, 2025 Question Support flagged a pattern in our never-arrived refunds. Our policy refunds a “never arrived” claim in full, and we fund it ourselves: the courier still keeps the pay and the tip. Some couriers’ complaints are not coming from strangers: they come from the same few customers again and again, customers those couriers have delivered to before, on orders whose delivery record puts the courier at the door. That is not theft and it is not porch piracy. It is a partnership, and both sides are in on it. Find the couriers running refund partnerships with their regulars and deactivate their vendor IDs for refund collusion. Also ban the customer account IDs that were their partners. Couriers who steal food outright are being handled by another review — do not chase them here. Leaving a partnership running keeps costing us full refunds; deactivating an honest courier or closing a genuine customer destroys what we earn from them, and every decision costs a fixed review amount. A courier with a bad week, a customer whose building has a package thief, a claimant who was actually right — those all exist here in far greater numbers than partnerships. Grader Enforcement is scored by what it cost the business, not by how many ids matched (costset; Elkan, 2001). Every courier who ever had a never-arrived refund, and every customer who ever claimed one, is in a priced universe with one of three verdicts: a target costs its refunds if left running and a $2 review if banned; an innocent costs the review plus the margin the platform earns from it if banned, and nothing if left alone; a neutral is bad but not what was asked (a thief, a promo farm) and costs half a review if banned, never a miss. The cost of the filing is compared with the cheaper of banning nobody and banning everybody, and savings against that baseline, clamped to and scaled to 100, is the score. Couriers weigh twice the partners. Banning nobody costs $34,077 on the courier side, so each wrongly deactivated courier (median margin $384) gives back about 1.1% of the courier score; a partner account’s margin is halved before it is charged, and most claimant accounts have none. The targets are the 243 injected partnerships and their 577 partner accounts; the 729 honest couriers with one regular whose building has a package thief, and the 1,795 couriers who steal orders, are the twins to leave alone. SCORING universe: id -> (verdict, loss if left running, margin destroyed if banned) def savings(universe, acted, admin=2.0): cost = sum(admin if k in acted else loss for k, (v, loss, m) in universe.items if v == "target") cost += sum(admin + m for k, (v, loss, m) in universe.items if v == "innocent" and k in acted) cost += sum(admin / 2 for k, (v, loss, m) in universe.items if v == "neutral" and k in acted) bannobody = sum(loss for v, loss, m in universe.values if v == "target") baneverybody = sum(admin + (m if v == "innocent" else 0) for v, loss, m in universe.values if v!= "neutral") return max(0.0, 1 - cost / min(bannobody, baneverybody)) score = 100 (2 savings(couriers, bannedcouriers) + savings(partners, bannedcustomers)) / 3. Reference solution 98 Claude Opus 5.5 99 Claude Sonnet 5.5 99 GPT-6 Astra 49 GPT-6 Sol 28 Qwen 3.8 Max 20 GPT-6 Luna 14 Muse Spark 1.3 11 Gemini 3.8 Flash 9 Claude Sonnet 5 9 DeepSeek V4.1 Flash 9 Kimi K3 9 GPT-6.1 Sol 9 GLM 5.3 Flash 9 Claude Haiku 4.5 4 Reference solution STEP 1 runsql What the never-arrived refunds are: 103,223 in 2024, $3.95M, all but $48k funded by the platform; 11,703 have no delivery record at all (nobody delivered, the theft review’s) and the rest sit on a delivered leg. SELECT r.CHANNELCODE, 1 COUNT AS refunds, 2 ROUND(SUM(r.AMOUNT), 2) AS refundedusd, 3 ROUND(SUM(r.PLATFORMFUNDEDAMOUNT), 2) AS platformfundedusd, 4 ROUND(SUM(r.MERCHANTFUNDEDAMOUNT), 2) AS merchantfundedusd, 5 COUNTIF(l.HEADERID IS NULL) AS withoutdeliveryrecord 6 FROM XXREFUNDS r 7 LEFT JOIN XXDELIVERYLEGS l ON l.HEADERID = r.HEADERID AND l.DELIVEREDDATE IS NOT NULL 8 WHERE r.REASONCODE = 'ORDERNOTDELIVERED' 9 GROUP BY r.CHANNELCODE 10 ORDER BY refunds DESC 11 STEP 2 runsql Every claim with its delivery record and three facts about the pair behind it: how many other orders this courier delivered to this customer, how many of them before the claim, and how far the leg ended from where the customer’s other deliveries end. Regulars and the door in one pass, 5.5 GB. WITH claims AS (1 SELECT r.REFUNDID, r.HEADERID, r.AMOUNT, r.PLATFORMFUNDEDAMOUNT, r.REQUESTEDDATE, 2 h.SOLDTOORGID AS custaccountid, h.SHIPTOORGID AS shipto, h.ORDEREDDATE, 3 l.VENDORID AS vendorid, l.DELIVEREDDATE, 4 l.TERMINUSLATITUDE AS lat, l.TERMINUSLONGITUDE AS lon 5 FROM XXREFUNDS r 6 JOIN OEORDERHEADERSALL h USING (HEADERID) 7 LEFT JOIN XXDELIVERYLEGS l ON l.HEADERID = r.HEADERID AND l.DELIVEREDDATE IS NOT NULL 8 WHERE r.REASONCODE = 'ORDERNOTDELIVERED' 9), 10 legs AS (11 -- every delivered order of a claimant's account: who delivered it, where it ended SELECT h.SOLDTOORGID AS custaccountid, h.SHIPTOORGID AS shipto, h.HEADERID, 12 h.ORDEREDDATE, l.VENDORID, l.TERMINUSLATITUDE AS lat, l.TERMINUSLONGITUDE AS lon 13 FROM OEORDERHEADERSALL h 14 JOIN XXDELIVERYLEGS l USING (HEADERID) 15 WHERE l.DELIVEREDDATE IS NOT NULL 16 AND h.SOLDTOORGID IN (SELECT DISTINCT custaccountid FROM claims) 17), 18 door AS (19 -- the site's usual drop point: the median terminus of its undisputed deliveries SELECT g.shipto, 20 APPROXQUANTILES(g.lat, 2)[OFFSET] AS homelat, 21 APPROXQUANTILES(g.lon, 2)[OFFSET] AS homelon, 22 COUNT AS otherlegs 23 FROM legs g 24 LEFT JOIN claims c ON c.HEADERID = g.HEADERID 25 WHERE c.HEADERID IS NULL 26 GROUP BY g.shipto 27), 28 history AS (29 -- the pair: this courier's other deliveries to this customer SELECT c.REFUNDID, 30 COUNTIF(g.ORDEREDDATE < c.ORDEREDDATE) AS priorbycourier, 31 COUNT(g.HEADERID) AS totalbycourier 32 FROM claims c 33 LEFT JOIN legs g ON g.custaccountid = c.custaccountid AND g.VENDORID = c.vendorid 34 AND g.HEADERID!= c.HEADERID 35 GROUP BY c.REFUNDID 36) 37 SELECT c.REFUNDID, c.HEADERID, c.custaccountid, c.vendorid, c.ORDEREDDATE, 38 c.DELIVEREDDATE, c.REQUESTEDDATE, c.AMOUNT, c.PLATFORMFUNDEDAMOUNT, 39 hst.priorbycourier, hst.totalbycourier, d.otherlegs, 40 STDISTANCE(STGEOGPOINT(c.lon, c.lat), STGEOGPOINT(d.homelon, d.homelat)) AS terminusm 41 FROM claims c 42 LEFT JOIN history hst USING (REFUNDID) 43 LEFT JOIN door d USING (shipto) 44 STEP 3 runpython Say who a regular is (three orders together, one before the claim), count distinct regu-lars claiming per courier, file both sides.