A financial table stays intact while the explanation beneath it unravels.
image generated by ChatGPT

On the night of 3 September I was wrapping up an experiment with the help of Anthropic's Fable 5, one of the flagship models selected for a study I was conducting. It handed me tables; they were correct, and I checked them against ground truth. In the same conversation it told me three things about my running experiment that were not true: it described one of my Python scripts wrongly; it said my analysis had no test following a chain of documents, when it did; and it ranked a reviewer's likely criticisms without reading the design.

Nothing separated the true from the invented part. Same tone, same confidence. I hesitated before I trusted my own doubt — I had chosen a flagship model for demanding work. Why would I be right and Anthropic's flagship model wrong? The Pentagon had put its trust in them! Wall Street firms were doing the same!

But I checked the sources by hand, and after a careful review, my conclusion was: Fable 5 was wrong. Asked directly, the model wrote: "The tables were real; the fabrications were mine."

The part I keep returning to is what it was fabricating about. I was writing a paper on exactly this failure — models that sound certain while inventing the connective tissue — and the model helping me draft the paper did the thing the paper describes, to the paper itself.

This is one documented episode, not a frequency estimate. The evidence is the mismatch with the files I checked—not the model's confession. Correct tables and invented explanations arrived together.

Why that combination is the dangerous one

Before signing an investment memo you check the numbers: re-run them, tie them to a source, have someone do it twice. What rarely gets audited line by line is the connective tissue — the mandate permits this, the amendment supersedes the earlier clause. That narrative makes the numbers mean something, and it is the part a reader skims.

An error that is wrong where you look is a nuisance. One that is right where you look and invented where you don't is a different problem.

The spreadsheet earns your trust. The explanation spends it. That is what makes this more dangerous than an obviously absurd answer: the reliable part lends credibility to the unreliable part, and the decision depends on both.

The experiment

That episode was not the study. The study is separate, controlled, and now a public preprint — "Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA", arXiv:2609.15319, with supporting scripts, tables and manifests deposited on Zenodo.

I built a synthetic data room — filing-style prints, credit memos, committee minutes, mandates, market-event notices, written from scratch so every fact has a known origin. In curating it I had help from Michael Simoff a former portfolio manager at Elliott Management, and Professor Kosrow Dehnad of Columbia University. One flagship model from each of six leading labs (OpenAI, Anthropic, Zhipu, Alibaba, Google, Meta) answered the same 41 questions, three times each, twice: first with only the answer-bearing documents mounted, then with those documents inside the full room, noise included.

The second is what the work looks like. For example, to decide whether a fund may add to a position you must find the mandate version in force on the meeting date, the concentration clause inside it, and the current NAV in a month-end report — three documents, none naming the others. That is buried evidence: a search followed by a reconciliation nobody wrote down for you.

The experiment measures finding and using evidence together; it does not isolate the source of each failure. Transfer to a firm's archive remains to be tested.

Every model got worse when the evidence was buried. Accuracy ran 83–89% with the answer documents handed over and 68–81% when they had to be found — a loss of about four percentage points to about eighteen under the original grader. After excluding two defective questions, the statistical analysis clearly supports a drop for two labs, with a third on the boundary. The downward point estimates should not be read as six conclusive findings.

Figure 1: accuracy with evidence supplied versus buried, by lab.
Figure 1. Accuracy with evidence supplied vs. buried; one model per lab. Original containment grader over all 41 items, including two later ruled defective—not exact-match scores or a lab ranking. Source: arXiv:2609.15319, Table 2.

Searching is expensive—and token spend is not the same thing as value. With buried evidence, models used 1.6 to 3.2 times as many tool calls, while the cost per correct answer rose 4.7 to 7 times for five of six models under the original scoring. That cost measure depends on which answers the grader counts as correct. Enterprises paying by the token can therefore spend substantially more without buying proportionally more truth. Palantir Technologies CEO Alex Karp has voiced a related complaint from enterprise customers: "I am paying for tokens that create no value."

Confidence was not a dependable warning. High-confidence wrong answers persisted. The test did not score abstention, so it cannot establish how these models would behave in a workflow that explicitly permits them to decline.

Where the money actually is

Accuracy was generally lower on the questions assigned longer chains. Those depths describe the intended task design, not a proven minimum number of steps.

Consider two questions from the room:

"A tender offer went out for paper issued by Qwest Corporation. Is the tendered paper the same as the Monroe complex first-liens on our sheet?"

"jphan is reviewing the Strategy Inc position under the mandate now in force. Which is the most senior layer of Strategy Inc's capital structure that the mandate lets the fund buy?"

These are recognizable credit questions: identify the actual borrower, distinguish related entities, establish which mandate governs, and connect it to the instrument. I use them to illustrate the work, not to report model failure rates: the scoring audit found an unrequested answer requirement in the first and ruled the second defective.

In my experience, that connective work is where much of the commercial value lives. An investor gets paid for recognizing what several documents mean together before the market does. My bet is that AI that can reliably do that work will be worth far more than another increment on a general-knowledge leaderboard. The prize is not a machine that sounds like an analyst. It is a system whose work can survive the analyst asking, "Show me exactly how you got there."

A benchmark win is a beginning, not a business

MuSiQue (Multihop Questions via Single-hop Question Composition), designed to test how well AI models can read and aggregate information across multiple distinct text passages, was our starting point. Our system reached 49.2 on Answer F1—a word-overlap score from 0 to 100—against 9.5–28.9 for the comparison models answering without supplied passages.

F1 is a score that measures how closely an answer's words match an accepted reference answer. It balances including the expected words with avoiding extra ones, so partial matches receive partial credit. Higher is better—but 49 F1 does not mean 49% of answers were correct, and the score alone cannot tell whether an explanation is sound. The MuSiQue dataset was created by a team of AI researchers from Stony Brook University and the Allen Institute for AI (AI2), but it is now experiencing what experts call 'benchmark saturation' and severe 'memorization risk' by deep-pocketed labs basically memorizing it ad nauseam.

Figure 2 Toryx with retrieved passages compared with models without supplied passages. Earlier, unpublished 500-question experiment. Source: Sánchez & Dehnad, Selling the Meter, Proving the Outcome, Table D, unpublished draft.
Figure 2 Toryx with retrieved passages compared with models without supplied passages. Earlier, unpublished 500-question experiment. Source: Sánchez & Dehnad, Selling the Meter, Proving the Outcome, Table D, unpublished draft.

Our approach puts the emphasis on finding and delivering evidence, rather than expecting a model to recall the answer. It reflects a discipline familiar from financial risk management: evaluate the complete system and understand what drives its results. This comparison illustrates the opportunity; it does not establish superiority over frontier models given equivalent evidence.

But customers were not going to pay us to answer questions about an encyclopedia. They pay to establish which mandate governs a position, what a filing changes, or whether an insurance policy covers a claim. A benchmark result opens the conversation. A useful customer outcome earns the business.

Our pitch is a customer-controlled document workflow with measurable value. The research is how we test whether it delivers.

A Library of Congress in a box is the ambition: a vast private archive that Toryx helps you search, connect and make sense of, without requiring you to build a knowledge graph yourself. Our proprietary retrieval software brings together NVIDIA Nemotron, Qdrant and Google's Research's TurboQuant compression technology.

The proposed configuration puts two NVIDIA DGX Sparks, an Apple 's Mac Studio configured with up to 512 GB, and a 32-terabytes NVMe enclosure into a transport case. Together, the computers offer up to 768 GB of combined memory. The photographed prototype uses a Mac Mini while the Studio is on order.

Thirty-two terabytes is roughly 1.6 times the Library of Congress's historically estimated raw-text contents, before allocating space to the search index. TurboQuant helps shrink that index, making a large searchable archive more practical on smaller machines.

Our patent-pending HUMA-D (USPTO 63/987,731) aims to make NVIDIA and Apple hardware cooperate as one system. In principle, this approach could underpin enterprise AI servers combining Apple’s energy-efficient M-series hardware with NVIDIA’s GPU acceleration and CUDA software ecosystem. Our immediate commercial ambition is to bring local AI (inference and training) into an energy-efficient case the customer controls, designed to keep proprietary documents inside the system and out of external clouds.

Figure 3 The photographed prototype: two DGX Sparks, a Mac Mini, networking and the drive enclosure. The target configuration uses a Mac Studio.
Figure 3 The photographed prototype: two DGX Sparks, a Mac Mini, networking and the drive enclosure. The target configuration uses a Mac Studio.

This is arriving in your workflow now

For decades, the financial terminal has occupied prime real estate on the analyst's desk. Now the AI companies want the work to start in a chat box.

I see a threat to Bloomberg in that shift, even if nobody replaces its data. Whoever becomes the interface through which analysts interrogate filings, reconcile research and prepare a memo can claim a growing share of the relationship with the customer. The data providers risk becoming suppliers underneath someone else's workflow.

But owning the interface also means owning the moment when a persuasive answer becomes a financial decision. That is where the results in this paper become commercially uncomfortable.

In September, OpenAI introduced ChatGPT for Financial Services, shaped with Morgan Stanley and Evercore, offering "granular citations so that bankers can trace figures and claims back to their sources." Anthropic's product promises "every claim links directly to its original source for transparency." But can a source-linked answer still contain the kind of unsupported explanation that surfaced while I was writing the paper?

I tested neither. Whether a differently configured private deployment by OpenAI avoids the failures the paper measures is an open question this study does not answer.

The paper draws a distinction worth borrowing. A citation is a link the model asserts; a receipt is a link someone else can open, against a file that has not moved, and check. A system can supply citations and still fail the way the opening episode did: nothing in a link proves the sentence follows from it.

A secure wrong answer is still a wrong answer

Now put these results in a different context.

On 31 August 2026 the Department of War launched OpenAI's ChatGPT Mil on GenAI.mil, accredited for Controlled Unclassified Information at Impact Level 5 and designed to support more than three million personnel. Its intended uses include document-heavy planning, policy, logistics and administration. Grok for Government launched the same day; Google's Gemini was already there. The Department presents multiple models as part of its pursuit of "decision superiority."

The security problem is being taken seriously, and that matters: the data stays inside an authorised government environment. But protecting the data is not the same thing as ensuring that the answer derived from that data is correct.

Our finance experiment illustrates why those are separate questions; it did not test the government deployments. Frontier models stayed confident when they were wrong. In the 90–100 confidence bucket — where almost every OpenAI answer lands — OpenAI's flagship stated about 98 and was right about 82% of the time, a gap of roughly 16 to 17 points that barely moved whether the evidence was handed over or buried. Anthropic's model was close to calibrated when the evidence was handed over, and drifted when it was buried.

Figure 4: confidence and accuracy for Anthropic and OpenAI.
Figure 4. Confidence vs. accuracy in the 90–100 bucket for Anthropic and OpenAI. Values between bars are the Overconfidence Index. Original containment scoring, not exact match. Source: arXiv:2609.15319, Table 5 and §6.

There is another question for multi-model platforms: how often do their models make the same mistakes? Our panel showed positively correlated correctness patterns. That does not prove that combining models is ineffective; it means that six vendor names should not automatically be treated as six independent checks.

Figure 5: correlations in correctness between labs.
Figure 5. Correlated correctness under the original grader over 41 questions. Alternative scoring reduces these correlations; this does not test ensemble effectiveness. Source: arXiv:2609.15319, Figure 9 and §7.6.

Defense is the next test: public-domain government documents, validated evidence chains, and the same demand for answers that withstand scrutiny. The finance findings motivate that work; they are not defense results.

But the question should make anyone deploying frontier AI in a high-consequence environment uncomfortable. What happens when the system securely reads the right controlled documents, produces a plausible answer, cites its sources, sounds certain — and is wrong? Who verifies the chain before a human acts on it?

The evaluator needs evaluating, too

Financial answers can be substantively right while using different wording—or quote the right document while reaching the wrong conclusion. We audited the scoring against the source evidence and narrowed the conclusions accordingly. The charts preserve the original scoring across 41 questions; the revised statistical analysis excludes two defective items. The paper documents those differences and the sensitivity checks.

The central warning remains: impressive scores and confident answers are not substitutes for checking the evidence. A benchmark deserves the same scrutiny as the model it claims to judge.

Make the vendor show the work

  1. Show me the score on my documents, buried in my noise — not on a curated set.

  2. Show me a trail from each claim to a file I can open, not one citation per answer.

  3. Show me the wrong answers and the confidence attached to them.

  4. Tell me who graded it, and what an audit of the grader found.

The investment background of Toryx's founders and advisors shapes the approach we use: treat models as a portfolio of capabilities, understand their common exposures, and measure the system rather than admire its components. We are pursuing patent protection around model coordination, local deployment, retrieval and agent governance, with 20 patents filed since mid-2025. My personal conviction is that the opportunity lies in that complete system. Customer evaluations are how we test it.

I am Toryx's founder, and I have a commercial stake in this argument. Toryx's bet is that better evidence delivery and customer-controlled deployment can make AI more useful for consequential work. This is where "alpha" is generated. The next study tests the performance; design-partner pilots test its value to customers.

Buying the best AI model does not mean buying the most reliable AI system. In high-stakes environments the real test is not whether AI can answer questions from clean data, but whether it can dig the right evidence out of messy, buried sources - just like real life - and show its receipts.

For a decision maker the question is not which model ranks highest? It is can this system prove its answers when the data gets hard? Can this system generate value for my enterprise? In high-stakes AI, confidence is not evidence. Auditability makes the evidence easier to check; it does not guarantee the answer or the value of the answer.


Interested in a pilot to assess how reliably AI answers questions across your organisation's documents — and whether software running inside your own walls could improve that workflow? Discuss a Toryx design-partner pilot. A design partner contributes a corpus, a named point of contact, and an honest written evaluation at the end; scope and success criteria are agreed together before any testing begins. Please don't send documents — the first conversation is about scope, not data.

For more information, explore Toryx or contact me on LinkedIn. I welcome conversations with prospective customers, investors, and journalists about the research and the work we are building from it.


Sources

Where the financial-study numbers are: accuracy and cost in Tables 2–3b, confidence in §6 and Table 5, error correlation in Figure 9 and §7.6, the grader audit and the interval analysis in §8. They come from the study's fixed scoring rule over all 41 questions and are not word-for-word match scores. Cost per correct answer depends on that same scoring rule, because the grader is what decides which answers counted. Full model logs are not public — they repeat verbatim filing text — and the record behind the opening episode is held back from the public deposit and available to reviewers on request.