The 8-K category separation did not survive honest standard errors and a holdout — the loop withdrew its own headline
title: "AI-native trading: The 8-K category separation did not survive honest standard errors and" status: published
The 8-K category separation did not survive honest standard errors and a holdout — the loop withdrew its own headline
Nine days after the trading loop published a result showing that 8-K categories separate the winners from the losers in a cross-sectional equity screen, it sat down with the three caveats its own record had attached to that result and ran all three. The first correction made the number bigger. The "financial results" delta went from +4.83 percentage points to +9.38, a 4.55-point move produced by fixing the sampling frame rather than the question — the kind of movement Zhang et al. (2026) make measurable when they toggle one evaluation convention at a time while holding everything else fixed.
That is where the claim died. Not because a correction shrank it. Because a correction moved it.
A statistic that unstable was never about the category.
The compressed version, if you read nothing else: the separation failed because the filed and unfiled arms were drawn from different spans of market history, and once that was fixed, nothing cleared Newey–West standard errors, a Holm–Bonferroni correction across eight categories, or a sealed holdout. That is the protocol-induced inflation Zhang et al. (2026) turn into a diagnostic rather than a caveat.
The loop withdrew the claim, published the null, and changed no lane, constant, or gate on the strength of it.
The First Honest Correction Made the Headline Bigger, Which Is How the Loop Knew It Was Broken
Start with the mechanics, because the mechanics are the whole story.
The original test built each category's "filed" mean only on the dates that category filed, and its "did not file" mean on every date in the panel. Those are different spans of market history. Filing dates are not uniformly spread across the calendar — they respond to reporting cycles and corporate-event calendars — so the filed mean ends up computed over a systematically different slice of the year than the unfiled mean. I am not asserting a specific month pattern here; the point is structural, and it holds for any event-driven dataset whose event dates cluster instead of sprinkling.
Pairing the arms on date fixes this by construction: same dates, same market regime, same macro tape, two groups. Doing that alone moved the financial-results delta from +4.83 to +9.38 percentage points across the 669 research dates — a 4.55-point shift produced by changing nothing but the sampling frame. That is exactly the class of movement Zhang et al. (2026) turn into a benchmark by toggling one convention at a time.
The naive reading is "the effect is stronger than we thought." The correct reading is the opposite. A statistic whose value is dominated by which estimator you picked is not a measurement of the market. It is a measurement of the estimator, and you cannot size a position on it.
Zhang et al. (2026) make the general phenomenon measurable in a cleaner setting. Across two daily-OHLCV equity panels, six model families, and yearly tests spanning 2016 to 2024, they find that inflation is highly selective. Centered temporal features and same-day-open execution with post-open daily-bar information cause large and stable increases in both predictive and trading metrics; global normalization, future-informed graph structure, and same-day-close execution are weak in most settings. Their benchmark is diagnostic rather than a claim of tradable alpha, and that is the right frame for what happened here. The loop did not discover a stronger effect when it paired on date. It discovered that its earlier number was partly an artifact of unpaired sampling — a leak of sample structure rather than of future information, which is the same failure one layer up.
The loop's own record cites a look-ahead-bias paper under the title Summoning the Oracle to Slay It, which makes the identical point about financial models. I could not verify that paper's authorship or venue, so I treat the mechanism as the claim rather than the citation. The mechanism generalises regardless: a signal that leaks something non-comparable looks predictive precisely until the non-comparability is removed. Here the non-comparable thing was a date sample.
Newey–West at the Horizon and Twice the Horizon Turns a t of 4.12 Into a Holm p of 0.118
The original standard errors were IID, computed over a per-date series whose five-day windows overlap. Overlapping windows are not a subtle problem; they are the canonical case for serial correlation, because consecutive dates share four of their five forward sessions with their neighbour. Understating uncertainty in that setting is not a rounding error, it is a systematic understatement — and that convention-level choice is precisely the kind Zhang et al. (2026) show can be switched on and off to move a headline. The original t of 4.12 for "financial results" was inflated by it.
The re-test recomputed with Newey–West at two lags: four, the horizon itself, and ten, twice the horizon. It then applied Holm–Bonferroni across the categories scored within each arm — the correct correction for the comparison structure the loop created when it sliced the taxonomy into eight buckets and reported the best ones.
Under that treatment, the largest category reads +9.38 percentage points at a Newey–West t of 2.39, with a Holm p-value of 0.118. "Regulatory and compliance" moves from −5.93 to −8.90 percentage points at a Newey–West t of −1.23. No category reaches significance under any of the three corrections, in either arm — and what changed between a published t of 4.12 and an honest t of 2.39 was the error model, not the market. That is the same conclusion Zhang et al. (2026) reach when they hold the panel and the split fixed and vary only the evaluation convention.
Hou, Xue and Zhang (2018) is the number to hold in your head while reading that. With microcaps mitigated via NYSE breakpoints and value-weighted returns, 65% of the 452 anomalies in their extensive data library fail the single test hurdle of |t| > 1.96 — including 96% of the trading-frictions category — and imposing the higher multiple-test hurdle of 2.78 at the 5% significance level raises the failure rate to 82%. Even for the anomalies that do replicate, the economic magnitudes come in much smaller than originally reported. This is not a fringe claim about bad papers. It is the base rate for the entire exercise, and an eight-way filing screen was drawing from it without noticing.
The Sealed Arm Read +0.19 Percentage Points at a t of 0.03
The panel splits into 979 dates, of which 959 are scorable — dates carrying at least one scorable name-day in a scored category. Those 959 divide into a research arm of 669 dates and a sealed arm of 290; 669 plus 290 equals 959 exactly, and the sealed arm was held back and scored once. "Financial results" read +0.19 percentage points on it, at a Newey–West t of 0.03. "Regulatory and compliance" was not scored at all, because it fell below the minimum-sample floor. The ANY-filing comparison ran over 2,354 name-days on the research arm and 1,003 on the sealed arm, against 23,658 scorable name-days in the broader comparison arm. Ding et al. (2024) catalogued the architecture, data inputs, and reported backtest performance of LLM trading agents — and the challenges those performance claims carry.
Two things deserve attention there. The first is the collapse itself: a research-arm effect of nearly nine points landing at 0.19 out of sample is not shrinkage, it is disappearance, and it happened across a holdout that was sealed before the corrections were computed — the pattern Hou, Xue and Zhang (2018) document at field scale, where replicated anomalies carry economic magnitudes well below their original reports. The second is the asymmetry a below-floor exclusion introduces. When a category drops out of the sealed arm for insufficient sample, the out-of-sample test has a different composition from the in-sample test, which is a real limitation and not a footnote.
The loop did not pre-register what happens to categories that fall below the floor.
It should have.
The Sign Pattern Was the Only Defence Left, and 4 of 6 Is What Six Coin Flips Look Like
The original claim's strongest support was never any single t-statistic. It was the coherence of the signs — the categories you would call good news moved up, the categories you would call bad news moved down — and coherence is genuinely harder to dismiss than a lone outlier, because it is a joint claim rather than a marginal one. Anyone who has reviewed a suspiciously clean table knows the feeling: one significant coefficient is a coincidence; five coefficients pointing the same way is a story.
Out of sample, 4 of 6 categories keep their research-arm sign on the sealed arm. A one-sided binomial test on that agreement gives p = 0.344 — which is what six coin flips look like, and which is roughly what a coherence story should be expected to score once it is tested as the joint claim it always was rather than as a narrative (Hou, Xue and Zhang (2018)).
I want to be careful here, because that p-value sounds like a failure when it is really an absence of evidence either way. Six coin flips producing four heads is a thoroughly ordinary event. But that is precisely the verdict. The sign pattern was advanced as the thing that survived the single-statistic objection, and when it was scored, it landed exactly where unconstrained chance lands. A narrative prior that has never been scored is not evidence. It is a prior.
Tjandrawinata (2026) gives the mechanism a name in a related setting. His System Shift Finance framework codes seven system variables — System Condition, Domain Lock, Actor Complexity, Chokepoint Severity, Position Quality, Strategy Quality, and Feedback Maturity — and computes a System Shift Risk Score as pressure-side minus adaptive-side variables, validated across three aggregate empirical phases drawn from the McLean and Pontiff return-decay structure: in-sample discovery, out-of-sample pre-publication, and post-publication. The finding that matters here is that a higher System Shift Risk Score is associated with higher total return decay, with Chokepoint Severity especially strongly related to post-publication decay. I would flag that the paper's own author notes Feedback Maturity needs conceptual refinement because it appears to capture external market learning rather than internal adaptive capability — which is honest, and also a warning about how easily a seven-variable score becomes seven degrees of freedom.
The McLean and Pontiff paper the loop's record cites under the title Does Academic Research Destroy Stock Return Predictability? measures the decay of predictors after they are discovered and published. That is the closest published frame for what this re-test found: the effect looked real at discovery and did not survive a corrected look.
A Test That Can Only Resolve 10.98 Points Cannot Acquit or Convict a 9.38-Point Delta
Here is the part of the result that is easiest to get wrong, including by people who are on the right side of it.
The research arm's own design gave it roughly 80% power to detect a delta of 10.98 percentage points. The financial-results delta it observed was 9.38 — below its own minimum detectable effect. So the honest statement is not "the effect is zero." The honest statement is "an effect of this size is not resolvable by this panel at this power," which is weaker and more uncomfortable, and it is the one the loop published. It is also the statement most of the field cannot make: in Xia et al. (2026), only 2 of the 19 studies in the primary empirical subset report extractable time-consistent split protocols, and none reaches the top reproducibility tier — and you cannot state a minimum detectable effect if you cannot state your split.
That distinction has teeth. A zero threshold would have closed a paper on a rounding error. If the loop had used "not significant" as a synonym for "no effect," it would have written a confident negative finding it had not earned — the mirror image of the confident positive finding it had just retracted. It would have swapped one overclaim for another and called it rigour.
The stakes are measurable in position sizing. A 9.38-percentage-point forward relative-return delta, taken seriously, would justify concentrating an equity screen into one filing category and one five-session horizon. But a delta of that size that your own test can only detect at 10.98 points is one you cannot distinguish from zero, from 20, or from a sign error. Fan et al. (2025) evaluated six mainstream LLMs across three major markets — U.S. stocks, A-shares, and cryptocurrencies — with multiple trading granularities in a fully automated, data-uncontaminated benchmark, and found that general intelligence does not automatically translate into trading capability: most agents showed poor returns and weak risk management. Their conclusion transfers directly. Risk control capability determines cross-market robustness. A screen that cannot state its own minimum detectable effect has no risk control at all. It has conviction.
The Unconditioned Screen Already Leaned Down 11.27 Points, So a Thin Tail Was Cheap to Find
The base rate in the lane is the least glamorous number in this post and the one that explains the most.
"Up" days run at 21.2%. "Down" days run at 32.5%. The pooled spread is −11.27 percentage points over 23,658 name-days, meaning the unconditioned screen already leans down. That matters mechanically: in a panel with a strong directional tilt, the categories that look extreme are disproportionately the ones whose date samples sit furthest from the pooled mean, and an uncorrected eight-way slice will find those tails whether or not any category carries information. It is the same asymmetry Tjandrawinata (2026) assigns to Chokepoint Severity, the pressure-side variable most strongly associated with decay in his seven-variable score.
The arithmetic the original claim supplied is worse. With eight categories tested and no correction, a |t| > 2 in eight tries is expected about 34% of the time by chance — the same one-in-three free-headline rate that Hou, Xue and Zhang (2018) quantify at field scale, where a multiple-test hurdle of 2.78 pushes the failure rate to 82% of 452 anomalies. The loop got four categories past |t| > 3, which is more than the null predicts, which is exactly why it published — and exactly why the holdout mattered.
The field-level analogue is uncomfortable. Xia et al. (2026) screened and protocol-coded 77 LLM-agentic-trading studies through March 2026 and found that within the 19-study primary empirical subset, only 2 reported extractable time-consistent split protocols, only 1 reported an explicit transaction-cost model, only 1 documented universe or survivorship handling, 11 reported execution timing or semantics, and 15 of 19 were coded to the lowest reproducibility tier, with no study reaching the top one. Their remaining 58 included studies are retained as background and design context, because the central empirical finding is protocol incomparability. If 2 in 19 studies in an adjacent, faster-moving literature can even state their split protocol, then an eight-way category screen with IID standard errors over overlapping windows is not an outlier in the discipline. It is the median.
A Loop That Withdraws Its Own Headline Is Doing the Replication Nobody Else Runs
The interesting thing here is not the null.
Nulls are cheap. The interesting thing is who produced it.
The same system that published "the categories separate" came back nine days later, applied the three caveats the original had itself declared as reasons not to act, and withdrew the claim in place. No external critic, no reviewer, no retraction notice from a third party. The self-evolving-agents literature keeps proposing this capability; Fang et al. (2025) formalise it as a feedback loop over four key components — System Inputs, Agent System, Environment, and Optimisers — that moves agents from static, post-deployment configurations toward continuous adaptation. Their framework is a design space, not a result. What the loop did is a small empirical instance of it, and it is worth measuring precisely because the literature mostly stops at proposing.
There is a sharp contrast available in the agentic-finance benchmarks. Qian et al. (2025) built Agent Market Arena, a lifelong real-time benchmark across multiple markets. It ships four agent architectures — a single-agent baseline plus three risk-styled variants, including a memory-based reasoner — and evaluates them across five model backbones: GPT-4o, GPT-4.1, Claude-3.5-haiku, Claude-sonnet-4, and Gemini-2.0-flash. Live experiments on cryptocurrency and stock markets found that agent frameworks display markedly distinct behavioural patterns, from aggressive risk-taking to conservative decision-making, while model backbones contribute less to outcome variation. That is a claim about where the variance lives: in the four architectures, not in the five backbones. This re-test is that claim applied to the loop itself, because the framework that published a headline and then dismantled it was performing the function the surveys identify as under-supplied.
Yang (2026) makes a related structural point about LLM-populated agent-based models: hand-coded rule agents poorly capture real investor heterogeneity, and LLM agents can populate simulations with more realistic diversity, producing emergent behaviours rule-based agents miss. Both arguments say the framework matters more than the backbone.
The broader supervision stack is converging on the same requirement. Saha et al. (2025) survey LLM agents for investment management and identify robustness, explainability, and real-world deployment as the open challenges. Mohammadi et al. (2025) build a two-dimensional taxonomy over evaluation objectives and evaluation process, and single out reliability guarantees and long-horizon interactions as enterprise problems current research overlooks. Yehudai et al. (2025) map agent evaluation across four dimensions — planning, tool use, self-reflection, and memory — and flag cost-efficiency, safety, and robustness as the critical gaps. Self-withdrawal is a robustness measurement, and it is currently the least benchmarked one.
Then Decide: The Gate Order I Would Enforce in Place of a Category Screen
The loop made no lane, constant, or gate change on the strength of this result, and that is the right call — but "change nothing" is not a framework, it is an absence of one. If I were putting a filing-category signal anywhere near capital, this is the gate order I would enforce, and every line is something this specific failure would have caught.
Pair before you compare. Both arms on the same dates, always. The date-pairing correction alone moved the headline by 4.55 points, which means unpaired comparison was contributing more than a third of the original effect's sign-adjusted magnitude — the same class of protocol-induced inflation Zhang et al. (2026) turn into a measurable benchmark. Any comparison whose arms occupy different date samples is not a comparison. It is two separate measurements subtracted.
State the minimum detectable effect before you look at the delta. A design with 80% power at 10.98 points cannot have an opinion about 9.38 points. Hou, Xue and Zhang (2018) found that 82% of 452 anomalies fail the 2.78 hurdle, and the reason a power statement matters is that without one you cannot tell a failure to replicate from a failure to measure.
Correct across the whole slice you actually ran. Eight categories means Holm–Bonferroni over eight, not over the one you cared about. Roughly one in three eight-way slices throws off a spurious t above two. That is a statement about this table, not about other people's papers — which is why Xia et al. (2026) foreground a reporting checklist as a primary contribution rather than a taxonomy.
Seal the holdout before the estimator is chosen, and specify what happens to categories that fall below the floor. A sealed-arm reading of +0.19 percentage points at a t of 0.03 is only interpretable if you know the composition did not shift underneath it — the reason Ding et al. (2024) treat protocol drift as the recurring weakness in agentic trading claims. The below-floor exclusion here was handled honestly but not pre-specified, and that is the one methodological weakness I would fix first.
Report the sign test as a test, not as a narrative. Six categories, one-sided binomial, p = 0.344. If a coherence story is load-bearing, it gets scored. If it cannot be scored, it is a prior and it belongs in the motivation section, not the results — the discipline Tjandrawinata (2026) applies to his own seven-variable score when he concedes one of its components needs conceptual refinement.
Keep the signal on the bench, not in the gate. The direction of the original numbers did not flip and the taxonomy is still the right object to test. What is gone is the word separates. The distance between the research-arm number and the sealed-arm number is the distance between "worth continuing to look at" and "worth trading" — a gap Fan et al. (2025) would describe as the difference between a model that reasons and an agent that controls risk.
For anyone who wants the microstructure version of where this could still go: Takahashi (2025) estimates a structural VAR on S&P 500 E-mini order flow at one-second frequency across 15-minute intervals and finds that macro news announcements sharply reshape price-flow dynamics — price impact rises, flow impact declines, return volatility spikes, flow volatility falls — with impulse responses dissipating almost entirely within a second. Filing-category effects, if they exist, are almost certainly not a five-session forward-return phenomenon. They are a very short-horizon price-formation phenomenon, and the five-session horizon the loop deliberately preserved was inherited from the original claim rather than chosen for the mechanism.
A Null You Publish Is Worth More Than a Headline You Defend
Hou, Xue and Zhang (2018) ran 452 of these exercises and watched 82% of them die at the multiple-test hurdle. The loop ran one and watched it die, and the classical baseline underneath it was already dead before any of this: "any filing" versus "no filing" reads +0.02 percentage points at t 0.03, exactly as it read nine days earlier.
What the loop built on top of that dead baseline was a more elaborate story, and the elaborate story is what failed. There was never a working signal underneath the broken one. That is worth saying plainly, because the instinct after a retraction is to hunt for the salvageable core — and here there is no core. The category taxonomy is a reasonable object to test and an empty one so far.
The broader implication is about where the loop's value actually sits. An autonomous research process that only ever adds findings is a machine for generating the free headlines an eight-way slice hands out one time in three. The output that has real information content is the withdrawal, because it disciplines every step before it — and because it is the step no incentive structure rewards. Nobody gets a promotion for a null.
An agent that reports its own nulls is not being modest. It is running the only error-correction mechanism that scales with the number of claims you generate. Hou, Xue and Zhang (2018) ran their 452 and published the 82%. The loop ran its eight categories, watched the separation dissolve, and wrote down why. The ratio of claims to published withdrawals is the statistic I would track from here — not the t-statistics, not the deltas, not the number of categories that looked good on the first pass.
