Autonomous trading: extreme-move alpha survives a liquidity cut
Every robustness test is a confession. It names the failure mode its author fears most, then tries to kill it. The test behind this record was aimed at the most respectable fear in cross-sectional equity work: that a screen ranking stocks on how violently they trade is a small-cap artifact wearing a ranking's clothes. The surprise is not that the screen survived the knife. The surprise is how little blood the knife drew — and what that reveals about which statistics are worth robustness-testing in the first place.
The research board this record comes from screens thousands of names each day. One lane ranks them by intraday high-low range as a percentage of price and asks whether its top names concentrate five-day extremes: an up-tail of fifty percent over five days and a log-symmetric down-tail at minus one-third. The obvious objection is the one any quant makes on sight: rank on realised volatility and you surface the smallest, thinnest names that clear the gates, and those names move fifty percent for reasons that have nothing to do with the ranking being informative.
How often does anyone bother to test that objection instead of asserting it? An audit-oriented evidence map of 77 LLM-trading studies found that of the 19 that met a closed-loop evaluation bar, only 1 documents universe or survivorship handling at all (Xia et al., 2026). Universe handling is the unglamorous act this whole measurement exists to perform — and the literature audit says publishing it is the exception, not the default.
So here is the answer in the form the question deserves: the extreme-move lift persists after excluding the smallest, least-liquid names. It is not a small-cap illusion; it survives a liquidity filter. What matters is the size of that survival — and why the survival is so much larger than the standard small-cap story would predict.
The Standard Remedy Had Never Been Run on the Statistic It Was Meant to Police
The board's standing response to "small-cap artifact" is blunt and conventional: drop the two lowest dollar-volume deciles of each date's screened names, by count, and recompute. That remedy had been applied to every mean-spread lane the board runs — a family of lanes measuring average returns, where thin names are known to carry the book — and it had never been applied to a tail statistic. The earlier record that produced the 20.57× up-lift ends with a caveat that belongs in a museum of honest footnotes: "whether this was a composition effect or a ranking effect was unmeasured."
Applied at last, and applied correctly, the test crossed two cut depths — none, and the two-decile drop — with the two spans the board has ever scored. The full record reaches back to 2022-10-10 and holds 975 scored dates. A rebuild in between had grown grouped days from 531 to 1,303 and screened dates from 236 to 980, so both arms were recomputed from code rather than one being read off the older log. Documenting that panel surgery is precisely the disclosure Xia et al. (2026) found missing from all but one of 19 closed-loop studies: a number without its universe history is a number without a parent.
The uncut arm on the trailing window reproduced the earlier record field for field: 141 hits on 5,307 name-days, hit rate 2.6569%, lift 20.5747, down-lift 30.3211, excess 2.4853 percentage points, 98 dates with a hit. That is the most boring result in this entire post, and the most important. A rebuild that prepended history and disturbed no recent row is the precondition Zhang et al. (2026) demand before their one-switch protocol tests mean anything: hold the data panel fixed, then toggle one convention at a time. Field-for-field reproduction is what holding the panel fixed looks like when it actually happens.
A Rebuild Masqueraded as a 13% Lift — the Span Moved, Not the Cut
Here is the trap that would have inverted this article's conclusion.
The earlier record said 20.57. This run's full-span uncut arm says 23.23. Line those two numbers up and the difference is 13% — and anyone who believed the only thing that changed was the robustness cut would write down that the cut raised the lift by 13%. That conclusion is exactly backwards. The cut was never in that comparison. The 13% is the span: 975 dates versus 231, a different experiment that happens to share a lane's name and a lane's output format.
The fix is as dull as the error is seductive: report both axes, every time, and never read one arm of a rebuilt panel against another arm of an older one. Zhang et al. (2026) formalised exactly this discipline as a one-switch benchmark: toggle a single evaluation convention around a clean reference while holding the data panel, split, model family and cost convention fixed, and you can see which protocol choices actually move results. Across two daily-OHLCV equity panels, six model families and yearly tests from 2016 to 2024, they found protocol-induced inflation is highly selective — moving execution to a same-day open with post-open daily-bar information caused large, stable inflation, while global normalisation and same-day-close execution were weak in most settings.
The lesson cuts in the other direction too. When you fail to hold the panel fixed, a convention that changed nothing can look like it moved everything. A reader who trusted a single span would have published a headline built from a rounding artifact of the rebuild calendar — and the cost of that error is not embarrassment. It is a published conclusion that is the mirror image of the truth.
What Survives the Cut: 92.2% Up, 92.0% Down, Asymmetry Intact
Now the numbers the headline is about.
On the trailing 231 dates at K=25, the two-decile cut took the up-lift from 20.57 to 18.96 — retaining 92.2% — and the down-lift from 30.32 to 27.88 — retaining 92.0%. The down/up asymmetry, the shape that makes this lane strange in the first place, moves from 1.474 to 1.470. That is not a survival against the odds; it is an almost perfectly elastic response to having a fifth of the population removed.
The literature gives a reason to expect the opposite. Fan et al. (2025), benchmarking autonomous agents live across US stocks, A-shares and crypto, found that AI strategies achieve excess returns more readily in highly liquid markets than in policy-driven ones — and that general intelligence does not automatically translate into trading capability. Read that as a regime statement and a liquidity cut should hurt whatever edge a mechanical screen has, because the cut moves you toward the market conditions where edges are supposed to be real. Instead, the tail concentration does not need the illiquid tail. It holds at depth: at K=50 the retention is 95.1% up and 92.5% down; at K=100 it is 93.9% up and 84.5% down. A fifth of the universe, counted by the thinnest dollar volume, was removed, and the ranking still concentrates extremes with the same directional asymmetry.
One-sentence summary for anyone skimming: the cut barely bites this lane, and the lane that most resembles a volatility-rank small-cap story is the lane that barely bleeds.
A Rubber Knife Proves Nothing — This One Cut the Reversal Lane Nearly in Half
A robustness test that never bites anything is a rubber stamp. So before reading the 92% retention as meaningful, ask whether this cut has teeth. It does — and the lanes it bites are the lanes the small-cap story was always about.
On the same run, the same cut, the one-month reversal lane retained only 65% of its up-lift and 57% of its down-lift at K=25, on 94 up hits and 175 down hits. On the mean board, the identical cut had already erased that lane outright: from −119.8 bps to −11.0 bps, with 91% of the pre-cut effect sitting in bottom-two-decile names. The settlement-fails lane retains 71.6% up and 63.2% down — damaged, though not demolished.
That pattern is the signature of a real instrument, and it is the same shape Zhang et al. (2026) document at the protocol level: a single toggled convention produces large, stable effects in some settings and almost none in others. Convention effects do not spread evenly across signals. They pick on particular structures. The two-decile cut picks on the reversal lane and the settlement-fails lane, and barely registers on the intraday-range lane. A cut that discriminates this cleanly between lanes is a test, not a ritual — which is what makes its verdict on the range lane worth reading at all.
The Lift Is a Ratio, and Ratios Eat Robustness Cuts for Breakfast
Why does one lane retain 92% while its neighbour retains 57%? The honest answer is not that the range lane is more real. It is that the statistic under test is a ratio — and the cut removes tail events from both the numerator and the denominator before the division happens.
Watch what the cut does to the underlying rates. The screened universe's own up-hit rate falls from 0.1291% to 0.1177% after the two-decile drop; its down-rate falls from 0.1703% to 0.1334%. The top-K name-days the lift is computed over shrink by the same stroke. A lift divides one rate by the other, so a cut that removes rare events from both sides attenuates the ratio far less than it attenuates either level.
The level statistic proves the point. The per-date excess hit rate — a level, not a ratio — retains only 84.8% of its value, falling from 2.4853 to 2.109 percentage points, while the lift retains 92.2%. Even the Newey-West t-statistic at bandwidth 10 slips only from 6.14 to 5.65. A robustness cut that removes a fifth of the population costs the ratio about eight percent of its value and costs the level twice that. The mechanism is arithmetic, not alpha.
The practical moral follows directly: if the alternative hypothesis is "this signal is carried by the cohort you are about to delete," choose the statistic that is a ratio before you run the cut — because a level will overstate the damage and a ratio will overstate the safety. Zhang et al. (2026) showed that how much a convention toggle hurts depends on the convention. This run shows it also depends on the metric: same toggle, same lane, 84.8% retention on the level and 92.2% on the ratio. Report both, or you are reporting whichever one flatters your prior.
Composition Effects Leave Fingerprints, and This Lane's Are Clean
The cut alone cannot say where the hits came from; it only says the lift survives their removal. Composition effects, however, leave fingerprints — and in this run they were measured directly on the uncut top-25 rather than inferred from the cut's aftermath.
In the intraday-range lane, 24.7% of top-25 name-days are bottom-two-decile names. Those names contribute 24.8% of its up hits and 31.0% of its down hits. That is near-proportional: the cohort the liquidity objection points at is present in the lane at roughly the rate it appears everywhere, and it is not producing a disproportionate share of the tails.
Now look at the lanes the cut demolished. The settlement-fails lane's top-25 is 62.7% bottom-two-decile names — nearly two-thirds of the ranking drawn from the thinnest cohort. The one-month reversal lane is 43.0% such names but takes 58.9% of its down hits and 52.1% of its up hits from them. One-third of the population producing three-fifths of the events is what a composition effect looks like when it is real. The intraday-range lane does not have one.
There is an irony worth letting sit. The range lane is the board's worst lane on mean spread — the lane a mean-spread researcher would most suspect of being a micro-cap mirage. By direct composition measurement, it has the cleanest tail attribution on the board. So, to answer the question the SEO briefs and the skeptics keep asking: is this extreme-move premium driven by size or by liquidity? Measured on dollar volume, the answer is neither. It is a property of the ranking, and the ranking survives the removal of the cohort that was supposed to be doing all the work.
Decompositions that locate an effect inside a subgroup are the tool that exposes mechanism. Elangovan (2026) documents a strategy whose entire apparent edge sits inside the retrospectively-known rank of an observation: rank 1 returns +0.13% mean net with a 51.8% win rate, while rank 2 returns +0.003% with a 43.8% win rate, even though rank-2 observations remain substantially more extreme than a randomly sampled bar. That is an effect hiding in a subgroup, exposed by decomposition. The decomposition here does the opposite work: it shows the lift is not hiding in the subgroup the liquidity cut was designed to purge.
Four Times the History Reproduced the Same Asymmetry — While Thin Lanes Stayed Noise
The earlier record made a falsifiable prediction: after the next panel rebuild, the K=25 down hit count would stay strictly above the up count, and the down-lift above the up-lift. The rebuild added 744 dates. What came back, over the full 975 scored dates, was down hits at 1,169 against 573 up hits, and a down-lift of 33.80 against an up-lift of 24.01. The asymmetry that showed on 231 dates survived a panel more than four times as long. That is the direction of travel a real ranking effect should show — a pre-registered directional bet that cleared the extension.
Noise deserves the same honesty as signal. The short-interest-flow lane rests on 11 up hits and 3 down hits at K=25; the days-to-cover lane rests on 6 and 8. Their retention swings are arithmetic on single-digit numerators, and they should be recorded, not read. Two lanes on the board are absent rather than zero: their coverage reads 0.00 because their source does not reach back over this panel, and no backfill can touch that. And the same-date base rate includes the top-K rows themselves: at 25 to 100 names out of roughly 2,800 screened, the contamination is under 4% of the denominator, and it biases the excess downward — against the finding, not for it. The disclosure of that downward bias is the kind of universe-construction detail Xia et al. (2026) found in exactly one of nineteen audited closed-loop studies.
Survival Is Not an Acquittal: the Trap Paper Proves Why
This is the paragraph that keeps the headline honest. A signal surviving a liquidity cut is evidence about one specific alternative explanation — and nothing more.
Elangovan (2026) reports a systematic, pre-registered-style search over 32 candidate short-horizon equity strategies on the Singapore Exchange. One intraday reversion strategy passed walk-forward validation, an untouched holdout, liquidity-capped position sizing and a ten-item adversarial audit — and is provably untradeable, because its entry rule requires information that does not exist at decision time: whether a larger deviation will occur later in the same session. The edge collapses from rank 1 (+0.13% mean net, 51.8% win rate) to rank 2 (+0.003%, 43.8%) even though rank-2 observations are still substantially more extreme than a randomly sampled bar. Every robustness gate the industry trusts said "pass." The market said "you were looking backwards."
Apply that discipline here. This lane's hit rate is not a strategy: no holding period, no turnover, no compounding, no drawdown, no cost model. The intraday-range lane remains the worst lane on the board's mean spread. Only one cut depth was run — none and two deciles — and a dose-response over deeper cuts is the difference between "the cut does not bite" and "this cut depth does not bite." The record already queued its own falsifiable prediction: at a half-universe cut, the K=25 up-lift should stay above 12 and the down/up ratio inside [1.2, 1.7]. If the lift collapses below 8, or the asymmetry inverts, the two-decile cut was too shallow and the ranking reading offered here is wrong. That test exists so this article can be wrong on evidence rather than defended on rhetoric. The disqualifying question from Elangovan (2026) applies before any of that: does any rule in this signal require information that does not exist at decision time? A hit rate is not a holding period, a cost model, or a drawdown curve — and if the question is "can I trade this," the answer on the evidence available here is: not yet measured, and the lane's mean-spread record gives no reason to assume yes.
No p-value was stamped on the record, no significance flag set. Point estimates sit next to three named standard-error conventions and one bootstrap interval, and the Newey-West t is printed next to its standard error so the reader concludes — deliberately. For a number that a human has an incentive to believe, leaving the verdict unautomated is the only honest reporting.
The Machine That Ran the Cut on Its Own Number Is the Story
The measurement here was chosen, specified, run and recorded by an autonomous loop that ticks on a timer, selects one open question from its own board, and is forbidden to write verdict adjectives into prose. Where both legs of a test come out clean, it records the legs and the evidence and queues the adjective for a human. The result is a robustness test the agent ran against a number the agent itself had produced — with a pre-registered falsifiable prediction attached and no verdict emitted.
Put that under the taxonomy Fang et al. (2025) propose for self-evolving agents — a feedback loop over System Inputs, Agent System, Environment and Optimisers — and the consequential act is the Optimiser declining to change the Agent System. The loop measured its own Environment, found the objection did not hold, and changed nothing. The self-evolving-agent literature catalogues techniques for changing agents and treats the evaluation of those changes as an open problem. A loop that can interrogate its own output, attach a falsifiable prediction to it, and then refuse to upgrade itself on the basis of its own measurement is one concrete answer to that open problem.
The same system shows how far that discipline extends. A learned fill-hazard model carries genuine skill — Brier skill against the base rate of +0.0784 and +0.0731 on its two instruments — yet it still beats "just cross the spread" on 0 of 8 panels, because it acts on only 0.82% and 5.60% of fills and is identical to the baseline everywhere else. Deleting every learned model leaves the baseline intact. The estimate is real; the decision has nowhere to go. Two independent measured refusals is a pattern, not an anecdote — and in both cases, the system that ran the removal test is the system being removed from. That is the discipline Fang et al. (2025) describe as an open problem in agent evaluation, executed rather than theorised.
It is also why the dull disclosure matters. Of the 19 closed-loop studies Xia et al. (2026) audited, only 2 report time-consistent split protocols, only 1 an explicit transaction-cost model, only 1 documents universe or survivorship handling, and 15 are coded R0 — meaning nothing is reproducible from the paper alone. The field's scarcest resource is not a better model. It is the act of reporting the universe you screened, the cut you ran, and the prediction you attached before you looked.
What You Actually Do With This: Five Questions Before You Trust a Screen
The findings resolve into a short decision procedure for anyone with a screened signal and a liquidity objection pointed at it.
First, ask what statistic you are defending. If it is a level — an excess return, an excess hit rate — expect a liquidity cut to hurt far more than the underlying signal warrants, because the cut removes events from the top-K side without a matching removal from the base rate. If it is a ratio, expect attenuation in the single digits to the mid-teens, and do not read that as proof of robustness by itself. Compute both before you cut; the gap between their retentions is information.
Second, verify the cut has teeth on your own board. Run the identical cut on a lane you have independent reason to distrust: here, the one-month reversal lane lost roughly a third to over two-fifths of its lift, while the range lane lost about eight percent. A cut that bites nothing tells you nothing — and a lane that survives a cut that visibly maims its neighbours has passed the only meaningful version of the test.
Third, count raw hits before reading retention percentages. Lanes resting on 11 and 3 hits, or 6 and 8, are noise with a percentage sign attached. Report them or drop them; never interpret them.
Fourth, hold the panel fixed and recompute every arm. If the underlying data was rebuilt, a 13% swing can materialise out of the calendar alone, and comparing an old arm to a new arm will manufacture either false hope or false doom. Both axes, every time — that is not bureaucracy, it is the entire difference between the 92.2% result and its mirror image.
Fifth, ask the disqualifying question before the validating one: does any rule in this signal require information that does not exist at decision time? The hindsight-selection trap of Elangovan (2026) passed every standard gate in the book and failed only on that single question — a question too few pipelines ever ask. Survival is not an acquittal. It is an invitation to the next test.
The 92% retention is real, the asymmetry is stable across four times the history, and the composition fingerprint is clean. The extreme-move lift on this lane is a ranking effect, not a small-cap illusion.
What I trust most in this record, though, is not those numbers. It is the constraint the loop operated under while producing them: no verdict adjectives, no significance flag, a falsifiable prediction written before the cut was run, and the knife held by the same system that had produced the number under test. The verdict on whether any of this is tradable was deliberately left unautomated — which, on the evidence here, is precisely where it should stay.
