Skip to main content

Autonomous trading: news does not predict stock direction

· 22 min read
Vadim Nicolai
Senior Software Engineer

Twelve combinations of event feed and window, measured: a catalyst does not separate the up tail from the down tail on this equities screen. The one arm that crossed t = 2 was logged as a lead and refused promotion.

Does a catalyst separate winners from losers? The measurement cannot say. No separation was detected across the twelve feed-and-window combinations, but the test can only rule out an effect larger than about 2.5pp — see the correction below. Measured: catalyst and no-catalyst names posted nearly identical spreads. The one signal that looked like it did — insider purchases within five days — was logged as a lead, priced against its own sixteen-test background, and refused promotion.

What is a feed-and-window combination? The feed is the event class — any event, an insider filing, an 8-K, or news. The window is the post-event lookback over which the outcome is measured. Four feeds crossed with one-, two- and five-day windows made the twelve combinations swept.

Most write-ups of a signal search end with the winning signal deployed. This one ends with the winning signal in a drawer, and the drawer is the point.

The system that ran the search is an autonomous loop. It picks its own questions, measures them, records verdicts, and is allowed to change its own lanes, constants, modules and gates. It surfaced one genuinely significant result in its own investigation: open-market insider purchases within five days, a +3.11pp spread delta at t = 2.76. Then it declined to promote that result. It logged the result as a lead, cited its own coverage gate, and changed nothing. Fang et al. (2025) frame self-evolving systems as a feedback loop over System Inputs, Agent System, Environment and Optimisers, and note that most deployed agent systems stay static after deployment. This is that loop closing in the direction nobody writes up: the optimiser ran, produced a candidate, and refused it.

Two measured refusals later, I trust the refusal branch more than I trust any catalyst in the data.

The Lane Itself Finds Volatility, Not Direction: 20.5× Up, 24.9× Down

The intraday-range lane on this cross-sectional screen has a known defect, and the defect is why the catalyst question was worth asking. It is nearly as good at finding a name that falls 33% over five sessions as one that rises 50%: up lift 20.5× against down lift 24.9×, with the per-date spread for the lane as a whole sitting near −2.8 percentage points.

The lane does not select direction. It selects extremity, and the extremity leans down — a screen calibrated to catch a −33% collapse will also catch a +50% run, because both clear the same volatility threshold. Every numeric fix attempted on the lane has failed, which is why the question was always about conditioning: if any event subset could tell the two tails apart, this was the lane that needed it. The conditioning-information literature makes that exact promise: some instrument subsets do improve out-of-sample pricing, but only after adjusting for the horse-race over all the subsets tried. Other subsets underperform the unconditional model badly.

Any Catalyst Within Five Days Moved Nothing: +0.02pp, t = 0.03

Until the day before the run, the question was untestable. The screen needed structured event data stretching back 400 days, and the panel covering news, 8-K disclosures and insider Form 4 filings did not land until the previous day — all 236 panel dates of it. Stating that span matters more than it should have to: in Xia et al. (2026), only 1 of 19 primary agentic-trading studies documents its universe or survivorship handling at all.

The headline is a null the test could not resolve. Coverage was 35.0%; the per-date spread with any catalyst within five days was −2.84pp (standard error 0.71); the spread without one was −2.86pp (standard error 0.52). The delta is +0.02pp and the t-statistic is 0.03.

Corrected after publication. This paragraph originally called that "a clean, properly powered null", and that was wrong. Coverage and power are different things. The 95% confidence interval on the delta is [−1.70, +1.74]pp and the minimum detectable effect at 80% power is 2.46pp, against a lane whose own spread is −2.84pp. The test therefore rules out a catalyst effect larger than about 2.5pp, and rules out nothing smaller: an effect of 1.5pp — more than half the lane's entire skew, and commercially very large — sits inside the interval. The honest status is unresolved, not refuted. A t-statistic with no interval beside it is absence of evidence dressed as evidence of absence.

The two groups are the same population wearing different labels. That flatness is exactly what the conditioning horse-race warning predicts before you correct for all the subsets you searched. Some subsets will look like they matter, and most will not. The unadjusted average over the sweep will be a shrug. The paper that disagrees with our headline still concedes the essential point: this is exactly the adjustment the 12-combination sweep did not make, and no subset can be called a winner without it.

The sweep itself was small but complete: four feeds — any event, insider, 8-K, news — crossed with windows of 1, 2 and 5 days. Nothing reached |t| = 2. The highest value in the grid was the insider feed at 1.37. That grid is precisely the horse-race over subsets whose unadjusted winner the conditioning literature declines to credit.

A board that promoted catalyst feeds anyway, on a directional hunch, would have taken a lane already losing −2.8pp per date and layered a 35%-coverage filter on top of it: the spread would not have changed, but the sample on which the board accepted that spread would have shrunk by 65%. That is the cost of being wrong in the optimistic direction — a known defect, confidently re-deployed with less data. The Hindsight-Selection Trap study (Zenodo 21754442) documents exactly how much damage a near-signal can do when it survives at exactly one point in a search space: its pre-registered-style hunt over 32 candidate strategies found the edge concentrated almost entirely at rank 1 of the observation — +0.13% mean net return and a 51.8% win rate — before rank 2 collapsed to +0.003% and 43.8%, indistinguishable from random. A result that lives at exactly one window, or exactly one rank, is the shape that paper names.

The Unfiltered Insider Feed Was 94.8% Not a Purchase — and the First Sweep Proved That, Not the Null

Here is where the first measurement embarrassed my first conclusion. The insider arm of the sweep looked like a null — best |t| of 1.37, below every threshold that matters — and it deserved to look that way. On a single measured day, the Form 4 feed contained 989 rows: 406 open-market sales, 239 option exercises, 202 grants, 91 tax withholdings and only 51 open-market purchases. The feed I had labelled "insider activity" was 94.8% something else — 938 of 989 rows were sales, exercises, grants and withholdings — and the signal the literature actually cares about, insider buying, was 5.2% of the rows being averaged into it. Surveys of these systems, Ding et al. (2024) among them, catalogue architectures and data inputs side by side; the input definition is the part that decided this result.

This is the definitional trap that makes a feed-and-window result untrustworthy before statistics even enter. Changing which rows count as "insider" changed the entire answer; keeping sales in the denominator was quietly asserting that insider selling and insider buying carry the same information. They do not. The trading-agent surveys keep framing these systems as decision pipelines over data inputs, cataloguing architectures and backtests side by side — Ding et al. (2024) summarize the common architecture and inputs in exactly that frame — but an input definition is not a plumbing detail. It is the strategy.

The same hierarchy showed up in the live multi-market Agent Market Arena benchmark. Qian et al. (2025) ran 4 agent frameworks across 5 model backbones and found the frameworks displayed markedly distinct behavioral patterns while the model backbones contributed far less to outcome variation. The frame around the data dominated the window and the event type here in the same way. The first sweep was not proof of the insider null. It was proof that a signal can be drowned by its own feed definition.

Purchases Within Five Days Was the Only |t| Above 2 in the Investigation — and the Table Explains Why It Stays a Lead

Filter the insider feed to transaction code P with acquisition — open-market purchases — and one window lights up.

WindowCoveragenSpread with purchaseSpread without purchaseDeltat
5 days1.7%91+0.34pp−2.77pp+3.11pp+2.76
10 days2.8%152−1.22pp−2.78pp+1.56pp+0.91
20 days4.9%269−1.79pp−2.82pp+1.03pp+0.55
30 days6.7%364−1.17pp−2.92pp+1.75pp+0.88
The wider windows are the corroboration test the Hindsight-Selection Trap study argues is mandatory before an edge is believed.

At five days, the spread with a purchase is +0.34pp against −2.77pp without one. The delta is +3.11pp at t = 2.76, the only |t| above 2 in the entire investigation. That number is only interpretable against its siblings — the same reading that made the Hindsight-Selection Trap result diagnosable, where rank 1 returned +0.13% mean net return and rank 2 collapsed to +0.003%. The delta halves by the 10-day window (+1.56pp, t = 0.91), sinks to +1.03pp by 20 days (t = 0.55), then drifts back up at 30 days (+1.75pp, t = 0.88). The early decay from 3.11 to 1.56 to 1.03 has the shape of a short-lived catalyst rather than noise — but the 30-day uptick breaks the clean monotonic story, and the t-statistics in every wider cell stay below one. A strictly decaying effect dissolving into noise is the pattern you would expect from a brief, real information event. It is also the pattern a single lucky cell would show. The table cannot tell those apart; that is what the other three rows are for.

The mechanism, at least, is specific. Insider buying does not create upside. It removes the downward skew that defines this lane, moving the per-date spread from roughly −2.8pp to roughly zero — it neutralises the defect rather than flipping it into alpha.

The date semantics caveat is inseparable from this result. Insiders file late and variably, so the board used filing dates — the only date knowable at ranking time — while the underlying event is older by an unmeasured amount. Which date you treat as knowable at decision time is exactly the kind of one-switch convention that separates a measured effect from an artifact, the same convention Zhang et al. (2026) toggle one at a time around a clean t+1-open reference in their leak-detection benchmark: across two daily equity panels, six model families and yearly tests from 2016 to 2024, they found protocol-induced inflation is highly selective, with centered temporal features and same-day-open execution inflating large and stable amounts while other conventions stay weak. Filing-date-is-not-event-date is not a footnote. It is the whole game.

The Promotion Test Failed Itself: 1.7% Coverage, a 15% Gate, and One 2.76 in a Field of 16

The board's own record gives three reasons the +3.11pp stayed a lead — and the third is the one that separates this loop from most published research.

Coverage fails the gate. The 5-day purchase arm covers 1.7% of panel dates — 91 picks across 236 dates — against the board's own 15% power gate. The gate was set before the sweep ran. A signal that cannot cover enough dates to matter has significance that is decorative. A gate fixed in advance is the kind of protocol Xia et al. (2026) found almost entirely missing: 2 of 19 studies report extractable time-consistent split protocols, and none reaches their top reproducibility tier.

The multiple-testing context is unforgiving. Roughly 16 variants were run in this investigation. Under a complete null, 16 independent tests will produce at least one |t| above 2 by chance about 56% of the time — 1 − 0.95^16, a back-of-envelope figure that makes a single 2.76 entirely unsurprising. The Hindsight-Selection Trap study (Zenodo 21754442) passed walk-forward validation, an untouched holdout and a ten-item adversarial audit, and its edge was still provably untradeable because it lived at exactly one rank of the observation. A signal that survives at exactly one window is the same shape.

The effect lives at exactly one window. The purchase arm is null at 10, 20 and 30 days. No corroborating cell, no monotonic confirmation across the full table, no second place where the mechanism shows up — the single-rank signature the Hindsight-Selection Trap study showed can pass walk-forward validation, an untouched holdout and a ten-item adversarial audit and still be untradeable.

The cost of a laxer gate is not abstract. A threshold of |t| = 1.37 — the best the unfiltered insider feed could offer — would have promoted a feed that was 94.8% sales, exercises, grants and withholdings. Promotion at that bar would not have been optimism; it would have been an accounting error with a t-statistic attached. The board changed nothing. No lane, constant, module or gate default was touched.

Filing Date Is Not Event Date, and the Control Arm Has a Hole

Two recorded caveats would survive even if the multiple-testing problem vanished, and both cut against the finding rather than for it.

The first is the date problem already named. The second is a control-arm defect: planned-sale filings are not excluded from the no-purchase control group, so the "without purchase" arm contains names with insider sales — and those sales may be doing part of the work that makes purchases look special. The board's own record notes only that those sales may be doing some of the work; it puts no bound on how much. The direction is inferable — a control arm holding a bearish signal flatters the treatment arm — so +3.11pp is more likely an overstatement than an understatement, but that is my inference, not a measured quantity. Which rows count as the control is itself a one-switch convention of the kind Zhang et al. (2026) isolate. There is a version of this investigation where the control is clean, the event dates are known, and the 5-day result vanishes — and the board refused to spend capital betting against that version on 91 observations.

Convention sensitivity shows up in the lane's standing nulls too. The short-interest flow lane clears the engine's floor only on the IID standard-error convention: t = 2.60 under IID, versus t = 1.85 under Newey-West at lag 4. This board has itself measured the IID convention to be understated. That a threshold is cleared under one standard-error convention and missed under another is exactly the protocol-induced selectivity Zhang et al. (2026) measure by toggling one convention at a time. Neither number is a win, and the fact that one clears a threshold while the other does not is the entire argument for treating evaluation conventions as part of the hypothesis. Microcap specificity was never tested — no size split was run. The down tail is defined as −33% over five sessions, an inherited convention. The spread is a per-date rate difference, not a return. Every one of those choices is a switch that could be toggled, and none of them was.

The Second Refusal Retired a Claim the System Had Already Believed

If the catalyst refusal were the only one, it could be dismissed as one underpowered investigation. It is not, and the second refusal is the stronger evidence that the loop's governance works.

The board's one frozen learned exemplar — a gradient-boosted fill-hazard model — was refit on real exchange data against a classical logistic baseline in the same script. The learned model's out-of-sample skill advantage over the classical baseline shrank as data grew: +0.028 at 8,000 panel rows, +0.024 at 24,000, +0.006 at 72,000. The advantage shrinks toward the classical baseline as evidence accumulates — the same lesson Fan et al. (2025) draw across six models and three markets, that raw model capability does not automatically become trading capability.

An earlier internal claim had attributed a "more data, more negative" pathology to that model. On real data, the claim did not replicate: the model's skill was stable at +0.185, +0.208 and +0.175 across the same three data sizes. The pathology was a fixture property, not a market property — and the older claim is now retired. That refusal branch, the optimiser that runs, evaluates, and declines to apply its own candidate, is precisely the branch that Fang et al. (2025) catalogue in their self-evolving-agents survey as designed but almost never reported.

This is the reversal nobody wants to publish. A "more data makes the model worse" result would have justified freezing data collection and excused every stagnant system; a fixture bug producing that illusion would have gone unnoticed forever in a loop that only reports wins. Fan et al. (2025), benchmarking 6 mainstream LLMs across 3 markets — U.S. stocks, A-shares, and cryptocurrencies — found that general intelligence does not automatically translate into trading capability and that risk control, not raw skill, determines cross-market robustness. The parallel is exact: the property that looked like market structure was actually model plumbing, and the system only discovered that because it was built to test its own refusals.

Two Refusals Are a Pattern: the Self-Evolving Loop That Says No

The self-evolving-agents literature frames these systems as a feedback loop over System Inputs, Agent System, Environment and Optimisers — and notes that most deployed agent systems remain static after deployment. What it rarely reports is the loop closing in the negative direction: the optimiser producing a candidate improvement, evaluating it, and declining to apply it.

Two measured refusals on different kinds of evidence — one an underpowered catalyst arm, one a retired model claim — are a pattern, not an anecdote. The contrast with the surrounding field is stark. Xia et al. (2026), in their evidence map of 77 agentic-trading studies, had to disqualify 58 studies for failing the minimum boundary of Action Output plus Closed-Loop Evaluation before they could call 19 a primary subset. Of those 19, only 2 report extractable time-consistent split protocols, only 1 documents an explicit transaction-cost model, 15 are coded R0 for reproducibility, and none reaches R3. The field's immediate bottleneck is comparable evaluation, not architecture — a point echoed from the evaluation side by Yehudai et al. (2025), who argue that LLM agents need new evaluation methodologies, benchmarks, environments and metrics beyond standard text-to-text tasks. Mohammadi et al. (2025) make the same point from a different angle: they split agent evaluation into evaluation objectives and evaluation process, and flag dynamic, long-horizon interactions as an under-addressed gap.

The breadth of the architecture work makes the gap worse. Surveys cataloguing LLM trading agents — Saha et al. (2025) map portfolio optimization, risk management, information retrieval and automated strategy generation; Ding et al. (2024) survey the common architectures and data inputs and report the performance achieved in backtests — describe a field rich in designs and poor in the governance that would tell a designer whether a design earned promotion. A loop that records its own coverage gate, its own variant count and its own refusal is doing the thing the surveys find almost nobody does. The danger of self-evolving systems was always that they would optimise their way into self-deception: fitting the evaluation, promoting the marginal, deleting the null. What the board measured is the opposite failure mode — the refusal branch, where the system's integrity actually lives.

Decision Framework: Six Gates Before Any Signal Earns Promotion

What generalises from this investigation is not the insider-purchase result — that stays a lead until more data arrives. What generalises is the governance. The sequence that produced two correct refusals reduces to a checklist any systematic trader can run, and every step is grounded in a number from this investigation.

Gate 1 — Set the coverage gate before you look at the t-statistic. The 15% gate refused the 1.7%-coverage purchase arm on 91 names. Fixing it in advance is what makes it a gate rather than a rationalisation — the discipline Xia et al. (2026) found in 2 of 19 studies. If a signal cannot cover enough dates to matter, its significance is decorative. Compute coverage first; if it fails the gate, stop.

Gate 2 — Count your variants and price in the multiple testing. Sixteen variants will produce a spurious |t| above 2 about 56% of the time under a true null. Record the variant count next to the result, the way the Hindsight-Selection study priced its rank-2 collapse against the 32 candidates it searched.

Gate 3 — Toggle every date convention one switch at a time. Filing date versus event date, t+1 open versus same-day open, IID versus Newey-West at lag 4: t = 2.60 became t = 1.85 on the short-interest lane under the latter switch. If the result dies under a defensible convention change, it was never alive — the one-switch method Zhang et al. (2026) build their leakage benchmark around.

Gate 4 — Test the control arm for contamination. Planned sales in the no-purchase control most likely flatter the insider-purchase effect. Ask what your control group secretly contains before you credit your treatment group.

Gate 5 — Retire failed claims in writing. The "more data, more negative" pathology was withdrawn only because the refit ran against a classical baseline on real data. Publish retractions at the same production bar as promotions. The self-evolving-agent surveys, Fang et al. (2025) among them, catalogue optimisers that improve a system; the branch that withdraws one of its own beliefs goes unreported.

Gate 6 — Log refusals with the same ceremony as findings. The 5-day purchase arm was recorded as a lead to re-run on more data, not as an effect. That record is what makes the next, better-powered run possible.

A system that cannot refuse is not a system. It is a bias with a loop around it.

Questions This Measurement Answers

What is a feed-and-window combination in algorithmic trading? A feed is the event class — any event, an insider filing, an 8-K, or news. A window is the post-event lookback over which the outcome is measured. Four feeds crossed with one-, two- and five-day windows made the twelve combinations swept here.

Why would a system refuse to promote a signal that shows predictive power? Because predictive power is not enough. The 5-day insider-purchase cell showed a +3.11pp delta at t = 2.76 — and still failed the 15% coverage gate on 1.7% coverage, sat inside a ~16-test field where one spurious |t| above 2 appears about 56% of the time by chance, and survived at exactly one window — the single-cell signature the Hindsight-Selection Trap study showed can clear every standard robustness check and still be untradeable.

How many backtest combinations do you need before trusting a signal? Twelve is a meaningful but limited grid: it produced nothing above |t| = 1.37 until an insider-purchase filter produced a single 2.76. A confident conclusion requires more than a wider grid — it requires the coverage gate, the multiple-testing correction, the convention toggles and a clean control arm. The conditioning-information literature makes the same point from the other side: subsets of instruments do earn their place, but only once the horse-race over all of them is priced in.

The Lead Stays a Lead

Catalysts did not separate winners from losers. The one signal that perked up — open-market insider purchases within five days — was refused promotion because 91 observations cannot carry a lane. The board changed nothing, and that is the result.

What the autonomous loop demonstrated instead was the harder skill: it measured a candidate improvement, priced in its own variant count, caught its own control-arm contamination, retired its own false claim, and wrote all of it down. The next time the insider-purchase arm is tested, it will be tested on more than 236 dates, against a control that excludes planned sales, with the filing-date lag measured rather than assumed. If the +3.11pp survives that run, it will have earned promotion the hard way. If it does not, the record of this refusal will still be there — which is more than most published signals can say.

Leads are for logging, not for deploying. The system that knows the difference is the one worth building.