An AI That Audits Trading Alpha
Take a statistic that cannot exist and give it a p-value that means something else. That is what one paper in the queue did: it reported Spearman rho = 0.94, p = 0.017 over five assets. On five untied ranks, rho lives on a finite grid spaced exactly 0.1 apart. The smallest two-sided p the test can produce is 0.0167, and 0.017 is the exact p-value of a perfect ranking. The nearest attainable rho, 0.90, carries p = 0.0833 — not significant at 5%. The claim is not subtly wrong; it is printed arithmetic that could not have come from the test the paper claims to have run.
The system that caught it is not another return-predicting model. It is an auditor: a loop over a local corpus of 21,305 quant-finance paper abstracts, with 21,119 still queued, 29 papers read end to end by a human, and 82 machine screens completed. Each tick claims one paper, asks a language model two questions about it, runs deterministic nulls against real market data, and records a verdict under a schema that refuses records which certify themselves. The most important thing I can tell you about this loop is not that it found fake alpha. It is that its ceiling is the corpus, not the model — and that honesty about that ceiling is the actual product.
