AI-First Trading Principles for Crypto Perpetuals
An optimistic backtest can show Sharpe 5 where live reality is negative, and an AI graded by it will optimise into the gap with total conviction — because inside that simulator the strategy genuinely works.
The simulator is not a test. It is the model's reward function. That is the sharpest trap I know, and the rest of this post is what I found while walking into it.
Here is the shape of what I found. Give a router a third action — quote, cross, or abstain — and it takes the third. always_abstain scores exactly 0.0 and is the best arm in 11 of 11 panels, without consulting the signal at all. The fitted policies that do trade pick about 1.6% of rows and still end below zero. A model that has learned to almost-not-play is not broken. It is reporting the absence of an edge correctly, which is the only honest output available and the one every trade-count metric punishes.
Everything that decides whether an AI makes money on a crypto perp lives in the coupling between model and market: fees, funding, regime, and the evidence you are willing to accept. A round trip costs 4–14 bps before the model says a word.
I went looking for that edge at sub-minute horizons and did not find it. What the search produced instead was fifteen principles about how to run a model against a market — each ending with the condition that breaks it, each carrying the measured number behind it, and several carrying the number that killed an earlier version of the same claim. They are worth more than the strategy would have been.
Where these come from
Every first-person number below traces to a named experiment; the appendix at the end lists which. They come from a research programme aimed at sub-minute perpetual-futures trading, which the arithmetic closed — against a maximum gross of 1.4457 bps anywhere measured, break-even needs a taker fee of 0.8078 bps per side at best, against a 1.7 bps published floor and a 5 bps VIP0 charge. The principles are what survived that result, and none of them depends on it. The screen is a two-leg removal test: delete the AI and the principle must go vacuous, delete the market and it must go inapplicable. I ran the same test over my own codebase, where 0 of 108 modules survived both legs. Everything below is enforced in a running system or measured in one, and the numbers are mine unless a paper is named.
I. What the market does to the model
Edge is denominated in fees, and the model inherits the denomination
Before any model runs, the venue has already priced the game. On a Binance USDT-M perpetual at VIP0, maker is ≈2 bps a side and taker ≈5 bps (published fee schedule). The full round trip therefore lands between 4 and 14 bps depending on the leg mix: maker-in/maker-out ≈4–5 bps, maker-in/taker-out ≈8–9, taker both ways 12–14.
My minimum-viable-edge rule is that the expected move must clear 1.5× the worst-case round trip. That is how the design geometry became +15 bps take-profit against −6 bps stop. The realised numbers follow from the fee table: a maker-exit win nets 15 − 4 = +11 bps, a taker-exit loss nets −6 − 9 = −15 bps. Breakeven is therefore 15/(11+15) = 57.7%.
It is tempting to file this under costs and move on, but the market-design literature says it is closer to the rules of the game. Budish, Cramton and Shim (2015) argue that the continuous limit order book is not a neutral venue but a mechanism whose arms race is structural. Their proposed replacement is frequent batch auctions — uniform-price double auctions run roughly every tenth of a second. The fee schedule is the visible surface of that mechanism.
Nor should you expect it to improve on its own. Budish, Lee and Shim, modelling exchange competition and the rents from selling speed, find a wedge between private and social incentives to innovate. That wedge "supports the persistence of an inefficient market design in equilibrium." The venue's economics are not a transient you can wait out.
Even in slower asset classes the erosion is severe. Implicit market-impact costs "may substantially erode a strategy's expected excess returns," in Li, Chow, Pickard and Garg's accounting for factor strategies. A high-turnover strategy pays that toll on every one of far more round trips.
The AI-first consequence is that this arithmetic is not advice the model receives. It is the shape of the model's vocabulary. My proposer may move exactly three numbers. Their walls were computed from the fee table before any model was ever called: tp_bps in [10, 25], sl_bps in [4, 10], time_stop_s in [20, 120]. A proposal outside a wall is refused by name, never clipped. The AI cannot propose the trade the fee schedule already killed, because the fee schedule was compiled into its cage.
Breached when the model's proposal space contains geometries that lose to fees — or when a band is widened without redoing the fee arithmetic that set it.
A longer horizon buys beta, not edge
When cost dominates a signal, the reflex is to hold longer. Fees are charged per round trip and the move should grow with time, so somewhere out along the horizon axis the arithmetic ought to flip. It does flip. It flips for the wrong reason, and the way it flips is the most expensive near-miss in my record.
I swept the holding horizon and the table read like a solved problem. Gross per fill runs −0.7512 bps at 800 ms, −0.6876 at 10 s, −0.7126 at 60 s, then turns: +0.3965 at ten minutes, +2.7464 at one hour, +8.3162 at six hours. The six-hour cell clears the 3.4 bps round-trip toll — the first and only cell in the entire investigation to do so — at a t-ratio of 35.7.
It is beta. On 2024-03-28 the market moved +187.8 bps; on 2024-03-29 it moved −132.0. Decomposed against that drift, the six-hour cell is 99.0% market and 0.0831 bps of residual, and the residual does not resolve. At one hour it is 88.3% market; at ten minutes, 50.1%. The apparent edge was the market's direction arriving inside my holding period, sampled by a strategy that happened to be holding.
The selection underneath it is the part worth stealing. The rising day produced 89,750 fills against the falling day's 60,708 — 48% more — so the fill-weighted drift is positive even though the two days nearly cancel. A strategy does not sample days uniformly. It samples them in proportion to how often its own conditions fire, and those conditions fire more when the market is moving its way. Beta does not arrive as an obvious market-long position. It arrives as a sampling bias in which days you traded.
What survives the decomposition is the actual answer, and it is flat. The residual reads −0.7424 bps at 800 ms, −0.7076 at 10 s, −0.7719 at 60 s — resolved at 73.8×, 25.0× and 12.2× its own bar — and stays flat across a 75× horizon range. That is the signature of a fixed cost, not a decaying signal: adverse selection is about −0.72 bps paid at the fill, and everything above it was the market. The best residual anywhere is +0.3202 bps at one hour, which is 10.6× short of the toll and only 1.10× its own error bar, on two days.
So the horizon lever does not work, and it fails in a way that would have published. A sweep across horizons is a search, and the cell it returns is the one where beta is largest — reported with a t-statistic that grows as the contamination grows. A t of 35.7 on an unadjusted number is not strong evidence. It is a very precise measurement of the market.
The correction is procedural rather than statistical, and it is cheap: a directional claim names its horizon, its raw figure, the benchmark, and the benchmark's unconditional move over the same window. The literature has had the machinery for this for years — Liu, Liang and Cui identify market, size and momentum as common risk factors across 78 cryptocurrencies, and a return not adjusted against at least the market factor is not a result about a strategy. What is new here is only how fast an automated sweep finds the contaminated cell and hands it to you with a compelling t-statistic attached.
Breached when a horizon sweep reports its best cell without decomposing it against the benchmark's unconditional move over the same window — or when a rising-market sample is allowed to contribute more fills than a falling one without that asymmetry being measured.
The trade's sign is an experiment, not a parameter
My strategy template enters after a liquidity sweep — a burst of aggressive orders clearing one side of the book — and follows it, betting on continuation. The best live evidence on the venue says the money is in counter-trading it — reversal. That disagreement is genuine and current. It is also exactly the kind of question an AI operator will try to resolve by quietly flipping a parameter. So the sign is structurally not proposable. direction carries no clamp band, and the proposer's no-band branch refuses it by name the moment a proposal mentions it.
The reason a flip cannot be a tuning step is arithmetic, and it is easy to get wrong. Long gross is dmid − 2·half; short gross is −dmid − 2·half. Inverting the signal reflects the mid move and keeps the cost — the round-trip spread is paid in both directions. So a large negative result is not a large positive result waiting to be turned around. It is mostly the cost, and the cost does not have a sign.
I have the measurement that makes this concrete, and it is a trap I would otherwise have walked into. On XRPUSDT, CVD comes back resolved negative in 10 of 10 cells, at −1.4073 to −1.7343 bps — which reads irresistibly like an inverted edge of the same magnitude. It is not. XRPUSDT's round-trip spread is 1.6076 bps, so the mid move actually hiding behind that −1.4073 is 0.2003 bps.
Flip the signal and the best short gross is −1.4809 bps: still negative, because you kept the spread and gained two-tenths of a basis point. Across all four instruments the best short gross is +0.0423 on BTCUSDT, −0.0158 on ETHUSDT, −0.0427 on DOGEUSDT and −1.4809 on XRPUSDT. The short side is worse than the long side on three of four, and the single positive is 80× short of the 3.4 bps floor it would need to clear.
The underlying signal is real, which is what makes the trap live rather than academic. Cont, Kukanov and Stoikov (2013) established that order flow imbalance relates linearly to price change over short intervals. The slope is inversely proportional to market depth. The result holds across time scales, stocks and intraday seasonality. Volume imbalance predicts the sign of the next market order on Nasdaq data, and HFT trading direction predicts price changes over seconds while correlating with book imbalance.
Against that, the live Binance experiment cited earlier concludes the opposite for a maker: counter-trade the imbalance. Both can be true. Direction-of-price and profitability-of-a-resting-order are different questions, and the depth-scaled slope means the same imbalance implies a different move on a book of a different thickness. So the sign gets settled the only way a sign should: one pre-registered experiment, continuation and reversal as each other's control, spending one holdout read, on my venue at my latency. Until that run, both templates exist and neither is believed. The model may move numbers inside declared walls.
It may not change what the strategy is while nobody is looking — because a flipped sign is not a tuned strategy. It is a different strategy wearing the old one's track record.
Breached when the model can flip long/short semantics — or when a sign is adopted from a paper instead of from a pre-registered run on your own venue — or when a large negative result is read as an edge in waiting, without first subtracting the cost that is paid in both directions.
Crypto's native state is an input advantage, not an edge
The inputs that make crypto crypto have no equities analogue: the funding countdown, the mark-versus-index basis, liquidation flow — and a strange algorithmic clock. The Quarter-Hour Effect (arXiv:2607.09426) documents that clock as periodic bursts of volatility and volume at the 1-, 5- and 15-minute marks on Binance perpetuals, measured across six contracts. Opening order imbalance predicts returns over a 4–12 hour horizon, with "much weaker effects at finer clock-time frequencies." These are the only inputs that could put a trading model at the AI × crypto intersection, so they are the ones worth testing properly.
The mechanism itself is well understood, which is the first reason to be suspicious of easy edges in it. Fundamentals of Perpetual Futures derives explicit no-arbitrage prices for linear, inverse and quanto contracts, and identifies the funding specifications that guarantee spot-futures price coincidence and dynamic replication. Later work builds replicating portfolios directly and proposes path-dependent funding as a practical way to hold prices to target. Meanwhile the exploitable gaps have been closing for years: crypto arbitrage opportunities are well documented, but their magnitude fell sharply from April 2018 onward and existing price differences are now barely exploitable.
Then the trap, and it is the most instructive negative result I know in this space. Funding-Aware Optimal Market Making for Perpetual DEXs (arXiv:2605.06405) treats the funding rate as a stochastic state variable. The coupling it adds to classical market making is that inventory now creates both mark-to-market exposure and a state-dependent funding cash flow. A reduced inventory-funding control problem is formulated and solved with a monotone finite-difference Hamilton-Jacobi-Bellman scheme, with bid and ask quote offsets recovered from discrete inventory value differences. The control problem is solved with a monotone finite-difference HJB scheme. No learned model anywhere.
The holdout runs 100 seeds, calibrated on Hyperliquid ETH, BTC and SOL data. The result: "the funding-aware HJB improves mean ETH/BTC performance while lowering inventory RMS relative to classical Avellaneda-Stoikov" (arXiv:2605.06405). That is better performance and less inventory risk than the 2008 baseline — from adding one state variable to forty-year-old machinery. So the crypto-native mechanism my specs had carried as a design, and never tested, was finally built and measured. The task: forecast the next 8-hour Binance funding print for BTCUSDT and ETHUSDT. The data: 2,091 settlements each from the public funding-rate archives, on six funding lags plus settlement hour and day-of-week. Persistence, trailing mean, AR(1), ridge and gradient-boosted trees all ran on an identical expanding-window five-fold walk-forward.
Before the data was touched, the bar was fixed: beat persistence on both out-of-sample MAE and directional accuracy in ≥4 of 5 folds, pooled sign test p < 0.05. BTCUSDT returned 3/5 and 3/5 at p = 0.095 — fails. ETHUSDT returned 5/5 on MAE but 3/5 on direction — fails the conjunction. Ridge carried the better MAE in 4 of 5 BTCUSDT folds. The classical arm won the disjunction my own spec had left open.
The economics are the part that should have been checked first, and this is the discipline the whole result taught me. Mean funding is 0.4577 bps per 8 hours on BTCUSDT (std 0.586) and 0.448 on ETHUSDT. Against a hurdle of one side's fee, every policy nets ≤ 0 at 2, 5 and 10 bps per side — including the oracle.
With perfect foreknowledge, BTCUSDT makes zero position changes and ETHUSDT's oracle loses 5.02 bps net over roughly two years. The mechanism is economically inert at any realistic fee before forecast quality even enters the question. The forecasting bar was carefully specified, pre-registered and run — against a quantity that could not have paid for itself however well it was predicted.
There is a sharper caveat sitting under any funding model, and it deserves to be load-bearing rather than a footnote. Funding is a print, not a price — and prints have an attack surface. In commit-then-reveal encrypted mempools, the funding signal sets a transfer rate while receiver-side open interest is the transfer base, and a corrective transaction cannot enter an already-committed batch. Skill at predicting a manipulable print is not the same thing as skill at predicting a market, and only one of them survives contact with whoever is doing the manipulating.
The same question went somewhere else entirely, in case the failure was the instrument rather than the thesis. The data: 833,098 resolved pump.fun launches, launch-time features only, chronological split. AUPRC was pre-registered as the metric, because a 0.18–0.23% base rate makes AUC misleading.
Both models beat random by four to eight times, so the structure is genuinely there. But the classical logistic scored 0.01778 AUPRC against the boosted trees' 0.00878 — the learned arm 50.6% worse in relative terms, and failing its +10% bar by a wide margin.
The metric choice inverts the answer: on AUC the trees look better, 0.824 to 0.696. Rank the whole distribution and the learned model wins; ask which candidates to actually act on and it loses by half. That is the textbook imbalance lesson, and it is exactly the lesson a dashboard reporting AUC would have hidden.
What did finally occupy the intersection is worth stating precisely, because it is not an edge. The first genuine crypto-native × learned object in my corpus is a measurement instrument. It is a cost-attribution probe that needed a model reading a classified funding column in order to run at all. It ran to completion for the first time here.
The model powering it is the same one that had just failed its own skill bar. Both facts belong in the ledger together. A crypto-native input does not make an AI-native edge. The model leg has to be earned separately, on data, with the classical scheme as the baseline it must beat.
Breached when "we use funding/on-chain/liquidation data" is offered as evidence the strategy is AI-first without a classical control fitted on the same inputs — or when a forecasting bar is set on a quantity whose economics were never checked against the fee — or when a ranking metric is chosen that flatters the model on the part of the distribution nobody trades.
Leverage makes the tail endogenous — and mine is unmeasured
Every principle so far has carried a number I measured. This one does not, and saying so is the point of including it.
The mechanism is specific to this asset class. Crypto perpetuals carry leverage up to 100×, and BitMEX-era platforms were averaging over $3 billion in daily volume on exactly that product. Perpetuals now account for roughly 93% of crypto futures volume (Zhivkov, 2026). At that leverage, price moves force liquidations, and liquidations are themselves market orders.
The selling is caused by the price and then causes more of it. In 2021, nearly **200 million a day](https://doi.org/10.1016/j.ejor.2022.07.037). Tran, Nguyen, Le and Pham study Binance USDT-margined perpetual swaps. They find bitcoin open-interest changes are the strongest crash predictor, with an odds ratio of 1.48. It models the cascades directly against circuit-breaker counterfactuals (Studies in Economics and Finance, 2026).
One consequence deserves to frighten a risk model more than it usually does. Under auto-deleveraging, a solvent, correctly-stopped position can be closed by the venue because somebody else's losses have exhausted the margin pool. Formalised as an optimisation, ADL socialises losses among surviving participants, and a minimax-leverage policy — minimising the maximum participant leverage — is optimal under monotone risk measures in isolated margin. Read that as a trader and it says: your exit is not entirely yours.
A stop-loss models the price reaching your level. It does not model the exchange reaching into your position because the queue of insolvencies got long enough. That tail does not appear anywhere in a price series, so no model trained on price series can anticipate it.
Now the honest part. My archive holds bookTicker and aggTrades. There is no funding, mark-price or premium-index data on disk at all. So I cannot measure liquidation intensity, cascade timing or ADL exposure from my own data. Every number in the two paragraphs above is somebody else's.
A system that quietly let this pass would be making this post's error in reverse. Not believing an unsupported claim, but treating an unmeasured risk as an absent one. Those look identical on a dashboard and are opposites in a drawdown.
So it is carried as a named gap rather than a silent one. The risk layer's response to a regime it cannot price is a refusal, not a forecast. Position limits and a hard kill, sized on the assumption that the cascade is unmodelled rather than mild. The research question it generates is a purchase decision, not a modelling one. This data would have to be acquired before any claim about cascade behaviour on my venue could be graded.
Breached when a leveraged strategy's tail risk is estimated from a price series that contains no liquidation or funding state — or when "we found no cascade effect" is reported from an archive that could not have recorded one.
II. The AI as trader: regime gates, no-trade and sizing
The best trade the AI makes is no trade
The regime filter is the highest-leverage decision in the whole stack, and it is a decision not to act. In LOW volatility the expected move is smaller than the round-trip cost, so every signal is noise you pay fees to hold. In EXTREME volatility, slippage and stop-gap risk dominate the geometry. Both block. Fees are fixed, opportunity scales with volatility, and the tradable band is the middle.
The strongest evidence I have for this principle arrived by accident, when I gave a router a third action. Alongside "quote" and "cross", it could now abstain — and the baseline changed underneath the question. Abstaining is free, every arm loses money, so always_abstain scores exactly 0.0 and beats everything the investigation had produced, without consulting the signal at all.
At 2 bps maker and 2 bps taker it is the best arm in 11 of 11 panels; at zero maker and 5 bps taker, 10 of 11. Across 22 cells, the number of routers that beat abstaining by a resolved margin after correction is zero.
The way the routers failed is more informative than the fact that they did. An earlier iteration predicted they would degenerate onto always_maker — collapse to the cheapest action and stop discriminating. Given a third choice they degenerated the other way, onto always_abstain. The fitted policies trade about 1.6% of rows, at median abstention rates of 0.9897 and 0.9841. They still end below zero, at −0.0051 and −0.0149 net.
Even the sliver they select is unprofitable. The model learned to almost-not-play, and the residue of playing was still a loss.
That result needs its boundary stated with it, because it is easy to over-quote. It is not a claim that no strategy works. It is a statement about this cost structure at this horizon: the measured policies lose money, and not playing beats playing. An AI that reports "no trade" against that background is not failing to find the edge. It is reporting the edge's absence correctly, which is the only honest output available and the one a metric like fill count or trade count will punish.
The mechanics of getting the filter wrong are worth as much as the filter itself, because they fail silently in both directions. An unwarmed volatility estimate substituted as 0.0 ranks LOW. My entry gate treats LOW as merely one blocked regime.
The volatility percentile gate reads that same zero as calm, and lets trades through elsewhere. The identical fabricated number blocks one path and unblocks another. So every feature refuses when it does not know: None, never zero.
I paid for that lesson twice, measured. A p99 threshold computed over ten samples is pinned to its top order statistics: 0.99 × 9 = 8.91 interpolates between the 9th and 10th values. So my sweep detector's bar — the threshold above which a burst of aggressive flow counts as an event — was effectively "the largest bucket seen so far," and ordinary flow cleared it.
42 sweeps before the warmup floor, 30 after — every vanished one inside the warmup window. Meanwhile a volatility EWMA fed into its own percentile distribution during warmup labelled 93% of quotes EXTREME, silently refusing every sweep the detector had found. A detector and a filter, each plausible alone, cancelling to zero trades.
The regime literature is mostly an argument about how to detect the states. It has moved well past two-state Gaussian switching. There are continuous HMMs where the regime chain governs autocorrelation, per-regime emissions carry the heavy tails, and a regime-conditional VaR falls out (arXiv, 2026). There are bidirectional-LSTM hybrids for joint break identification and volatility modelling (2026), and network-community methods for early warning (PLOS ONE, 2025). A crypto-specific literature classifies latent BTC states through these structural breaks (2026).
Useful — but note what the abstain result says about the whole genre. A better regime classifier improves which rows you decline. It cannot manufacture a profitable subset out of a population whose best member is refusal.
Breached when an unwarmed statistic reaches the entry decision as a number — or when "don't trade" is treated as a failure mode instead of the most profitable output the AI has — or when a router's abstention rate is read as underfitting rather than as the answer.
