A pre-analysis research note on whether a public derivatives venue can expose a persistent negative-alpha cohort, and whether a cost-aware inverse strategy can harvest it.
Version
v0.3, July 2026
Status
Thesis and test design
Claim
Not yet a live-profit audit
Abstract
Main claim
The claim is not that every losing wallet should be copied in reverse. The defensible claim is narrower: if a wallet can be classified ex ante as a member of a persistent negative-alpha cohort, then the inverse of its future trades may have positive expected value after execution costs.
This hypothesis is grounded in a large behavioral-finance literature. Active retail trading has been shown to reduce net returns [1], investors display a tendency to realize winners and retain losers[2], aggregate individual trading losses can be economically large [3], and day-trader skill is highly dispersed with a reliably weak lower tail [4]. The open question is whether similar weak-tail behavior is measurable and tradeable in Hyperliquid wallets.
Equation (4) is the estimand: the expected net fade edge of the top Fool Score decile in excess of a matched baseline, measured out of sample.
1. Design
Estimand and notation
The estimand is the out-of-sample net fade edge of wallets selected by a pre-trade score: the inverse markout net of all execution costs, written π in the equation block. This note uses that one term throughout. The unit of observation is a closed episode, not an account screenshot. An episode starts when net exposure leaves zero and ends when exposure returns to zero or flips.
Table 1 · Notation
Wallet index.
Closed trading episode: position leaves zero, then returns to zero or flips.
Target direction of the episode. Long = +1, short = -1.
Future markout over horizon t after the target trade.
Inverse markout: the markout earned by taking the opposite side.
All-in cost: fees, spread, slippage, funding and execution failure cost.
Fool Score computed strictly before the validation window.
Net fade edge: inverse markout net of all execution costs.
| Symbol | Definition |
|---|---|
| Wallet index. | |
| Closed trading episode: position leaves zero, then returns to zero or flips. | |
| Target direction of the episode. Long = +1, short = -1. | |
| Future markout over horizon t after the target trade. | |
| Inverse markout: the markout earned by taking the opposite side. | |
| All-in cost: fees, spread, slippage, funding and execution failure cost. | |
| Fool Score computed strictly before the validation window. | |
| Net fade edge: inverse markout net of all execution costs. |
A positive research result requires that top-score wallets have worse future original-direction markout than matched controls, and that the inverse strategy remains positive after all costs. This separates behavioral selection from execution quality. The horizon parameter t is not free: Figure 1 sketches the target term structure, in which gross inverse markout must clear the all-in cost stack before some horizon for a tradable window to exist at all.
Figure 1
Inverse markout by holding horizon
Illustrative target term structure. Gross inverse markout must clear the all-in cost stack before some horizon for the strategy to have a tradable window; here the modeled breakeven sits near one hour.
Inverse markout (bps) vs horizon after target entry
Data table
| Horizon | Net of costs (bps) | CI ± | Gross (bps) |
|---|---|---|---|
| 5m | -10 | 4 | +6 |
| 30m | -4 | 5 | +12 |
| 1h | 0 | 6 | +16 |
| 4h | +8 | 8 | +24 |
| 1d | +15 | 11 | +31 |
| 3d | +12 | 14 | +28 |
2. Data
Sample construction
The sample is built by a two-stage public funnel. A daily discovery pass screens the venue's leaderboard for wallets that are active enough to score and not structurally unfadeable; survivors enter a scoring queue, where each wallet's trailing 90 days of fills are reconstructed into closed episodes and scored. The top 1,024 wallets by Fool Score form the deck — simultaneously the tradable cohort and the pre-registered test cohort. Every input is a public API endpoint [7]: fills by time, account snapshots and L2 book. Anyone can rebuild the sample.
Table 2 · Sampling and cohort parameters, pre-registered (v0.3)
| Parameter | Value | Purpose |
|---|---|---|
| Universe | Hyperliquid public leaderboard, daily snapshot | Every address visible on the venue's leaderboard each day; no privileged data. |
| Day volume floor | ≥ $10,000 traded on the snapshot day | Data sufficiency: enough fills for the window to be scoreable. |
| Account value floor | ≥ $1,000 | Excludes dust and abandoned wallets, which cannot be faded at size. |
| Turnover ceiling | monthly volume < 100× account value | Excludes HFT and market-making flow, whose fills are not behavioral. |
| Scoring window | trailing 90 days of fills | Long enough to accumulate episodes; short enough to track the current regime. |
| Episode floor | ≥ 20 closed episodes per wallet-window | Statistical floor below which a wallet-level estimate is noise. |
| Deck size | top 1,024 wallets by Fool Score | The tradable cohort, and therefore the pre-registered test cohort. |
| Matched controls | 4 per target wallet | Nearest neighbours on coin universe, turnover, account value and liquidity bucket. |
| Validation window | next 90 days, zero overlap | Walk-forward: the score is re-frozen at each window boundary. |
The screening floors and ceiling are not neutral: they define the population the thesis speaks about. The volume and account-value floors remove wallets that cannot be scored or faded at size; the turnover ceiling removes high-frequency and market-making flow, whose fills reflect inventory management rather than belief. All claims in this note are therefore claims about the screened cohort, and the published audit must report the funnel count at every stage — wallets seen, wallets screened, wallets scored, wallets decked — so that selection is visible rather than silent.
Because episodes within one wallet are correlated, wallets — not episodes — are the clustering unit for every standard error in Section 5. Wallets that go inactive during a validation window remain in the denominator until their exit is observed, which is what the exit column of Figure 2 makes explicit.
3. Score
The Fool Score
The score is deliberately boring. Each wallet-window is measured on five factor families, each anchored in a documented behavioral regularity; family measurements are winsorized, z-scored against the screened cohort, and combined with equal weights into a single monotone score. Higher means more exploitable. There is no fitted model and no interaction terms in v0.3 — a score this simple is harder to overfit and easier to falsify.
Table 3 · Fool Score factor families
| Family | Behavioral signature | Primary measurement | Anchor |
|---|---|---|---|
| Disposition | Realizes winners quickly, rides losers | PGR − PLR: proportion of gains realized minus proportion of losses realized per window | Odean (1998) [2] |
| Overtrading | Turnover far above what the account justifies | Monthly turnover; trade count per unit of account value | Barber & Odean (2000) [1] |
| Persistence | Losses repeat across windows instead of mean-reverting | Trailing original-direction markout in prior windows; decile stability | Barber, Lee, Liu & Odean (2014) [4] |
| Timing | Systematically enters just before adverse moves | Short-horizon (5m–4h) original-direction markout after entry | Barber, Lee, Liu & Odean (2009) [3] |
| Sizing | Escalates size and leverage while losing | Size-after-loss ratio; leverage trajectory inside drawdowns | Barber & Odean (2001) [5] |
Two discipline rules protect the score from its own authors. First, the recipe is frozen and hash-stamped before each validation window opens; the walk-forward test in Figure 6 is only meaningful if the score cannot move after the fact. Second, any change to a feature, weight or winsorization bound re-registers the study and restarts the validation clock. The features themselves are chosen to be expensive to fake: the only way for a wallet to stop looking foolish on these measures is to stop trading foolishly.
4. Literature
Why this is plausible
Barber and Odean document that households with the highest turnover materially underperform after costs[1]. Their overconfidence study also links higher trading activity to lower net returns[5]. Odean's disposition-effect paper shows that investors are more willing to realize gains than losses, and that the losers they keep do not subsequently justify that choice [2].
The most relevant empirical bridge is the day-trading skill paper. Barber, Lee, Liu and Odean sort traders by past performance and test the next period; the bottom-ranked tail continues to lose after fees, while only a very small top tail earns reliable positive abnormal returns [4]. This is exactly the shape a fade engine needs: avoid the skilled tail and target the persistently weak tail.
Crypto markets do not remove behavioral bias. Schatzmann and Haslhofer find evidence of the disposition effect in Bitcoin transaction data, with varying intensity across regimes [6]. That does not prove whosthefool's edge, but it supports the premise that on-chain markets can contain measurable behavioral regularities.
What the literature predicts, concretely, is a sticky transition matrix. If skill and its absence persist, a wallet's score decile in one window should forecast its decile in the next; Figure 2 shows the target shape, including the exit column that makes survivorship an explicit object of the audit rather than a silent bias.
Figure 2
Decile transition matrix
Illustrative target shape for score persistence: rows are Fool Score deciles in one window, columns the decile occupied in the next. The fade thesis needs a dark lower-right corner — bad wallets staying bad — while high exit rates in D10 flag the survivorship control this audit must include.
Data table
| From \ To | D1 | D2 | D3 | D4 | D5 | D6 | D7 | D8 | D9 | D10 | Exit |
|---|---|---|---|---|---|---|---|---|---|---|---|
| D1 | 39.2 | 18.4 | 12.5 | 8.5 | 5.8 | 4.0 | 2.7 | 1.8 | 1.2 | 0.8 | 5.0 |
| D2 | 15.6 | 32.0 | 15.6 | 10.6 | 7.2 | 4.9 | 3.4 | 2.3 | 1.6 | 1.1 | 5.7 |
| D3 | 9.6 | 14.1 | 28.0 | 14.1 | 9.6 | 6.6 | 4.5 | 3.0 | 2.1 | 1.4 | 7.0 |
| D4 | 6.1 | 9.0 | 13.2 | 25.6 | 13.2 | 9.0 | 6.1 | 4.2 | 2.8 | 1.9 | 8.8 |
| D5 | 4.0 | 5.9 | 8.6 | 12.6 | 24.2 | 12.6 | 8.6 | 5.9 | 4.0 | 2.7 | 11.0 |
| D6 | 2.6 | 3.9 | 5.7 | 8.3 | 12.3 | 23.5 | 12.3 | 8.3 | 5.7 | 3.9 | 13.6 |
| D7 | 1.8 | 2.6 | 3.8 | 5.6 | 8.2 | 12.1 | 23.4 | 12.1 | 8.2 | 5.6 | 16.5 |
| D8 | 1.2 | 1.8 | 2.6 | 3.8 | 5.7 | 8.3 | 12.2 | 24.1 | 12.2 | 8.3 | 19.7 |
| D9 | 0.9 | 1.3 | 1.9 | 2.7 | 4.0 | 5.9 | 8.7 | 12.7 | 26.0 | 12.7 | 23.2 |
| D10 | 0.7 | 1.0 | 1.4 | 2.1 | 3.0 | 4.5 | 6.6 | 9.6 | 14.1 | 30.1 | 27.0 |
5. Tests
Pre-registered statistical tests
The central test is a decile sort (Figure 3). Freeze the score before the validation window, sort wallets into deciles, and measure the next-window net fade edge. The expected result is a monotonic relation: higher Fool Score implies lower original-direction markout and a larger net fade edge. The four pre-registered hypotheses, each with a numeric pass criterion, are collected in Figure 4.
Figure 3
Pre-registered decile test
Illustrative target chart for the live audit. Bars grow from a zero baseline; whiskers are the 95% interval that must be reported per decile. The required result is a positive, monotonic slope after costs — not merely a positive top decile.
Net fade edge (bps) by pre-trade Fool Score decile
Data table
| Decile | Net fade edge (bps) | CI low | CI high |
|---|---|---|---|
| D1 | -4 | -13 | 5 |
| D2 | -2 | -10 | 6 |
| D3 | +1 | -7 | 9 |
| D4 | +3 | -5 | 11 |
| D5 | +5 | -3 | 13 |
| D6 | +7 | -1 | 15 |
| D7 | +10 | 2 | 18 |
| D8 | +14 | 5 | 23 |
| D9 | +19 | 9 | 29 |
| D10 | +27 | 15 | 39 |
Figure 4
Hypotheses and pass criteria
Each hypothesis carries a numeric pass criterion fixed before the validation window opens. A criterion that is not met falsifies the corresponding claim — there is no post-hoc adjustment.
H1
Top-score wallets have lower next-window original-direction markout than matched controls.
PASS · One-sided p < 0.05 with wallet-clustered standard errors, against the matched-control mean.
H2
The net fade edge remains positive after fees, spread, slippage, funding and reject cost.
PASS · Block-bootstrap 95% CI of the top-decile mean net fade edge lies strictly above zero.
H3
The decile slope is monotonic: higher Fool Score implies a larger net fade edge.
PASS · Spearman ρ(decile, net fade edge) ≥ 0.8 at p < 0.05; Jonckheere-Terpstra as robustness.
H4
The result survives walk-forward validation and Benjamini-Hochberg FDR control.
PASS · Positive top-decile edge in ≥ 2 of 3 windows; every surviving test passes BH at q = 0.10.
6. Execution
From signal to realized PnL
Signal quality is not sufficient. The edge must survive exchange fees, builder fee, bid-ask spread, slippage, funding and rejected-order opportunity cost. Hyperliquid's public API surfaces the relevant evidence: user fills, fills by time, order status, L2 book snapshots and builder-fee fields [7].Figure 5 makes the requirement explicit as a bridge from gross markout to realized PnL, with a stress control on every execution cost.
Figure 5
Cost attribution bridge
Illustrative bridge from gross inverse markout to realized PnL. Drag the stress slider to scale every execution cost; the audit must replace these values with fill-level evidence.
Data table
| Step | Contribution (bps) | Running total (bps) |
|---|---|---|
| Gross inverse markout | +37.0 | 37.0 |
| Exchange and builder fees | -4.0 | 33.0 |
| Bid-ask spread and slippage | -7.0 | 26.0 |
| Funding paid while holding | -2.0 | 24.0 |
| Reject and delay opportunity cost | -3.0 | 21.0 |
| Net fade edge | +21.0 | 21.0 |
The audited deliverable is the walk-forward equity curve in Figure 6: each window re-freezes the score on trailing data only, and the fade portfolio is compared against a matched baseline, with the drawdown path reported alongside the average.
Figure 6
Walk-forward equity curve
Simulated illustration of the walk-forward audit output — not measured performance. Each window re-freezes the score on trailing data only; the published version must show this chart with real fills, the drawdown path and the matched baseline.
Cumulative net return (%) over 26 validation weeks
Data table
| Week | Fade top decile (%) | Matched baseline (%) |
|---|---|---|
| 0 | 0.0 | 0.0 |
| 1 | +0.4 | +0.2 |
| 2 | +0.9 | -0.1 |
| 3 | +0.6 | +0.3 |
| 4 | +1.3 | +0.1 |
| 5 | +1.8 | -0.2 |
| 6 | +2.4 | +0.2 |
| 7 | +2.1 | +0.5 |
| 8 | +2.9 | +0.3 |
| 9 | +3.4 | 0.0 |
| 10 | +3.0 | -0.3 |
| 11 | +2.2 | +0.1 |
| 12 | +1.6 | +0.4 |
| 13 | +2.0 | +0.2 |
| 14 | +2.8 | -0.1 |
| 15 | +3.5 | +0.3 |
| 16 | +4.1 | +0.6 |
| 17 | +4.8 | +0.4 |
| 18 | +4.4 | +0.1 |
| 19 | +5.2 | +0.5 |
| 20 | +5.9 | +0.3 |
| 21 | +6.3 | 0.0 |
| 22 | +5.8 | +0.4 |
| 23 | +6.6 | +0.7 |
| 24 | +7.2 | +0.5 |
| 25 | +7.8 | +0.2 |
| 26 | +8.2 | +0.4 |
A wallet entering watching is not a failed trade. It means the target wallet has no open exposure at that moment. The strategy has joined the portfolio and waits for a future tradable event. If the target already has exposure, execution can move to the queue immediately (Figure 7).
Figure 7
Validation pipeline
Watching is not a failed order: it is the event-triggered branch taken when the target is flat. Hover a branch to trace it.
7. Validity
Threats to validity
A pre-analysis plan is only credible if it names the ways it can be wrong before a critic does. Five threats dominate here; each carries a mitigation that is part of the design, not an afterthought.
T1
Reflexivity and crowding
If many accounts fade the same wallet, the inverse flow itself moves price against the fade, and an edge measured at audit scale decays at product scale. Mitigation: per-wallet notional caps, participation capped as a share of coin-level volume, and cohort-level edge-decay monitoring after launch.
T2
Selection at discovery
The screening floors in Table 2 condition on activity and account value, so the scored universe is not the wallet population. Mitigation: publish the funnel count at every stage and restrict every claim to the screened cohort.
T3
Survivorship and attrition
The worst wallets blow up or leave — the exit column of Figure 2 — and conditioning on survival inflates measured persistence. Mitigation: episode-level accounting keeps exiting wallets in the denominator until exit, and transition matrices are always reported exit-inclusive.
T4
Non-stationarity
Behavioral intensity varies across market regimes, as the Bitcoin disposition-effect evidence shows [6]. A score calibrated in one regime may decay in the next. Mitigation: walk-forward re-freezing every window, regime-split reporting, and monitored live decay.
T5
Execution adverse selection
The strategy observes target fills with delay and trades after them, so its entries are mechanically worse than the target's own. Mitigation: the cost model in Figure 5 is estimated from the strategy's own fills, never from quoted spreads.
8. Falsification
What would disprove the thesis
The thesis is pre-registered to be rejected, not defended. Each hypothesis in Figure 4 carries a numeric criterion, and any of the following outcomes falsifies the thesis or marks it unproven: top-decile wallets fail to underperform matched controls at one-sided p < 0.05 with wallet-clustered standard errors; the block-bootstrap 95% confidence interval of the top-decile mean net fade edge includes zero after all costs; the decile slope loses monotonicity, with Spearman ρ below 0.8 or insignificant; no test survives Benjamini-Hochberg control at q = 0.10; or fewer than two of the three walk-forward windows deliver a positive top-decile edge.
Minimum evidence for publication
- Scoring rule frozen and hash-stamped before the validation window opens.
- At least 1,024 scored wallets, each with ≥ 20 closed episodes per window.
- Four matched control wallets per target, plus a no-trade cash baseline.
- Net PnL after fees, spread, slippage, funding and rejection cost, measured from own fills.
- 95% confidence intervals, wallet-clustered t-statistics, Benjamini-Hochberg q ≤ 0.10.
- Failure cases, a capacity estimate, and the full drawdown distribution — not just the mean.
References