Can Weak Wallets Generate Alpha?

A pre-analysis research note on whether a public derivatives venue can expose a persistent negative-alpha cohort, and whether a cost-aware inverse strategy can harvest it.

Version

v0.3, July 2026

Status

Thesis and test design

Claim

Not yet a live-profit audit

Abstract

Main claim

The claim is not that every losing wallet should be copied in reverse. The defensible claim is narrower: if a wallet can be classified ex ante as a member of a persistent negative-alpha cohort, then the inverse of its future trades may have positive expected value after execution costs.

This hypothesis is grounded in a large behavioral-finance literature. Active retail trading has been shown to reduce net returns [1], investors display a tendency to realize winners and retain losers[2], aggregate individual trading losses can be economically large [3], and day-trader skill is highly dispersed with a reliably weak lower tail [4]. The open question is whether similar weak-tail behavior is measurable and tradeable in Hyperliquid wallets.

ri,e  =  si,ePoutPinPin(1)r_{i,e} \;=\; s_{i,e}\,\frac{P_{\mathrm{out}} - P_{\mathrm{in}}}{P_{\mathrm{in}}} \tag{1}
m~i,e,t  =  si,eme,t(2)\widetilde{m}_{i,e,t} \;=\; -\,s_{i,e}\, m_{e,t} \tag{2}
πi,e,t  =  m~i,e,t    ce(3)\pi_{i,e,t} \;=\; \widetilde{m}_{i,e,t} \;-\; c_{e} \tag{3}
αt  =  E ⁣[πi,e,t  |  F(i)D10]    E ⁣[πe,tbase](4)\alpha_t \;=\; \mathbb{E}\!\left[\,\pi_{i,e,t} \;\middle|\; F(i) \in \mathrm{D}_{10}\,\right] \;-\; \mathbb{E}\!\left[\,\pi^{\mathrm{base}}_{e,t}\,\right] \tag{4}

Equation (4) is the estimand: the expected net fade edge of the top Fool Score decile in excess of a matched baseline, measured out of sample.

1. Design

Estimand and notation

The estimand is the out-of-sample net fade edge of wallets selected by a pre-trade score: the inverse markout net of all execution costs, written π in the equation block. This note uses that one term throughout. The unit of observation is a closed episode, not an account screenshot. An episode starts when net exposure leaves zero and ends when exposure returns to zero or flips.

Table 1 · Notation

ii

Wallet index.

ee

Closed trading episode: position leaves zero, then returns to zero or flips.

si,e{+1,1}s_{i,e} \in \{+1, -1\}

Target direction of the episode. Long = +1, short = -1.

me,tm_{e,t}

Future markout over horizon t after the target trade.

m~i,e,t\widetilde{m}_{i,e,t}

Inverse markout: the markout earned by taking the opposite side.

cec_{e}

All-in cost: fees, spread, slippage, funding and execution failure cost.

F(i)F(i)

Fool Score computed strictly before the validation window.

πi,e,t\pi_{i,e,t}

Net fade edge: inverse markout net of all execution costs.

A positive research result requires that top-score wallets have worse future original-direction markout than matched controls, and that the inverse strategy remains positive after all costs. This separates behavioral selection from execution quality. The horizon parameter t is not free: Figure 1 sketches the target term structure, in which gross inverse markout must clear the all-in cost stack before some horizon for a tradable window to exist at all.

Figure 1

Inverse markout by holding horizon

Illustrative target term structure. Gross inverse markout must clear the all-in cost stack before some horizon for the strategy to have a tradable window; here the modeled breakeven sits near one hour.

Inverse markout (bps) vs horizon after target entry

net of costsgross
-20-100102030405m30m1h4h1d3dbreakeven ≈ 1hnet +12gross +28
Data table
HorizonNet of costs (bps)CI ±Gross (bps)
5m-104+6
30m-45+12
1h06+16
4h+88+24
1d+1511+31
3d+1214+28

2. Data

Sample construction

The sample is built by a two-stage public funnel. A daily discovery pass screens the venue's leaderboard for wallets that are active enough to score and not structurally unfadeable; survivors enter a scoring queue, where each wallet's trailing 90 days of fills are reconstructed into closed episodes and scored. The top 1,024 wallets by Fool Score form the deck — simultaneously the tradable cohort and the pre-registered test cohort. Every input is a public API endpoint [7]: fills by time, account snapshots and L2 book. Anyone can rebuild the sample.

Table 2 · Sampling and cohort parameters, pre-registered (v0.3)

ParameterValuePurpose
UniverseHyperliquid public leaderboard, daily snapshotEvery address visible on the venue's leaderboard each day; no privileged data.
Day volume floor≥ $10,000 traded on the snapshot dayData sufficiency: enough fills for the window to be scoreable.
Account value floor≥ $1,000Excludes dust and abandoned wallets, which cannot be faded at size.
Turnover ceilingmonthly volume < 100× account valueExcludes HFT and market-making flow, whose fills are not behavioral.
Scoring windowtrailing 90 days of fillsLong enough to accumulate episodes; short enough to track the current regime.
Episode floor≥ 20 closed episodes per wallet-windowStatistical floor below which a wallet-level estimate is noise.
Deck sizetop 1,024 wallets by Fool ScoreThe tradable cohort, and therefore the pre-registered test cohort.
Matched controls4 per target walletNearest neighbours on coin universe, turnover, account value and liquidity bucket.
Validation windownext 90 days, zero overlapWalk-forward: the score is re-frozen at each window boundary.

The screening floors and ceiling are not neutral: they define the population the thesis speaks about. The volume and account-value floors remove wallets that cannot be scored or faded at size; the turnover ceiling removes high-frequency and market-making flow, whose fills reflect inventory management rather than belief. All claims in this note are therefore claims about the screened cohort, and the published audit must report the funnel count at every stage — wallets seen, wallets screened, wallets scored, wallets decked — so that selection is visible rather than silent.

Because episodes within one wallet are correlated, wallets — not episodes — are the clustering unit for every standard error in Section 5. Wallets that go inactive during a validation window remain in the denominator until their exit is observed, which is what the exit column of Figure 2 makes explicit.

3. Score

The Fool Score

The score is deliberately boring. Each wallet-window is measured on five factor families, each anchored in a documented behavioral regularity; family measurements are winsorized, z-scored against the screened cohort, and combined with equal weights into a single monotone score. Higher means more exploitable. There is no fitted model and no interaction terms in v0.3 — a score this simple is harder to overfit and easier to falsify.

Table 3 · Fool Score factor families

FamilyBehavioral signaturePrimary measurementAnchor
DispositionRealizes winners quickly, rides losersPGR − PLR: proportion of gains realized minus proportion of losses realized per windowOdean (1998) [2]
OvertradingTurnover far above what the account justifiesMonthly turnover; trade count per unit of account valueBarber & Odean (2000) [1]
PersistenceLosses repeat across windows instead of mean-revertingTrailing original-direction markout in prior windows; decile stabilityBarber, Lee, Liu & Odean (2014) [4]
TimingSystematically enters just before adverse movesShort-horizon (5m–4h) original-direction markout after entryBarber, Lee, Liu & Odean (2009) [3]
SizingEscalates size and leverage while losingSize-after-loss ratio; leverage trajectory inside drawdownsBarber & Odean (2001) [5]

Two discipline rules protect the score from its own authors. First, the recipe is frozen and hash-stamped before each validation window opens; the walk-forward test in Figure 6 is only meaningful if the score cannot move after the fact. Second, any change to a feature, weight or winsorization bound re-registers the study and restarts the validation clock. The features themselves are chosen to be expensive to fake: the only way for a wallet to stop looking foolish on these measures is to stop trading foolishly.

4. Literature

Why this is plausible

Barber and Odean document that households with the highest turnover materially underperform after costs[1]. Their overconfidence study also links higher trading activity to lower net returns[5]. Odean's disposition-effect paper shows that investors are more willing to realize gains than losses, and that the losers they keep do not subsequently justify that choice [2].

The most relevant empirical bridge is the day-trading skill paper. Barber, Lee, Liu and Odean sort traders by past performance and test the next period; the bottom-ranked tail continues to lose after fees, while only a very small top tail earns reliable positive abnormal returns [4]. This is exactly the shape a fade engine needs: avoid the skilled tail and target the persistently weak tail.

Crypto markets do not remove behavioral bias. Schatzmann and Haslhofer find evidence of the disposition effect in Bitcoin transaction data, with varying intensity across regimes [6]. That does not prove whosthefool's edge, but it supports the premise that on-chain markets can contain measurable behavioral regularities.

What the literature predicts, concretely, is a sticky transition matrix. If skill and its absence persist, a wallet's score decile in one window should forecast its decile in the next; Figure 2 shows the target shape, including the exit column that makes survivorship an explicit object of the audit rather than a silent bias.

Figure 2

Decile transition matrix

Illustrative target shape for score persistence: rows are Fool Score deciles in one window, columns the decile occupied in the next. The fade thesis needs a dark lower-right corner — bad wallets staying bad — while high exit rates in D10 flag the survivorship control this audit must include.

decile in window t+1 →← decile in window tD1D2D3D4D5D6D7D8D9D10ExitD1391813D2163216D3142814D4132613D5132413D612231214D712231216D812241220D913261323D10143027
Data table
From \ ToD1D2D3D4D5D6D7D8D9D10Exit
D139.218.412.58.55.84.02.71.81.20.85.0
D215.632.015.610.67.24.93.42.31.61.15.7
D39.614.128.014.19.66.64.53.02.11.47.0
D46.19.013.225.613.29.06.14.22.81.98.8
D54.05.98.612.624.212.68.65.94.02.711.0
D62.63.95.78.312.323.512.38.35.73.913.6
D71.82.63.85.68.212.123.412.18.25.616.5
D81.21.82.63.85.78.312.224.112.28.319.7
D90.91.31.92.74.05.98.712.726.012.723.2
D100.71.01.42.13.04.56.69.614.130.127.0

5. Tests

Pre-registered statistical tests

The central test is a decile sort (Figure 3). Freeze the score before the validation window, sort wallets into deciles, and measure the next-window net fade edge. The expected result is a monotonic relation: higher Fool Score implies lower original-direction markout and a larger net fade edge. The four pre-registered hypotheses, each with a numeric pass criterion, are collected in Figure 4.

Figure 3

Pre-registered decile test

Illustrative target chart for the live audit. Bars grow from a zero baseline; whiskers are the 95% interval that must be reported per decile. The required result is a positive, monotonic slope after costs — not merely a positive top decile.

Net fade edge (bps) by pre-trade Fool Score decile

positivenegative
-20-10010203040D1-4D2D3D4D5D6D7D8D9D10+27
Data table
DecileNet fade edge (bps)CI lowCI high
D1-4-135
D2-2-106
D3+1-79
D4+3-511
D5+5-313
D6+7-115
D7+10218
D8+14523
D9+19929
D10+271539

Figure 4

Hypotheses and pass criteria

Each hypothesis carries a numeric pass criterion fixed before the validation window opens. A criterion that is not met falsifies the corresponding claim — there is no post-hoc adjustment.

H1

Top-score wallets have lower next-window original-direction markout than matched controls.

PASS · One-sided p < 0.05 with wallet-clustered standard errors, against the matched-control mean.

H2

The net fade edge remains positive after fees, spread, slippage, funding and reject cost.

PASS · Block-bootstrap 95% CI of the top-decile mean net fade edge lies strictly above zero.

H3

The decile slope is monotonic: higher Fool Score implies a larger net fade edge.

PASS · Spearman ρ(decile, net fade edge) ≥ 0.8 at p < 0.05; Jonckheere-Terpstra as robustness.

H4

The result survives walk-forward validation and Benjamini-Hochberg FDR control.

PASS · Positive top-decile edge in ≥ 2 of 3 windows; every surviving test passes BH at q = 0.10.

6. Execution

From signal to realized PnL

Signal quality is not sufficient. The edge must survive exchange fees, builder fee, bid-ask spread, slippage, funding and rejected-order opportunity cost. Hyperliquid's public API surfaces the relevant evidence: user fills, fills by time, order status, L2 book snapshots and builder-fee fields [7].Figure 5 makes the requirement explicit as a bridge from gross markout to realized PnL, with a stress control on every execution cost.

Figure 5

Cost attribution bridge

Illustrative bridge from gross inverse markout to realized PnL. Drag the stress slider to scale every execution cost; the audit must replace these values with fill-level evidence.

net edge +21.0 bps · breakeven ≈ 2.3× costs
-10010203040+37Gross-4Fees-7Spread-2Funding-3Reject+21Net
Data table
StepContribution (bps)Running total (bps)
Gross inverse markout+37.037.0
Exchange and builder fees-4.033.0
Bid-ask spread and slippage-7.026.0
Funding paid while holding-2.024.0
Reject and delay opportunity cost-3.021.0
Net fade edge+21.021.0

The audited deliverable is the walk-forward equity curve in Figure 6: each window re-freezes the score on trailing data only, and the fade portfolio is compared against a matched baseline, with the drawdown path reported alongside the average.

Figure 6

Walk-forward equity curve

Simulated illustration of the walk-forward audit output — not measured performance. Each window re-freezes the score on trailing data only; the published version must show this chart with real fills, the drawdown path and the matched baseline.

Cumulative net return (%) over 26 validation weeks

fade top decilematched baseline
W1W2W3-2024680510152025fade +8.2%baseline +0.4%−1.8 pp drawdown
Data table
WeekFade top decile (%)Matched baseline (%)
00.00.0
1+0.4+0.2
2+0.9-0.1
3+0.6+0.3
4+1.3+0.1
5+1.8-0.2
6+2.4+0.2
7+2.1+0.5
8+2.9+0.3
9+3.40.0
10+3.0-0.3
11+2.2+0.1
12+1.6+0.4
13+2.0+0.2
14+2.8-0.1
15+3.5+0.3
16+4.1+0.6
17+4.8+0.4
18+4.4+0.1
19+5.2+0.5
20+5.9+0.3
21+6.30.0
22+5.8+0.4
23+6.6+0.7
24+7.2+0.5
25+7.8+0.2
26+8.2+0.4

A wallet entering watching is not a failed trade. It means the target wallet has no open exposure at that moment. The strategy has joined the portfolio and waits for a future tradable event. If the target already has exposure, execution can move to the queue immediately (Figure 7).

Figure 7

Validation pipeline

Watching is not a failed order: it is the event-triggered branch taken when the target is flat. Hover a branch to trace it.

position open → pending_executionflat targettarget opens → queueWatchingarmed · event-triggered01Freeze scoreF(i) before window02Match controlsliquidity · turnover03Observe targettradable event?04Attribute PnLfills · book · funding

7. Validity

Threats to validity

A pre-analysis plan is only credible if it names the ways it can be wrong before a critic does. Five threats dominate here; each carries a mitigation that is part of the design, not an afterthought.

T1

Reflexivity and crowding

If many accounts fade the same wallet, the inverse flow itself moves price against the fade, and an edge measured at audit scale decays at product scale. Mitigation: per-wallet notional caps, participation capped as a share of coin-level volume, and cohort-level edge-decay monitoring after launch.

T2

Selection at discovery

The screening floors in Table 2 condition on activity and account value, so the scored universe is not the wallet population. Mitigation: publish the funnel count at every stage and restrict every claim to the screened cohort.

T3

Survivorship and attrition

The worst wallets blow up or leave — the exit column of Figure 2 — and conditioning on survival inflates measured persistence. Mitigation: episode-level accounting keeps exiting wallets in the denominator until exit, and transition matrices are always reported exit-inclusive.

T4

Non-stationarity

Behavioral intensity varies across market regimes, as the Bitcoin disposition-effect evidence shows [6]. A score calibrated in one regime may decay in the next. Mitigation: walk-forward re-freezing every window, regime-split reporting, and monitored live decay.

T5

Execution adverse selection

The strategy observes target fills with delay and trades after them, so its entries are mechanically worse than the target's own. Mitigation: the cost model in Figure 5 is estimated from the strategy's own fills, never from quoted spreads.

8. Falsification

What would disprove the thesis

The thesis is pre-registered to be rejected, not defended. Each hypothesis in Figure 4 carries a numeric criterion, and any of the following outcomes falsifies the thesis or marks it unproven: top-decile wallets fail to underperform matched controls at one-sided p < 0.05 with wallet-clustered standard errors; the block-bootstrap 95% confidence interval of the top-decile mean net fade edge includes zero after all costs; the decile slope loses monotonicity, with Spearman ρ below 0.8 or insignificant; no test survives Benjamini-Hochberg control at q = 0.10; or fewer than two of the three walk-forward windows deliver a positive top-decile edge.

Minimum evidence for publication

  • Scoring rule frozen and hash-stamped before the validation window opens.
  • At least 1,024 scored wallets, each with ≥ 20 closed episodes per window.
  • Four matched control wallets per target, plus a no-trade cash baseline.
  • Net PnL after fees, spread, slippage, funding and rejection cost, measured from own fills.
  • 95% confidence intervals, wallet-clustered t-statistics, Benjamini-Hochberg q ≤ 0.10.
  • Failure cases, a capacity estimate, and the full drawdown distribution — not just the mean.