A precise probability can still be fragile. A signal can be directionally correct and economically useless. It can have a strong historical average supported by too few independent cases. Good evaluation keeps those possibilities visible.
The anatomy of a useful crypto signal
A research signal should identify the asset, direction or risk state, horizon, timestamp, setup, eligible market regime, evidence, counter-evidence, invalidation condition, and expiry. If any of those are missing, the reader has to invent part of the claim.
The timestamp matters because market evidence changes. The horizon matters because a four-hour setup and a three-day setup face different noise, costs, and overlap. The invalidation condition matters because a signal that can never admit failure cannot be evaluated.
Read the base rate before the signal rate
The base rate is the frequency of the outcome among comparable eligible periods before the proposed setup adds information. If favorable outcomes occur 52% of the time normally and 56% after the setup, the measured lift is four percentage points, not 56 points.
lift = setup outcome rate - eligible-market base rate
Choose the comparison carefully. A broad all-market base rate can hide meaningful differences between upward trends, crisis states, high volatility, and quiet ranges. But slicing the data too finely creates tiny samples and unstable estimates.
Ask how many independent cases support the number
Always show the sample count beside a rate or probability. Then ask how much overlap exists. Twenty-four hourly observations of a 24-hour outcome are not twenty-four independent cases because most of their future windows share the same price path.
An effective sample size can be much smaller than the row count. Dependence-aware intervals, temporal folds, and block resampling help reveal that loss of information. When the sample is too thin, show the count and withhold the rate.
Displaying 0.623 instead of 0.62 does not create more evidence. Precision should reflect what the data can support, not what a formatter can print.
Calibration asks whether probabilities mean what they say
Accuracy asks whether a direction was correct. Calibration asks whether events assigned a probability near 60% actually occur about 60% of the time across comparable predictions. A model can rank cases reasonably while producing probabilities that are too confident.
Inspect calibration by probability range, not only as one average. Show an uncertainty interval around the estimated probability. A wide interval can make the best action "watch" even when the midpoint appears attractive.
Translate historical tendency into net utility
A favorable outcome must exceed the friction required to observe or act on it. At minimum, consider fees, spread, slippage, latency, funding, gas, and fill failure where relevant.
net outcome = gross outcome - fees - spread - slippage - funding - other friction
Use costs appropriate to the venue and horizon. Shorter horizons often have less movement available to absorb friction. Stress the estimate under higher costs and delayed decisions. A signal that disappears under a modestly worse profile should be labelled fragile.
Put counter-evidence beside supporting evidence
Evidence should come from meaningfully different sources rather than repeated versions of the same indicator. Price trend, derivatives positioning, market breadth, liquidity, and event risk can disagree. That disagreement is information.
A useful signal states what supports it, what argues against it, and which observation would invalidate it. Do not bury the opposing evidence in a separate screen after the confident headline.
Treat abstention as a valid result
A research system should be allowed to say that no governed signal exists. Missing data, stale inputs, an ineligible regime, weak confirmation, poor economics, or a wide uncertainty interval can all justify staying quiet.
Preserve the candidate and the first reason it stopped. That creates a suppression record you can evaluate later. Erasing near misses makes the system look cleaner while destroying evidence about its actual selectivity.
Crypto signal evaluation worksheet
| Question | What to look for |
|---|---|
| What exactly is the claim? | Asset, direction or risk state, horizon, timestamp, setup, and eligible regime |
| What is the denominator? | Resolved sample count, effective sample size, and missing or unfillable cases |
| What is normal? | A relevant eligible-market base rate or simple baseline |
| How uncertain is it? | Confidence interval, calibration evidence, and sensitivity across time |
| Does it survive friction? | Net outcome after venue- and horizon-relevant costs |
| What argues against it? | Independent counter-evidence, disagreement, and data-quality limits |
| What ends the claim? | Price, time, regime, freshness, or evidence invalidation condition |
| Can it stay quiet? | An explicit watch, nearest-miss, insufficient-evidence, or abstain state |
Continue the research path
See how these ideas are applied, including negative benchmark results, in the Crypto Signal Lab public methodology.