The Oil Anomaly Is a Calendar

TL;DR The strangest line in this website's machine-learning results has been sitting in the claims ledger for a month, filed under "standing anomalies": crude oil, which is not on the dashboard, outranked most indicators that are. Top-5 SHAP importance in 7 of 12 walk-forward rounds, ranked first twice, and present in 3 of the 5 healthiest models, all at the 30-day horizon. Four tests now resolve the anomaly. Oil's rank is real, and it is a calendar, not a signal. Pooled across 11 years, days with cheap oil saw bitcoin rise over the next 30 days 71.4% of the time, against roughly 51% otherwise: a 19.5-point edge, easily the kind of split a gradient-boosted tree loves. Control for the calendar year and the edge collapses to 0.0 points. Cheap-oil eras (2015-16, 2020-21) simply were bitcoin bull markets; within any year, oil says nothing about the next 30 days, and its flows carry no lead (r² under 1%). The model learned when each row happened, not what happens next. Keeping oil off the dashboard was right, and every SHAP ranking, here and elsewhere, now carries the lesson.

Since June, this website's claims ledger has carried a section most research sites would not publish, titled "standing anomalies (named, not hidden)." Its first entry was an embarrassment on purpose: crude oil, a feature included in the training data as macro context and never shown on the dashboard, ranked in the SHAP top-5 (SHAP works like an itemized receipt for a model's forecast: it scores how much each input moved the output) in 7 of the 12 walk-forward model runs (walk-forward meaning: train on the past, test on a future the model never saw, slide the window forward, repeat) and finished first in two of them. That placed it above most of the thirteen indicators the dashboard does use. The entry ended with a confession: the working hypothesis was "regime proxy," but no test distinguished that from noise, so exclusion was a judgment call, disclosed as such.

This article runs the test. It does three things: establishes that the anomaly is real rather than an artifact of the degenerate models this website has already disclosed, runs oil through four falsification tests (the same treatment that rejected the stablecoin indicator), and draws the conclusion that matters beyond oil: a feature can earn a high SHAP rank by telling the model what year it is. A calendar, not a signal.

The Anomaly Is Real

The lazy resolution was available and would have been wrong. Seven of the twelve XGBoost models collapsed to one or two trees under early stopping (disclosed in the SHAP analysis), and rankings from degenerate models are weak evidence. XGBoost builds chains of yes/no questions about the data, each chain patching the errors of the last: each tree is one such chain, and a healthy model stacks dozens. Early stopping is the training rule that halts a model the moment extra complexity stops paying, and a model that stops after one or two trees found almost nothing durable to learn: its opinions are closer to a shrug than a ranking. If oil's placements all came from those, the anomaly would dissolve into the existing caveat.

They do not. Of the five healthy models (ten or more trees, enough to have actually built a view of the data), oil is top-5 in three, and the three share a property: all are 30-day-horizon models. At that horizon oil appears in the top-5 in three of four rounds, spanning test windows from 2019 through 2026, ranked first in one of them (a 13-tree model, not a degenerate one). Whatever oil was doing, the healthiest models at the longest horizon kept doing it, across three different market cycles. That is a pattern demanding an explanation, not a shrug.

Two explanations were on the table. Either oil genuinely carries 30-day-ahead information about bitcoin (which would mean the dashboard is missing an indicator), or oil is a regime marker: a slow-moving level that tells the model which macro era a training row belongs to, letting the trees memorize era-specific base rates that do not transfer forward. The two hypotheses make different testable predictions, which is what makes this resolvable rather than a matter of taste.

Four Tests

The first test, trend contamination, clears oil rather than convicting it. Bitcoin against WTI at the level correlates at just +0.32 (+0.44 in logs, excluding the negative print of April 20, 2020). Unlike the stablecoin case, where a 0.84 correlation was two uptrends holding hands, oil mean-reverts (it gets pulled back toward its own average) while bitcoin trends. The SHAP rank is not a shared-trend artifact, which makes it more curious, not less.

The second test, flows, finds nothing. Across 618 weeks of log-returns, the strongest oil-leads-bitcoin correlation at any horizon out to 12 weeks is +0.095, an r² of 0.9 percent, and the placebo direction (bitcoin leading oil) produces values of the same size. If oil moves and bitcoin follows, 618 weeks of data cannot see it.

The third test, valuation-style, is the same oscillator treatment the stablecoin indicator failed: oil z-scored against its trailing year (rescaled to how unusual today's price is against its own past year), quintiled (every day sorted into five buckets from cheapest to dearest), checked against forward bitcoin returns. Pooled correlation at 30 days: −0.001. The quintile pattern is non-monotonic with its peak in the fourth quintile: if oil mattered, forward returns should step up or down steadily across the buckets; instead the pattern jumps around. Noise-shaped, again.

The fourth test is the one built to separate the two hypotheses, and it produces the article's headline numbers. Take every day in the sample and ask a simple question: over the next 30 days (the horizon where oil's SHAP rank concentrates), did bitcoin go up? Pooled across the full 2015-2026 sample, days with oil in the bottom third of its range answered yes 71.4 percent of the time. Days with oil in the middle or top third: roughly 51 percent, a coin flip. A 19.5-point spread on a binary outcome is enormous; a decision tree that splits on "oil below about $50" buys itself a massive, real, in-sample edge with a single threshold. This is, almost certainly, the split the models found.

Then run the identical comparison within each calendar year, so the model can no longer use oil to tell 2016 from 2022, and the edge does not shrink. It vanishes: the mean within-year spread across 11 years is 0.0 points. The 71.4 percent was never oil predicting bitcoin. Cheap-oil eras, 2015-16 after the shale glut and 2020-21 after the pandemic crash, happened to be the two great bitcoin bull markets, and expensive-oil eras (2022's inflation shock) happened to contain bitcoin's worst bear. Oil's level is a timestamp wearing a price. The trees read the timestamp, memorized each era's base rate, and SHAP faithfully reported that the timestamp was useful. A calendar, not a signal.

What This Buys Every Other Ranking

The immediate housekeeping is pleasant: the dashboard's exclusion of oil, previously a disclosed judgment call, is now a tested decision. The anomaly section of the claims ledger gets its first resolution, and the answer required no retraction because the uncertainty was stated at the time. This is what the ledger is for.

The transferable lesson is worth more than the housekeeping. This website's SHAP rankings already carried one caveat: most of the models are too small to trust individually. They now carry a second, sharper one: a feature can earn its rank by dating the row rather than predicting the outcome. Any slow-moving macro level in a training set that spans distinct market eras (oil, rates, M2, the dollar) is a candidate calendar, and a walk-forward SHAP table cannot tell you by itself which kind of importance it is reporting. Distinguishing them takes exactly the test run here: pool, then condition on time, and watch whether the edge survives. Readers who see feature-importance charts elsewhere in crypto research, where "macro feature X drives bitcoin" is a genre, may find that the question "does it survive year fixed effects?" is rarely asked and rarely survivable. Update: the question has now been asked of every feature in this pipeline, including the dashboard's own indicators; see The Dashboard Takes the Calendar Test.

The caveat also points inward, and honesty requires saying so. The dashboard's own weights are informed by these same rankings, and this website has already published that they add no timing edge over naive alternatives (they are an aggressiveness dial; see the backtest article). The oil result is consistent with that finding and explains part of its mechanism: some of what walk-forward SHAP rewarded, in every feature, was era-memorization that does not transfer. The defensible claims remain cluster-level and are documented as such.

What the reader takes away: one number, one test, one habit. The number is 71.4 percent pooled collapsing to a 0.0-point edge within-year, the cleanest example this website has produced of an in-sample pattern that is real, large, and worthless. The test is ml/research/oil_test.py, one command against the public dataset, reproducing every figure here. The habit is to ask, of any feature-importance claim, whether the feature predicts the outcome or merely dates the observation. The claims ledger has one fewer anomaly tonight, and the reason it could be resolved cleanly is that it was written down when it was still embarrassing.