Did anyone buy the return months? CL futures around the 2026 strategic-reserve loans¶
soroban, 2026-10-05. Data caches: build_spr_loan_caches.py. Every number below is rebuilt from the exchange's own data and the Energy Department's own postings.
Why the question matters. Between March and June 2026 the US Energy Department lent 133.6 million barrels of crude from the Strategic Petroleum Reserve to oil companies and trading houses. These were loans, which the Department calls exchanges: a company takes crude now and must give back the same amount plus a premium of 8 to 24% more barrels, on a schedule that runs from late 2026 into 2029. A company that sells the borrowed crude and owes crude later carries a price risk on the barrels it owes. The standard way to remove that risk is to buy crude futures for the months in which the crude must be returned. If the borrowers did that, the CL futures for the return months should show unusual buying close to the dates on which the loans were agreed. This notebook looks for that buying.
How. Each loan round has a published bid deadline and award date. For every CL contract I measure five things over the ten trade dates ending three trade dates after each award: how much the number of open contracts grew, how that growth compared with the neighbouring delivery months, how much more was traded than usual, how much of the trading happened away from the exchange's screen, and how the price moved against the neighbouring months. Each measurement is judged against what CL contracts at the same distance from delivery did over the ten years 2016 to 2025, by two machine-learning anomaly detectors (an autoencoder and an isolation forest) and one plain statistical score. The direction of screen trading, meaning whether buyers or sellers were the ones crossing the bid-ask spread, is measured separately over the past year, which is as far back as that kind of data is free. The contracts whose delivery months lie inside a round's return windows are compared with the same contracts on every other 2026 date.
The short answer. No anomalous buying is visible, and the test was not strong enough for that absence to mean much. In none of the five rounds did the contracts for the return months show more buying anomalies, or a larger share of buyer-initiated screen trading, than the same months showed on ordinary 2026 dates. The smallest p-value is 0.34 for the autoencoder and 0.29 for the plain score, and the buyer-initiated share was 49 to 51% in every round. Section 8 measures the test's reach by planting a synthetic hedge on every ordinary 2026 date and counting how often the test catches it. A hedge of a quarter to a half of the owed barrels, bought evenly across the return months inside the ten-day window, is caught on 21 to 54% of dates in rounds 1, 1.b and 2 by the plain score, and less often by the autoencoder. A full hedge in round 2 is caught on 93% of dates. A hedge placed in the heavily traded December contracts, or any hedge in round 1.a, is caught little or no more often than chance. So the result rules out a large, fast, evenly spread hedge in round 2 with some confidence, and little else.
import json, warnings
from pathlib import Path
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.dates as mdates
from sklearn.ensemble import IsolationForest
from sklearn.neural_network import MLPRegressor
warnings.filterwarnings("ignore")
pd.set_option("display.width", 170)
pd.set_option("display.max_columns", 30)
CACHE = Path("../data/databento/spr_loans") # built by build_spr_loan_caches.py
BLUE, GREY, ORANGE, RED, INK, MUTED = "#2a78d6", "#a8a69e", "#eb6834", "#e34948", "#0b0b0b", "#52514e"
plt.rcParams.update({
"figure.dpi": 110, "savefig.dpi": 110, "font.size": 10, "axes.titlesize": 11,
"axes.spines.top": False, "axes.spines.right": False, "axes.edgecolor": "#c3c2b7",
"axes.labelcolor": MUTED, "xtick.color": MUTED, "ytick.color": MUTED,
"axes.grid": True, "grid.color": "#ecebe7", "grid.linewidth": 0.8, "lines.linewidth": 2,
"legend.frameon": False,
})
stats = pd.read_parquet(CACHE / "stats.parquet")
bars = pd.read_parquet(CACHE / "bars_1d.parquet")
defs = pd.read_parquet(CACHE / "defs.parquet")
flow = pd.read_parquet(CACHE / "flow_daily.parquet")
prov = json.loads((CACHE / "provenance.json").read_text())
print(f"statistics through {prov['stats_last_ref']}, screen trades through {prov['flow_last_date']}; "
f"estimated Databento cost of every pull: ${prov['total_est_cost_usd']:.2f}")
statistics through 2026-10-02, screen trades through 2026-10-05; estimated Databento cost of every pull: $0.00
1. The loans¶
The table lists the five rounds that have been awarded, taken from the Energy Department's requests for proposal and award notices on spr.doe.gov. Round 4 closes for bids on 6 October 2026 and is not yet awarded. The return windows are the delivery months each round offered for giving the crude back, combined over the storage sites. Each borrower chose its window when bidding, and the award notices do not say which. Inside a chosen window the Department sets an even monthly delivery schedule, so a borrower cannot leave everything to the last month. Round 1 offered a single window per site, running to September 2028.
The column hedge_contracts is the number of 1,000-barrel CL contracts needed to cover every barrel owed, at the round's lowest minimum premium. The true amount owed is somewhat larger: minimum premiums at some sites ran to 22% (round 1) and 24% (round 2), and bidders offered extra premium to win, so the figure is a floor and the real one is probably 5 to 10% higher. Hedging could add at most about that many contracts. A borrower might not hedge at all, might hedge only part, or might hedge with Brent futures or privately negotiated swaps, none of which appear in CL.
Round 3 deserves a note. It offered up to 40 million barrels in June 2026 at minimum premiums of only 8 to 9%, and companies took 0.5 million. It therefore serves as a placebo round: if the methods below find as much hedging around round 3 as around the others, they are finding noise.
ROUNDS = pd.DataFrame([
# name, solicitation, bids due (Central time), award notice "as of", barrels awarded,
# lowest minimum premium offered, return windows (union over storage sites, by delivery month)
("Round 1", "DE-RP96-26PO00001", "2026-03-17", "2026-03-20", 45_220_000, 0.18, [("2026-09", "2028-09")]),
("Round 1.a", "DE-RP96-26PO00002", "2026-04-06", "2026-04-10", 8_480_000, 0.17, [("2027-01", "2027-11")]),
("Round 1.b", "DE-RP96-26PO00003", "2026-04-13", "2026-04-17", 26_030_000, 0.185, [("2028-01", "2028-12")]),
("Round 2", "DE-RP96-26PO00004", "2026-05-04", "2026-05-11", 53_330_000, 0.18,
[("2027-01", "2028-04"), ("2028-10", "2029-07")]),
("Round 3", "DE-RP96-26PO00005", "2026-06-15", "2026-06-22", 500_000, 0.08, [("2027-01", "2028-03")]),
], columns=["round", "solicitation", "bids_due", "award", "barrels", "min_premium", "windows"])
ROUNDS["bids_due"] = pd.to_datetime(ROUNDS["bids_due"])
ROUNDS["award"] = pd.to_datetime(ROUNDS["award"])
ROUNDS["hedge_contracts"] = (ROUNDS["barrels"] * (1 + ROUNDS["min_premium"]) / 1000).round(-2).astype(int)
def in_windows(delivery, windows, shift=0):
"""True where the delivery month lies inside any window (optionally shifted by whole months)."""
d = pd.Series(pd.to_datetime(np.asarray(delivery)))
m = pd.Series(False, index=d.index)
for a, b in windows:
lo = pd.Timestamp(a + "-01") + pd.DateOffset(months=shift)
hi = pd.Timestamp(b + "-01") + pd.DateOffset(months=shift)
m |= (d >= lo) & (d <= hi)
return m.to_numpy()
ROUNDS.assign(
bids_due=ROUNDS.bids_due.dt.strftime("%d %b"), award=ROUNDS.award.dt.strftime("%d %b"),
barrels=(ROUNDS.barrels / 1e6).map("{:.1f}M".format),
min_premium=(100 * ROUNDS.min_premium).map("{:.1f}%".format),
windows=ROUNDS.windows.map(lambda ws: ", ".join(f"{pd.Timestamp(a):%b %Y} to {pd.Timestamp(b):%b %Y}"
for a, b in ws)),
)[["round", "solicitation", "bids_due", "award", "barrels", "min_premium", "hedge_contracts", "windows"]]
| round | solicitation | bids_due | award | barrels | min_premium | hedge_contracts | windows | |
|---|---|---|---|---|---|---|---|---|
| 0 | Round 1 | DE-RP96-26PO00001 | 17 Mar | 20 Mar | 45.2M | 18.0% | 53400 | Sep 2026 to Sep 2028 |
| 1 | Round 1.a | DE-RP96-26PO00002 | 06 Apr | 10 Apr | 8.5M | 17.0% | 9900 | Jan 2027 to Nov 2027 |
| 2 | Round 1.b | DE-RP96-26PO00003 | 13 Apr | 17 Apr | 26.0M | 18.5% | 30800 | Jan 2028 to Dec 2028 |
| 3 | Round 2 | DE-RP96-26PO00004 | 04 May | 11 May | 53.3M | 18.0% | 62900 | Jan 2027 to Apr 2028, Oct 2028 to Jul 2029 |
| 4 | Round 3 | DE-RP96-26PO00005 | 15 Jun | 22 Jun | 0.5M | 8.0% | 500 | Jan 2027 to Mar 2028 |
The same windows drawn against the delivery calendar. Rounds 1 and 2 cover most months from late 2026 to mid-2029. Round 2 leaves a gap from May to September 2028, and no round reaches beyond July 2029.
fig, ax = plt.subplots(figsize=(10, 3.0))
for k, r in enumerate(ROUNDS.itertuples()):
for a, b in r.windows:
x0 = mdates.date2num(pd.Timestamp(a + "-01"))
x1 = mdates.date2num(pd.Timestamp(b + "-01") + pd.offsets.MonthEnd(0))
ax.broken_barh([(x0, x1 - x0)], (k - 0.32, 0.64), color=BLUE, linewidth=0)
ax.text(mdates.date2num(pd.Timestamp("2029-12-31")) + 15, k, f"{r.barrels / 1e6:.1f}M barrels",
va="center", fontsize=9, color=MUTED)
ax.set_yticks(range(len(ROUNDS)), [f"{r.round}, bids {r.bids_due:%d %b}" for r in ROUNDS.itertuples()])
ax.invert_yaxis()
ax.set_xlim(mdates.date2num(pd.Timestamp("2026-08-01")), mdates.date2num(pd.Timestamp("2029-12-31")))
ax.xaxis.set_major_locator(mdates.MonthLocator(bymonth=[1, 7]))
ax.xaxis.set_major_formatter(mdates.DateFormatter("%b %Y"))
ax.grid(axis="y", visible=False)
ax.set_axisbelow(True)
ax.set_xlabel("delivery month of the returned crude")
ax.set_title("Return windows offered in each round", loc="left")
plt.tight_layout(); plt.show()
2. Data¶
All market data come from Databento's copy of the CME Globex feed, inside the CME subscription bundle, at no metered cost. Three kinds are used.
The exchange's daily statistics give, for every CL contract and trade date, the settlement price, the cleared volume, and the open interest. Open interest is the number of contracts open at the end of the day. It rises when a new buyer meets a new seller and falls when two existing holders close out against each other. It cannot say who is long, because every open contract has one buyer and one seller. Cleared volume counts every contract that changed hands, on the screen or off it. These statistics are used from September 2015 to 2 October 2026.
Daily bars of screen trading give the volume traded on the exchange's electronic order book, both in single contracts and in calendar spreads, which are listed instruments that buy one delivery month and sell another in one trade. In the distant months most screen trading arrives through calendar spreads: for CL contracts delivering in 2027 and 2028, 64 to 70% of cleared volume between March and September 2026 came as legs of screen-traded spreads, and only 11 to 13% as single-contract screen trades. Volume that is cleared but not traded in either of those ways is mostly blocks (large trades negotiated privately and reported to the exchange), exchanges of futures for physical crude, and multi-month packages such as butterflies and strips. I call it volume away from the screen.
Individual screen trades carry an aggressor side: the side that crossed the spread by accepting the standing offer or bid. The share of screen volume initiated by buyers is the most direct sign of buying pressure that public data offers. These trades are free only for the trailing year, here 6 October 2025 to 2 October 2026, so the direction analysis cannot use the ten-year history. The trade cache also holds only contracts delivering in 2027 or later.
# a trade date is one on which most listed contracts carry a settlement price
cnt = stats.dropna(subset=["settle"]).groupby("ref_date").size()
SESSIONS = pd.DatetimeIndex(sorted(cnt[cnt >= 0.5 * cnt.rolling(20, min_periods=1).max()].index))
pos = {d: k for k, d in enumerate(SESSIONS)}
print(f"{len(SESSIONS):,} trade dates, {SESSIONS[0]:%Y-%m-%d} to {SESSIONS[-1]:%Y-%m-%d}")
2,788 trade dates, 2015-09-01 to 2026-10-02
The daily figures are arranged as one table per measure, with a row per trade date and a column per contract. Three corrections are applied on the way. First, the exchange reuses instrument ids, so statistics records dated outside a contract's possible life (more than twelve years before its expiry, or after it) belong to some other instrument and are dropped; there are 10,486 of them. Second, daily bars dated on a Sunday or holiday hold the evening session that belongs to the next trade date and are moved there (1.3% of screen volume). Third, a contract's open interest is carried forward over days without a record, and nothing is counted after its expiry. Screen volume in calendar spreads is assigned to both delivery months of the spread. On 5.7% of contract-days the screen volume so assigned exceeds the cleared volume, by a median of about 3%. The daily bars are cut at midnight UTC while the exchange's trade date begins the previous evening, so the two do not align exactly. Those days count as having no volume away from the screen.
meta = defs.set_index("instrument_id")[["raw_symbol", "delivery", "expiration"]]
meta["expiration"] = pd.to_datetime(meta["expiration"], utc=True).dt.tz_localize(None)
# the exchange reuses instrument ids: keep only records dated within a CL contract's possible life
life = stats.ref_date.between(stats.expiration.dt.tz_localize(None) - pd.DateOffset(years=12),
stats.expiration.dt.tz_localize(None)) if stats.expiration.dt.tz is not None else \
stats.ref_date.between(stats.expiration - pd.DateOffset(years=12), stats.expiration)
print(f"statistics records outside their contract's possible life (reused ids, dropped): {(~life).sum():,}")
s = stats[life & stats.ref_date.isin(SESSIONS)]
OI = s.pivot_table(index="ref_date", columns="instrument_id", values="open_interest").reindex(SESSIONS)
CLR = s.pivot_table(index="ref_date", columns="instrument_id", values="cleared_volume").reindex(SESSIONS)
SET = s.pivot_table(index="ref_date", columns="instrument_id", values="settle").reindex(SESSIONS)
first = OI.notna().idxmax()
alive = pd.DataFrame({i: (SESSIONS >= first[i]) & (SESSIONS <= meta.loc[i, "expiration"]) for i in OI.columns},
index=SESSIONS)
OI = OI.ffill().where(alive)
SET = SET.ffill(limit=5).where(alive)
CLR = CLR.fillna(0).where(alive)
bars["date"] = pd.to_datetime(bars["ts_event"], utc=True).dt.tz_localize(None).dt.normalize()
# bars dated on a Sunday or holiday hold the evening session that belongs to the next trade date
bars = bars[(bars["date"] >= SESSIONS[0]) & (bars["date"] <= SESSIONS[-1])].copy()
nxt = SESSIONS.searchsorted(bars["date"])
moved = ~bars["date"].isin(SESSIONS)
bars["date"] = SESSIONS[nxt]
print(f"screen volume on bars dated outside trade dates, moved to the next trade date: "
f"{bars.volume[moved].sum() / bars.volume.sum():.1%}")
b = bars
outr = b[b.symbol.str.fullmatch(r"CL[FGHJKMNQUVXZ]\d{1,2}")]
SCR = outr.pivot_table(index="date", columns="instrument_id", values="volume", aggfunc="sum") \
.reindex(index=SESSIONS, columns=OI.columns).fillna(0)
cal = b[b.symbol.str.fullmatch(r"CL[FGHJKMNQUVXZ]\d{1,2}-CL[FGHJKMNQUVXZ]\d{1,2}")].copy()
cal[["leg1", "leg2"]] = cal.symbol.str.split("-", expand=True)
symmap = outr[["date", "symbol", "instrument_id"]].drop_duplicates().sort_values("date")
def leg_ids(df, leg):
"""A spread leg's raw symbol on its date -> instrument_id. Raw symbols repeat across
decades, so take the nearest-dated outright print of that symbol whose contract is
alive on the spread's date."""
q = df[["date", leg]].rename(columns={leg: "symbol"}).reset_index().sort_values("date")
m = pd.merge_asof(q, symmap, on="date", by="symbol", direction="nearest", tolerance=pd.Timedelta(days=400))
exp = meta["expiration"].reindex(m["instrument_id"]).to_numpy()
m.loc[~(exp >= m["date"].to_numpy()), "instrument_id"] = np.nan
return m.set_index("index")["instrument_id"].reindex(df.index)
cal["id1"], cal["id2"] = leg_ids(cal, "leg1"), leg_ids(cal, "leg2")
print(f"calendar-spread volume with both legs resolved: "
f"{cal.volume[cal[['id1', 'id2']].notna().all(axis=1)].sum() / cal.volume.sum():.1%}")
legs = pd.concat([cal[["date", "id1", "volume"]].rename(columns={"id1": "iid"}),
cal[["date", "id2", "volume"]].rename(columns={"id2": "iid"})]).dropna()
SPL = legs.pivot_table(index="date", columns="iid", values="volume", aggfunc="sum") \
.reindex(index=SESSIONS, columns=OI.columns).fillna(0)
other = (CLR - SCR - SPL)
print(f"contract-days on which outright + spread-leg screen volume exceeds cleared volume: "
f"{(other < -0.5).sum().sum() / alive.sum().sum():.2%} (counted as zero away from the screen)")
statistics records outside their contract's possible life (reused ids, dropped): 10,486
screen volume on bars dated outside trade dates, moved to the next trade date: 1.3%
calendar-spread volume with both legs resolved: 99.9% contract-days on which outright + spread-leg screen volume exceeds cleared volume: 5.69% (counted as zero away from the screen)
Two checks on the data. Databento marks a handful of days as reduced quality. Two fall inside event windows (16 March and 10 April 2026), and on both the exchange published its usual number of open-interest records. The aggressor side was checked on one outright contract and one calendar spread over a week of trades. Buyer-initiated trades printed on an uptick 71% and 89% of the time, seller-initiated trades 27% and 12%. The side field therefore means what it says, for spreads as well as for single contracts.
cond = ["2026-01-31", "2026-03-15", "2026-03-16", "2026-03-21", "2026-04-10", "2026-05-24", "2026-08-29"]
cond = [pd.Timestamp(d) for d in cond if pd.Timestamp(d) in pos]
raw26 = stats[(stats.ref_date >= "2026-01-02") & stats.open_interest.notna()].groupby("ref_date").size()
print("2026 reduced-quality trade dates (Databento): open-interest records published that day, "
f"against the 2026 median of {raw26.median():.0f} a day:")
print({f"{d:%Y-%m-%d}": int(raw26.get(d, 0)) for d in cond})
side = json.loads((CACHE / "side_check.json").read_text())
pd.DataFrame({f"{s} {k}": v for s, sides in side.items() for k, v in sides.items()}).T \
.rename(columns={"share_uptick": "share of price-changing trades that were upticks"})
2026 reduced-quality trade dates (Databento): open-interest records published that day, against the 2026 median of 60 a day:
{'2026-03-16': 60, '2026-04-10': 61}
| trades | share of price-changing trades that were upticks | |
|---|---|---|
| CLZ7 B | 7096.0 | 0.714910 |
| CLZ7 A | 7845.0 | 0.274697 |
| CLZ7-CLZ8 B | 4798.0 | 0.891622 |
| CLZ7-CLZ8 A | 5243.0 | 0.118825 |
3. The five measures¶
Each observation is one contract over a window of ten trade dates. Only contracts between 6 and 42 months from delivery at the end of the window are used, which is the range the return windows span. Only contracts with at least 500 open contracts at the start are used, because below that ratios mean little. Several thin 2029 months (April, May, July and later) fail that test. The five measures are:
- Open-interest growth: the natural logarithm of open interest at the end of the window over open interest at the start, each plus 100 contracts so that small contracts do not produce huge ratios. A value of 0.10 means about 10% growth.
- Growth against the neighbours: the open-interest growth minus the median growth of the two nearest eligible delivery months on each side. A hedge placed in one month shows up here; a build across the whole curve does not.
- Volume surge: the logarithm of the volume cleared in the window over ten times the contract's median daily cleared volume in the 60 trade dates before the window.
- Change in the share away from the screen: the share of the window's cleared volume traded away from the screen, minus the same share over the 60 trade dates before the window. The share itself drifted from about 4% of volume in 2016 to about 25% in 2026. Part of that is more block and package trading and part is the bar alignment described above, so the measure compares each contract with its own recent level.
- Price against the neighbours: the change over the window in the contract's settlement price minus the average of the two adjacent delivery months, as a percentage of price. Buying concentrated in one month would push that month up against its neighbours. This measure is weak, because the exchange sets the settlement prices of thinly traded distant months largely from spread relationships, which smooths away local pressure.
Same distance from delivery, same kind of contract. Contracts behave differently depending on how far they are from delivery, and December and June contracts carry far more open interest than the other months. Each measure is therefore turned into a robust score within its group: tenor bands of six months (6 to 12 months from delivery, 12 to 18, and so on to 42) crossed with three contract kinds (December, June, other). The robust score is the value minus the group's median over 2016 to 2025, divided by 1.4826 times the group's median absolute deviation. This is a version of the usual standardised score that a few extreme windows cannot distort. Scores are capped at ±10. The table shows how many reference windows each group holds.
L_MAIN = 10 # window length in trade dates
LOOK = 60 # trailing trade dates for the normal-volume yardstick
TENOR = (6, 42) # months from window end to delivery
MIN_OI = 500 # open contracts at window start
FEATS = ["oi_growth", "oi_vs_neighbours", "volume_surge", "other_share", "price_vs_neighbours"]
deliv = meta["delivery"].reindex(OI.columns)
kind = deliv.dt.month.map(lambda m: "Dec" if m == 12 else ("Jun" if m == 6 else "other"))
order = list(deliv.sort_values().index)
def window_features(j, L, OI=OI, CLR=CLR, other=other):
"""The five measures for every eligible contract over trade dates j-L+1..j
(the panels can be swapped for modified copies; see the power check)."""
t1 = SESSIONS[j]
oi0, oi1 = OI.iloc[j - L], OI.iloc[j]
tenor = (deliv - t1).dt.days / 30.44
ok = (oi0 >= MIN_OI) & oi1.notna() & tenor.between(*TENOR)
elig = [i for i in order if ok.get(i, False)]
if len(elig) < 3:
return None
g = np.log((oi1 + 100) / (oi0 + 100))
clr = CLR.iloc[j - L + 1:j + 1].sum()
base = CLR.iloc[max(0, j - L - LOOK):j - L].median()
vol = np.log((clr + 1) / (L * base + 1))
oth = other.iloc[j - L + 1:j + 1].clip(lower=0).sum()
share = (oth / clr.where(clr > 0)).clip(0, 1)
pclr = CLR.iloc[max(0, j - L - LOOK):j - L].sum()
poth = other.iloc[max(0, j - L - LOOK):j - L].clip(lower=0).sum()
share = share - (poth / pclr.where(pclr > 0)).clip(0, 1) # against the contract's own recent share
nb = {}
for k, i in enumerate(elig):
nn = elig[max(0, k - 2):k] + elig[k + 1:k + 3]
nb[i] = g[i] - np.median(g[nn])
s1, s0 = SET.iloc[j], SET.iloc[j - L]
live = [i for i in order if pd.notna(s1.get(i)) and pd.notna(s0.get(i))]
px = {}
for k, i in enumerate(live):
if 0 < k < len(live) - 1 and ok.get(i, False):
a, c = live[k - 1], live[k + 1]
px[i] = 100 * ((s1[i] - (s1[a] + s1[c]) / 2) - (s0[i] - (s0[a] + s0[c]) / 2)) / s1[i]
return pd.DataFrame({
"end": t1, "j": j, "iid": elig, "tenor": tenor[elig].values, "kind": kind[elig].values,
"oi0": oi0[elig].values, "oi1": oi1[elig].values,
"oi_growth": g[elig].values, "oi_vs_neighbours": [nb[i] for i in elig],
"volume_surge": vol[elig].values, "other_share": share[elig].values,
"price_vs_neighbours": [px.get(i, np.nan) for i in elig],
})
TB = [6, 12, 18, 24, 30, 36, 42]
REF_JS = [pos[d] for d in SESSIONS if pd.Timestamp("2016-01-01") <= d <= pd.Timestamp("2025-12-31")][::5]
STUDY_JS = [pos[d] for d in SESSIONS if d >= pd.Timestamp("2026-01-02")]
def build(L):
"""Reference (2016-2025, every fifth trade date) and study (every 2026 trade date)
contract-windows of length L, each measure turned into a robust score within its
tenor band x contract kind, using the reference's median and spread."""
ref = pd.concat([window_features(j, L) for j in REF_JS], ignore_index=True)
study = pd.concat([window_features(j, L) for j in STUDY_JS], ignore_index=True)
for df in (ref, study):
df["band"] = pd.cut(df.tenor, TB, include_lowest=True,
labels=[f"{a}-{b} months" for a, b in zip(TB[:-1], TB[1:])])
df["grp"] = df.band.astype(str) + " | " + df.kind
ref["year"] = ref.end.dt.year
med = ref.groupby("grp")[FEATS].median()
mad = ref.groupby("grp")[FEATS].agg(lambda x: 1.4826 * (x - x.median()).abs().median())
mad = mad.where(mad > 0, 1.4826 * (ref[FEATS] - ref[FEATS].median()).abs().median(), axis=1)
norm = (med, mad)
for df in (ref, study):
zscore(df, norm)
return ref, study, norm
def zscore(df, norm):
med, mad = norm
z = (df[FEATS] - med.reindex(df.grp).values) / mad.reindex(df.grp).values
df[[f"z_{f}" for f in FEATS]] = z.fillna(0).clip(-10, 10).values
return df
ref, study, NORM = build(L_MAIN)
ZF = [f"z_{f}" for f in FEATS]
print(f"reference: {len(ref):,} contract-windows ending {ref.end.min():%Y-%m-%d} to {ref.end.max():%Y-%m-%d}; "
f"2026: {len(study):,} contract-windows ending {study.end.min():%Y-%m-%d} to {study.end.max():%Y-%m-%d}")
ref.groupby(["band", "kind"], observed=True).size().unstack()[["Dec", "Jun", "other"]]
reference: 14,973 contract-windows ending 2016-01-04 to 2025-12-26; 2026: 5,729 contract-windows ending 2026-01-02 to 2026-10-02
| kind | Dec | Jun | other |
|---|---|---|---|
| band | |||
| 6-12 months | 249 | 256 | 2524 |
| 12-18 months | 253 | 246 | 2503 |
| 18-24 months | 249 | 257 | 2516 |
| 24-30 months | 257 | 249 | 2252 |
| 30-36 months | 246 | 252 | 1640 |
| 36-42 months | 255 | 249 | 520 |
4. The anomaly detectors¶
The reference set is every eligible contract-window ending on every fifth trade date from January 2016 to December 2025. That is the ten-year history against which 2026 is judged. Three scores are computed for every window, and each is reported as a percentile of the same score over the reference set, so 50 is typical and 95 means more unusual than 95% of the history.
Autoencoder. An autoencoder is a small neural network trained to reproduce its input after squeezing it through a narrower layer. Here the five robust scores pass through a layer of 16 units, then a bottleneck of 3, then 16 again, and back out to five. Trained on the history, the network learns which combinations of the five measures usually occur together; for example, that a volume surge usually comes with open-interest growth. A window it rebuilds badly is one whose combination of measures the history did not contain, and the average squared rebuild error is the anomaly score. The history is split into five two-year blocks, and five networks are trained, each on four of the blocks. Each block's reference errors come from the network that never saw it. A 2026 window is rebuilt by all five networks. Each network's error is placed among those held-out reference errors, and the window's percentile is the median of the five. That way a 2026 error is always compared with errors made on data the network had not seen.
Isolation forest. A second opinion with a different mechanism. The isolation forest grows 500 random trees, each of which cuts the data at random thresholds on randomly chosen measures until every point stands alone. Points unlike the rest are cut off in few steps, and the average number of steps is the anomaly score.
Plain build score. The average of the robust scores for open-interest growth and for growth against the neighbours. It is not machine learning, and it measures only the thing a hedge would most directly produce. It is here so that the machine-learning results can be checked against something readable.
Both detectors flag unusual windows in any direction, including collapses in open interest. A window counts as a buying anomaly only if its score is in the top 5% of the history and its open interest grew faster than both its group's norm and its neighbours. For the plain build score, a buying anomaly is simply a score in the top 5%. The table shows how often each test fires in the history and in 2026.
BLOCKS = [(2016, 2017), (2018, 2019), (2020, 2021), (2022, 2023), (2024, 2025)]
def autoencoder(seed):
# five scores -> 16 -> 3 -> 16 -> five scores
return MLPRegressor(hidden_layer_sizes=(16, 3, 16), activation="tanh", alpha=1e-3,
learning_rate_init=1e-3, max_iter=400, early_stopping=True,
validation_fraction=0.1, n_iter_no_change=20, random_state=seed)
def pct(x, refvals):
r = np.sort(np.asarray(refvals))
return 100 * np.searchsorted(r, np.asarray(x), side="right") / len(r)
def score(ref, study):
"""Autoencoder (cross-fitted by two-year blocks), isolation forest, and the plain
build score; each as a percentile of its own reference distribution. Returns the
fitted models too, so new windows can be scored the same way."""
Xr = ref[ZF].to_numpy() / 3.0 # z / 3 keeps typical values inside tanh's range
ae_r, feat_r = np.full(len(ref), np.nan), np.full((len(ref), len(ZF)), np.nan)
nets = []
for k, (y0, y1) in enumerate(BLOCKS):
hold = ref.year.between(y0, y1).to_numpy()
m = autoencoder(seed=11 + k).fit(Xr[~hold], Xr[~hold])
nets.append(m)
er = (m.predict(Xr[hold]) - Xr[hold]) ** 2
ae_r[hold], feat_r[hold] = er.mean(axis=1), er
iso = IsolationForest(n_estimators=500, max_samples=512, random_state=7).fit(ref[ZF])
iso_r = -iso.score_samples(ref[ZF])
bld_r = (ref.z_oi_growth + ref.z_oi_vs_neighbours) / 2
ref["ae_err"], ref["ae_pct"] = ae_r, pct(ae_r, ae_r)
ref["iso_pct"], ref["build_pct"] = pct(iso_r, iso_r), pct(bld_r, bld_r)
ref[[f"err_{f}" for f in FEATS]] = feat_r
flags(ref)
models = dict(nets=nets, iso=iso, ae_ref=ae_r, iso_ref=iso_r, build_ref=bld_r)
return ref, score_new(study, models), models
def flags(df):
# a buying anomaly: top 5% of history AND open interest growing faster than both the
# contract's own group norm and its neighbours
df["builds"] = (df.z_oi_growth > 0) & (df.oi_vs_neighbours > 0)
df["ae_flag"] = (df.ae_pct >= 95) & df.builds
df["iso_flag"] = (df.iso_pct >= 95) & df.builds
df["build_flag"] = df.build_pct >= 95
return df
def score_new(df, models):
"""Score windows outside the history. Each of the five networks rebuilds them; each
network's error is placed among the history's held-out errors (errors made by networks
on years they never saw, so like is compared with like), and the autoencoder percentile
is the median of the five."""
X = df[ZF].to_numpy() / 3.0
errs = np.array([(m.predict(X) - X) ** 2 for m in models["nets"]]) # nets x windows x measures
df["ae_err"] = np.median(errs.mean(axis=2), axis=0)
df["ae_pct"] = np.median([pct(e.mean(axis=1), models["ae_ref"]) for e in errs], axis=0)
df[[f"err_{f}" for f in FEATS]] = errs.mean(axis=0)
df["iso_pct"] = pct(-models["iso"].score_samples(df[ZF]), models["iso_ref"])
df["build_pct"] = pct((df.z_oi_growth + df.z_oi_vs_neighbours) / 2, models["build_ref"])
return flags(df)
ref, study, MODELS = score(ref, study)
print(f"autoencoder rebuild error (mean squared, z/3 units): held-out history median "
f"{np.nanmedian(ref.ae_err):.3f}, 2026 median {study.ae_err.median():.3f}")
pd.DataFrame({
"reference 2016-2025": [ref.ae_flag.mean(), ref.iso_flag.mean(), ref.build_flag.mean()],
"2026": [study.ae_flag.mean(), study.iso_flag.mean(), study.build_flag.mean()],
}, index=["autoencoder buying anomaly", "isolation forest buying anomaly", "plain build score top 5%"]) \
.map("{:.1%}".format)
autoencoder rebuild error (mean squared, z/3 units): held-out history median 0.018, 2026 median 0.034
| reference 2016-2025 | 2026 | |
|---|---|---|
| autoencoder buying anomaly | 1.9% | 5.4% |
| isolation forest buying anomaly | 2.9% | 6.2% |
| plain build score top 5% | 5.0% | 9.6% |
2026 is an unusual year for the whole distant curve. Buying anomalies occur in 5.4% of 2026 contract-windows by the autoencoder, against 1.9% in the reference years, which ranged from 1.2% (2021) to 3.0% (2025). The rate peaked in March, the month the Strait of Hormuz closed, at 13.3%. The isolation forest and the plain score tell the same story (6.2% against 2.9%, and 9.6% against 5.0%). This is why each loan round is judged against other 2026 dates, not against the history alone. The history says what is unusual for a contract at a given distance from delivery, and the 2026 placebo dates say whether the loan windows were unusual for 2026.
study["month"] = study.end.dt.to_period("M")
fig, (a1, a2) = plt.subplots(1, 2, figsize=(11, 3.3), gridspec_kw={"width_ratios": [1.1, 1]}, sharey=True)
yr = ref.groupby("year").ae_flag.mean() * 100
a1.bar(yr.index.astype(str), yr.values, color=GREY, width=0.7)
a1.bar(["2026"], [study.ae_flag.mean() * 100], color=BLUE, width=0.7)
a1.axhline(ref.ae_flag.mean() * 100, color=INK, linewidth=1)
a1.set_title("Autoencoder buying anomalies, share of contract-windows", loc="left")
a1.set_ylabel("percent of windows")
mo = study.groupby("month").ae_flag.mean() * 100
a2.bar([p.strftime("%b") for p in mo.index], mo.values, color=BLUE, width=0.7)
a2.axhline(ref.ae_flag.mean() * 100, color=INK, linewidth=1)
a2.set_title("2026, by month in which the window ends", loc="left")
for a in (a1, a2):
a.grid(axis="x", visible=False)
a1.tick_params(axis="x", labelsize=8)
plt.tight_layout(); plt.show()
5. The test: return-window contracts around each round¶
For each round the event window is the ten trade dates ending three trade dates after the award notice, which covers the bid deadline in every case. The test counts buying anomalies among the contracts inside the round's return windows. The comparison is the same count, for the same return windows, on every 2026 date whose ten-day window does not overlap any event window: 112 placebo dates. The p-value is the share of placebo dates with a count at least as large as the event's, so a small p-value means the loan window stands out. The placebo dates are consecutive and their ten-day windows overlap, so they hold the information of only about a dozen independent windows; the p-values are coarse, and none below about 0.1 should be read as precise. The table also gives the plain build score's difference between the in-window and out-of-window contracts, in percentile points, with its own placebo p-value.
def evaluate(study, L, offset=3, shift=0, keep=None, anchor="award"):
"""For each round: the event window (L trade dates ending `offset` trade dates after the
award), the in-window vs out-of-window gap in mean percentile for each score, and the
same gap at every 2026 placebo window that does not overlap an event window."""
st = study if keep is None else study[keep(study)]
by_end = {e: g for e, g in st.groupby("end")}
if anchor == "award":
ends = {r.round: pos[SESSIONS[SESSIONS.searchsorted(r.award)]] + offset for r in ROUNDS.itertuples()}
else: # window starts on the bid deadline
ends = {r.round: pos[SESSIONS[SESSIONS.searchsorted(r.bids_due)]] + L - 1 for r in ROUNDS.itertuples()}
blocked = set()
for j in ends.values():
blocked |= set(range(j - L, j + L + 1))
placebo = [SESSIONS[j] for j in STUDY_JS if j not in blocked and SESSIONS[j] in by_end]
rows, pls, evs = [], {}, []
def gaps(g, m):
a, b = g[m], g[~m]
return dict(n_in=len(a), n_out=len(b),
ae_in=a.ae_pct.mean(), ae_gap=a.ae_pct.mean() - b.ae_pct.mean(),
iso_in=a.iso_pct.mean(), iso_gap=a.iso_pct.mean() - b.iso_pct.mean(),
build_in=a.build_pct.mean(), build_gap=a.build_pct.mean() - b.build_pct.mean(),
ae_flags_in=int(a.ae_flag.sum()), ae_flags_out=int(b.ae_flag.sum()),
iso_flags_in=int(a.iso_flag.sum()), build_flags_in=int(a.build_flag.sum()))
for r in ROUNDS.itertuples():
j = ends[r.round]
g = by_end[SESSIONS[j]].copy()
m = in_windows(meta.delivery.reindex(g.iid), r.windows, shift)
g["in_window"], g["round"] = m, r.round
evs.append(g)
obs = gaps(g, m)
pl = []
for e in placebo:
h = by_end[e]
mm = in_windows(meta.delivery.reindex(h.iid), r.windows, shift)
if mm.sum() >= 3 and (~mm).sum() >= 3:
pl.append(gaps(h, mm))
pl = pd.DataFrame(pl)
pls[r.round] = pl
row = dict(round=r.round, start=SESSIONS[j - L], end=SESSIONS[j], **obs)
for k in ("ae_gap", "iso_gap", "build_gap"):
row["p_" + k] = (1 + (pl[k] >= obs[k]).sum()) / (1 + len(pl))
# directional tests: buying anomalies inside the return windows, against the number
# inside the same windows at every placebo date
for k in ("ae_flags_in", "iso_flags_in", "build_flags_in"):
row["p_" + k] = (1 + (pl[k] >= obs[k]).sum()) / (1 + len(pl))
row["placebo_mean_" + k] = pl[k].mean()
row["placebo_windows"] = len(pl)
rows.append(row)
ev = pd.concat(evs, ignore_index=True)
ev["symbol"] = meta.raw_symbol.reindex(ev.iid).values
ev["delivery"] = meta.delivery.reindex(ev.iid).values
return pd.DataFrame(rows), pls, ev
RES, PLACEBO, EV = evaluate(study, L_MAIN)
RES[["round", "start", "end", "n_in", "n_out",
"ae_flags_in", "placebo_mean_ae_flags_in", "p_ae_flags_in",
"iso_flags_in", "placebo_mean_iso_flags_in", "p_iso_flags_in",
"build_flags_in", "placebo_mean_build_flags_in", "p_build_flags_in",
"build_gap", "p_build_gap", "placebo_windows"]] \
.assign(start=RES.start.dt.strftime("%d %b"), end=RES.end.dt.strftime("%d %b")).round(2)
| round | start | end | n_in | n_out | ae_flags_in | placebo_mean_ae_flags_in | p_ae_flags_in | iso_flags_in | placebo_mean_iso_flags_in | p_iso_flags_in | build_flags_in | placebo_mean_build_flags_in | p_build_flags_in | build_gap | p_build_gap | placebo_windows | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | Round 1 | 11 Mar | 25 Mar | 24 | 7 | 1 | 0.46 | 0.34 | 1 | 0.88 | 0.52 | 1 | 0.96 | 0.42 | -14.56 | 0.73 | 112 |
| 1 | Round 1.a | 31 Mar | 15 Apr | 11 | 19 | 0 | 0.20 | 1.00 | 0 | 0.54 | 1.00 | 0 | 0.30 | 1.00 | -11.11 | 0.77 | 112 |
| 2 | Round 1.b | 08 Apr | 22 Apr | 12 | 18 | 1 | 0.47 | 0.34 | 1 | 0.34 | 0.20 | 2 | 1.02 | 0.31 | 10.67 | 0.24 | 112 |
| 3 | Round 2 | 30 Apr | 14 May | 23 | 6 | 1 | 1.09 | 0.67 | 0 | 1.63 | 1.00 | 4 | 2.29 | 0.29 | -8.06 | 0.88 | 112 |
| 4 | Round 3 | 10 Jun | 25 Jun | 15 | 14 | 0 | 0.20 | 1.00 | 0 | 0.62 | 1.00 | 0 | 0.48 | 1.00 | -22.58 | 0.99 | 112 |
The same comparison for the detectors' raw percentiles, which are not directional (a contract whose open interest collapsed scores as high as one whose open interest doubled). The columns are the mean percentile of the in-window contracts, the difference from the out-of-window contracts, and its placebo p-value. Nothing stands out here either. Round 1 comes closest (p = 0.11 for the autoencoder), which, among ten p-values, is what chance alone produces.
RES[["round", "ae_in", "ae_gap", "p_ae_gap", "iso_in", "iso_gap", "p_iso_gap"]].round(2)
| round | ae_in | ae_gap | p_ae_gap | iso_in | iso_gap | p_iso_gap | |
|---|---|---|---|---|---|---|---|
| 0 | Round 1 | 75.24 | 0.57 | 0.11 | 67.27 | 9.59 | 0.17 |
| 1 | Round 1.a | 39.06 | -30.75 | 0.92 | 57.35 | -5.33 | 0.81 |
| 2 | Round 1.b | 62.32 | 4.80 | 0.23 | 48.11 | -17.07 | 0.50 |
| 3 | Round 2 | 53.41 | -22.92 | 0.97 | 56.57 | -16.36 | 0.92 |
| 4 | Round 3 | 66.51 | -6.55 | 0.39 | 70.12 | -7.89 | 0.83 |
Reading the counts. The autoencoder found one buying anomaly inside the return windows in round 1, none in round 1.a, one in round 1.b, one in round 2, and none in round 3. The 2026 placebo dates average 0.46, 0.20, 0.47, 1.09 and 0.20 for the same windows, giving p-values of 0.34, 1.00, 0.34, 0.67 and 1.00. The isolation forest gives the same picture (smallest p-value 0.20). The plain build score found 1, 0, 2, 4 and 0, against placebo averages of 1.0, 0.3, 1.0, 2.3 and 0.5 (smallest p-value 0.29). The difference in build score between in-window and out-of-window contracts is negative in four of the five rounds.
The figure shows each round's event window, one dot per eligible contract, plotted at its delivery month. Blue dots are inside the round's return windows, grey dots outside, and red rings mark the autoencoder's buying anomalies. One pattern stands out, and it is not tied to the loans. January 2029 is flagged in four of the five event windows, February 2029 in three and October 2028 in two, whether or not they lie in that round's return windows. Open interest in that part of the curve built throughout 2026. That fits producers selling future production forward at war prices better than it fits any single round of loans, though the data cannot say who the sellers were.
fig, axes = plt.subplots(len(ROUNDS), 1, figsize=(10, 1.9 * len(ROUNDS)), sharex=True)
for ax, r in zip(axes, RES.itertuples()):
w = EV[EV["round"] == r.round]
for m, c, lab in [(~w.in_window, GREY, "outside the round's return windows"),
(w.in_window, BLUE, "inside the round's return windows")]:
ax.scatter(w.delivery[m], w.ae_pct[m], s=34, color=c, edgecolor="white", linewidth=1, label=lab, zorder=3)
hits = w[w.ae_flag]
ax.scatter(hits.delivery, hits.ae_pct, s=95, facecolor="none", edgecolor=RED, linewidth=1.5, zorder=4,
label="buying anomaly (top 5%, open interest building)")
ax.axhline(95, color=INK, linewidth=0.8)
ax.set_ylim(0, 103)
ax.set_title(f"{r.round}: ten trade dates {r.start:%d %b} to {r.end:%d %b} 2026", loc="left")
ax.set_ylabel("percentile")
axes[0].legend(loc="lower left", ncol=3, fontsize=8)
axes[-1].xaxis.set_major_locator(mdates.MonthLocator(bymonth=[1, 7]))
axes[-1].xaxis.set_major_formatter(mdates.DateFormatter("%b %Y"))
axes[-1].set_xlabel("delivery month of the contract")
plt.tight_layout(); plt.show()
The counts against their placebo distributions. Grey bars show how often each count occurred on the 112 placebo dates; the blue bar marks the event window's count.
fig, axes = plt.subplots(2, len(ROUNDS), figsize=(12, 4.4))
for k, r in enumerate(RES.itertuples()):
for row, (col, lab) in enumerate([("ae_flags_in", "autoencoder"), ("build_flags_in", "plain build score")]):
ax = axes[row, k]
vals = PLACEBO[r.round][col]
cnt = vals.value_counts().sort_index()
ax.bar(cnt.index, 100 * cnt.values / len(vals), color=GREY, width=0.7)
obs = getattr(r, col)
ax.bar([obs], [100 * (vals == obs).mean()], color=BLUE, width=0.7)
ax.set_xticks(range(0, max(int(vals.max()), obs) + 1))
ax.set_title(f"{r.round}: {obs} found, p = {getattr(r, 'p_' + col):.2f}" if row == 0
else f"{obs} found, p = {getattr(r, 'p_' + col):.2f}", fontsize=10)
ax.grid(axis="x", visible=False)
axes[0, k].set_ylim(0, 100); axes[1, k].set_ylim(0, 100)
axes[0, 0].set_ylabel("autoencoder\n% of placebo dates")
axes[1, 0].set_ylabel("plain build score\n% of placebo dates")
fig.supxlabel("buying anomalies among the return-window contracts (blue: the event window; grey: 2026 placebo dates)",
fontsize=10, color=MUTED)
plt.tight_layout(); plt.show()
6. The individual contracts flagged¶
Every return-window contract that any of the three scores flagged during an event window. away_from_screen_change is measure 4 (negative means a smaller share away from the screen than the contract's recent norm), and main_driver is the measure the autoencoder rebuilt worst.
Most of these are flagged by one score only, and most are driven by a volume surge rather than by an isolated build of open interest. Three deserve a closer look because they were raised earlier in this investigation, all in round 2's return windows. December 2028 (CLZ8) added 19,900 contracts, 31% of its open interest, and passes the plain build score at the 96th percentile. The build began on 4 May, the bid deadline, which was also the day the July 2026 contract settled at $101.51 before falling about $10 over the next two sessions; producers selling forward into a price spike would leave the same mark. March 2028 (CLH8) and June 2029 (CLM9) each added 4,000 to 6,000 contracts and pass the plain build score (95th and 97th percentile). None of the three passes the autoencoder or the isolation forest.
FL = EV[EV.in_window & (EV.ae_flag | EV.iso_flag | EV.build_flag)].copy()
FL["oi_change"] = (FL.oi1 - FL.oi0).astype(int)
FL["main_driver"] = FL[[f"err_{f}" for f in FEATS]].idxmax(axis=1).str.replace("err_", "")
FL[["round", "symbol", "delivery", "oi0", "oi_change", "oi_vs_neighbours", "volume_surge", "other_share",
"ae_pct", "iso_pct", "build_pct", "main_driver"]] \
.rename(columns={"other_share": "away_from_screen_change"}) \
.assign(delivery=lambda d: d.delivery.dt.strftime("%b %Y"), oi0=lambda d: d.oi0.astype(int)) \
.sort_values(["round", "delivery"]).round(2)
| round | symbol | delivery | oi0 | oi_change | oi_vs_neighbours | volume_surge | away_from_screen_change | ae_pct | iso_pct | build_pct | main_driver | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 18 | Round 1 | CLJ8 | Apr 2028 | 2758 | 1027 | 0.17 | 2.24 | -0.23 | 94.04 | 94.96 | 96.89 | volume_surge |
| 8 | Round 1 | CLM7 | Jun 2027 | 113227 | 11589 | 0.05 | 1.32 | -0.06 | 96.53 | 94.96 | 77.39 | volume_surge |
| 1 | Round 1 | CLX6 | Nov 2026 | 60001 | 5822 | 0.06 | 1.38 | 0.13 | 89.11 | 96.33 | 69.75 | volume_surge |
| 77 | Round 1.b | CLH8 | Mar 2028 | 9192 | 5102 | 0.39 | 1.40 | 0.16 | 94.08 | 95.84 | 98.07 | volume_surge |
| 85 | Round 1.b | CLX8 | Nov 2028 | 4296 | 430 | 0.02 | 1.65 | -0.45 | 98.82 | 90.36 | 95.33 | oi_growth |
| 115 | Round 2 | CLZ8 | Dec 2028 | 63588 | 19922 | 0.21 | 0.90 | -0.01 | 14.80 | 82.72 | 95.89 | other_share |
| 116 | Round 2 | CLF9 | Jan 2029 | 2610 | 303 | 0.07 | 2.74 | -0.15 | 97.96 | 90.52 | 96.84 | volume_surge |
| 119 | Round 2 | CLM9 | Jun 2029 | 8820 | 5799 | 0.47 | 2.52 | 0.11 | 79.15 | 89.15 | 96.75 | volume_surge |
| 106 | Round 2 | CLH8 | Mar 2028 | 12192 | 4374 | 0.17 | 1.11 | -0.04 | 81.97 | 75.82 | 95.09 | volume_surge |
Open interest in March 2028 and June 2029 and their nearby months, indexed to 100 on 30 April 2026. The shaded band is round 2's event window and the vertical line its bid deadline. March 2028 rose 36% during the window, about as fast as June 2028 (31%), which lies outside every round 2 return window. April 2028 rose another 78% in the week after the window closed. Nothing about March 2028 sets it apart from its neighbours. June 2029 is the more interesting case. Its open interest rose 66% during the window and had doubled by 21 May, while March 2029 rose 5% and December 2029 fell 3%. A borrower hedging a return window around mid-2029 would buy exactly this contract. But one contract building over three weeks is also the kind of thing that happens several times a year in a busy curve, and the placebo counts above say that, counted across the return windows, round 2's window was not unusual for 2026.
alive26 = meta[meta.expiration >= pd.Timestamp("2026-01-01")].reset_index().set_index("raw_symbol")["instrument_id"]
cases = [("CLH8", ["CLG8", "CLJ8", "CLM8"]), ("CLM9", ["CLH9", "CLZ8", "CLZ9"])]
fig, axes = plt.subplots(1, 2, figsize=(11, 3.4), sharey=False)
span = OI.loc["2026-04-15":"2026-06-05"]
r2 = RES.set_index("round").loc["Round 2"]
for ax, (sym, nbs) in zip(axes, cases):
ax.axvspan(r2.start, r2.end, color="#e8eef8", zorder=0)
ends = []
for s_ in nbs + [sym]:
y = 100 * span[alive26[s_]] / span[alive26[s_]].loc[r2.start]
ax.plot(y.index, y, color=BLUE if s_ == sym else GREY, linewidth=2 if s_ == sym else 1.5,
zorder=3 if s_ == sym else 2)
ends.append([y.iloc[-1], s_])
ends.sort()
for k in range(1, len(ends)): # keep end labels at least 6 points apart
ends[k][0] = max(ends[k][0], ends[k - 1][0] + 6)
for yv, s_ in ends:
ax.annotate(s_, (span.index[-1], yv), xytext=(4, 0), textcoords="offset points", va="center",
fontsize=9 if s_ == sym else 8, color=INK if s_ == sym else MUTED)
ax.axvline(ROUNDS.set_index("round").loc["Round 2", "bids_due"], color=INK, linewidth=0.8)
ax.set_title(f"{sym} ({meta.delivery[alive26[sym]]:%b %Y}) against nearby months", loc="left")
ax.xaxis.set_major_formatter(mdates.DateFormatter("%d %b"))
axes[0].set_ylabel("open interest, 100 = 30 Apr")
plt.tight_layout(); plt.show()
7. Size: how much did return-window open interest grow beyond the rest of the curve?¶
A rough check of scale, independent of the detectors. For each round, the open interest added in the return-window contracts during the event window is compared with what those contracts would have added had they grown at the same percentage as the out-of-window contracts. If borrowers hedged everything they owed inside the window, the excess would approach the contracts owed.
The excess is negative in rounds 1, 1.a and 3. It is positive in round 1.b (12,000 contracts, against at least 30,800 owed) and round 2 (9,100 contracts, against at least 62,900 owed). The out-of-window group is small, six to 19 contracts depending on the round, so these figures are rough. At most they leave room for borrowers in rounds 1.b and 2 to have hedged about two-fifths and one-seventh of what they owed inside the window. They do not show that anyone did.
ex = []
for r in ROUNDS.itertuples():
w = EV[EV["round"] == r.round]
a, b = w[w.in_window], w[~w.in_window]
g_out = (b.oi1 - b.oi0).sum() / b.oi0.sum()
ex.append(dict(round=r.round, contracts_owed_if_fully_hedged=r.hedge_contracts,
open_interest_change_in_window=int((a.oi1 - a.oi0).sum()),
growth_outside_pct=round(100 * g_out, 1),
change_beyond_outside_growth=int(((a.oi1 - a.oi0) - a.oi0 * g_out).sum())))
pd.DataFrame(ex)
| round | contracts_owed_if_fully_hedged | open_interest_change_in_window | growth_outside_pct | change_beyond_outside_growth | |
|---|---|---|---|---|---|
| 0 | Round 1 | 53400 | 45630 | 5.3 | -7568 |
| 1 | Round 1.a | 9900 | 20666 | 7.1 | -11067 |
| 2 | Round 1.b | 30800 | 17778 | 3.7 | 11956 |
| 3 | Round 2 | 62900 | 60759 | 6.5 | 9148 |
| 4 | Round 3 | 500 | -15161 | 4.9 | -52403 |
8. Could these tests have seen a hedge?¶
A negative result says something only if the test could have found the effect. To measure that, I planted a synthetic hedge in the data and ran the same count test, once at each of the 112 placebo dates. Each time, the planted count is compared with the unplanted placebo counts, and the hedge counts as caught if its p-value is 0.05 or less. The share of dates on which it is caught is the test's power. The hedge is a fraction (0, 25%, 50% or 100%) of the contracts owed, added to open interest evenly over the ten trade dates and held, with the same amount added to cleared volume. Two placements are tried. In the first, the hedge is spread equally over every eligible contract inside the return windows, the natural hedge for an even return schedule. In the second, all of it goes into the December contracts inside the windows, the liquid months a hedger might use and later roll. The last two columns give, for the actual event window, the highest plain-score percentile among the planted contracts before and after planting.
The 0% rows are the false-alarm rate, and they come out at 0 to 4%, as they should for a test run at 5%. Spread evenly, the plain build score catches a quarter of the owed barrels on 21% (round 1), 29% (round 1.b) and 48% (round 2) of dates. It catches half on 38%, 41% and 54%, and the full amount on 41%, 42% and 93%. The autoencoder does worse, catching at most 39% and often under 20%. When every return month builds at once, no single month stands out from its neighbours, and growth against the neighbours is one of its main inputs. Round 1.a's hedge is too small to see at any size (900 contracts a month added to 2027 contracts holding 20,000 to 130,000), and its rates stay at the false-alarm level. Placed in December contracts, no hedge is caught more often than 12% of the time. The December contracts hold 60,000 to 250,000 open contracts and normally swing by several percent in ten days. A build there also lowers the growth against the neighbours of the adjacent months, which removes their flags, so the round's count hardly moves. In round 1.b the planted December 2028 contract itself rises from the 39th to the 96th percentile at half the owed barrels. In round 2 December 2028 was already at the 96th percentile before anything was planted, so that round's December case cannot be judged this way.
TBL = [f"{a}-{b} months" for a, b in zip(TB[:-1], TB[1:])]
def planted_window(j, windows, frac, owed, placement):
"""Score the window ending at trade date j after planting a synthetic hedge: frac x owed
contracts, bought evenly over the window and held, spread equally over the eligible
contracts inside `windows` ("even") or over the December contracts among them
("december"). The planted volume counts as screen volume, so it slightly lowers the
away-from-screen share. Returns the scored window, the in-window mask, the planted ids."""
base = by_end_all[SESSIONS[j]]
m0 = in_windows(meta.delivery.reindex(base.iid), windows)
tgt = list(base.iid[m0]) if placement == "even" else list(base.iid[m0 & (base.kind == "Dec").to_numpy()])
if not tgt:
return None
per = frac * owed / len(tgt)
k = np.arange(len(SESSIONS))
ramp = np.clip((k - (j - L_MAIN)) / L_MAIN, 0, 1)
daily = ((k > j - L_MAIN) & (k <= j)) / L_MAIN
OI2, CLR2 = OI.copy(), CLR.copy()
OI2[tgt] = OI2[tgt].add(per * ramp, axis=0)
CLR2[tgt] = CLR2[tgt].add(per * daily, axis=0)
g = window_features(j, L_MAIN, OI=OI2, CLR=CLR2, other=other)
g["grp"] = pd.cut(g.tenor, TB, include_lowest=True, labels=TBL).astype(str) + " | " + g.kind
g = score_new(zscore(g, NORM), MODELS)
return g, in_windows(meta.delivery.reindex(g.iid), windows), tgt, per
by_end_all = {e: g for e, g in study.groupby("end")}
placebo_js = {}
for r in RES.itertuples():
blocked = set().union(*[set(range(pos[e] - L_MAIN, pos[e] + L_MAIN + 1)) for e in RES.end])
placebo_js[r.round] = [j for j in STUDY_JS if j not in blocked]
prow = []
for r, rr in zip(RES.itertuples(), ROUNDS.itertuples()):
if r.round == "Round 3":
continue
pl = PLACEBO[r.round]
for placement, fracs in (("even", (0.0, 0.25, 0.5, 1.0)), ("december", (0.25, 0.5, 1.0))):
for frac in fracs:
# (i) planted at every placebo date (0.0 = nothing planted: the false-alarm rate): how often would the count test fire?
hits_ae = hits_b = n = 0
for j in placebo_js[r.round]:
out = planted_window(j, rr.windows, frac, rr.hedge_contracts, placement)
if out is None:
continue
g, m, _, _ = out
n += 1
hits_ae += (1 + (pl.ae_flags_in >= g.ae_flag[m].sum()).sum()) / (1 + len(pl)) <= 0.05
hits_b += (1 + (pl.build_flags_in >= g.build_flag[m].sum()).sum()) / (1 + len(pl)) <= 0.05
# (ii) planted in the event window itself, against the unplanted percentiles
out = planted_window(pos[r.end], rr.windows, frac, rr.hedge_contracts, placement)
if out is None:
continue
g, m, tgt, per = out
before = EV[(EV["round"] == r.round) & EV.iid.isin(tgt)].set_index("iid").build_pct
after = g.set_index("iid").build_pct.reindex(before.index)
prow.append(dict(round=r.round, placement=placement, share_of_owed=frac, contracts_each=int(per),
planted=len(tgt), detect_rate_autoencoder=hits_ae / max(n, 1),
detect_rate_build=hits_b / max(n, 1), placebo_dates=n,
event_build_pct_before=before.max(), event_build_pct_after=after.max()))
POWER = pd.DataFrame(prow)
POWER.round(2)
| round | placement | share_of_owed | contracts_each | planted | detect_rate_autoencoder | detect_rate_build | placebo_dates | event_build_pct_before | event_build_pct_after | |
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | Round 1 | even | 0.00 | 0 | 24 | 0.03 | 0.04 | 112 | 96.89 | 96.89 |
| 1 | Round 1 | even | 0.25 | 556 | 24 | 0.11 | 0.21 | 112 | 96.89 | 98.24 |
| 2 | Round 1 | even | 0.50 | 1112 | 24 | 0.25 | 0.38 | 112 | 96.89 | 98.91 |
| 3 | Round 1 | even | 1.00 | 2225 | 24 | 0.29 | 0.41 | 112 | 96.89 | 100.00 |
| 4 | Round 1 | december | 0.25 | 6675 | 2 | 0.03 | 0.03 | 112 | 66.13 | 77.43 |
| 5 | Round 1 | december | 0.50 | 13350 | 2 | 0.03 | 0.03 | 112 | 66.13 | 84.37 |
| 6 | Round 1 | december | 1.00 | 26700 | 2 | 0.04 | 0.05 | 112 | 66.13 | 92.21 |
| 7 | Round 1.a | even | 0.00 | 0 | 11 | 0.04 | 0.04 | 112 | 84.29 | 84.29 |
| 8 | Round 1.a | even | 0.25 | 225 | 11 | 0.04 | 0.03 | 112 | 84.29 | 84.62 |
| 9 | Round 1.a | even | 0.50 | 450 | 11 | 0.04 | 0.03 | 112 | 84.29 | 84.84 |
| 10 | Round 1.a | even | 1.00 | 900 | 11 | 0.04 | 0.05 | 112 | 84.29 | 85.21 |
| 11 | Round 1.b | even | 0.00 | 0 | 12 | 0.00 | 0.04 | 112 | 98.07 | 98.07 |
| 12 | Round 1.b | even | 0.25 | 641 | 12 | 0.19 | 0.29 | 112 | 98.07 | 97.91 |
| 13 | Round 1.b | even | 0.50 | 1283 | 12 | 0.39 | 0.41 | 112 | 98.07 | 98.95 |
| 14 | Round 1.b | even | 1.00 | 2566 | 12 | 0.17 | 0.42 | 112 | 98.07 | 99.35 |
| 15 | Round 1.b | december | 0.25 | 7700 | 1 | 0.00 | 0.04 | 112 | 39.24 | 88.57 |
| 16 | Round 1.b | december | 0.50 | 15400 | 1 | 0.00 | 0.04 | 112 | 39.24 | 95.70 |
| 17 | Round 1.b | december | 1.00 | 30800 | 1 | 0.00 | 0.04 | 112 | 39.24 | 98.36 |
| 18 | Round 2 | even | 0.00 | 0 | 23 | 0.04 | 0.01 | 112 | 96.84 | 96.84 |
| 19 | Round 2 | even | 0.25 | 683 | 23 | 0.00 | 0.48 | 112 | 96.84 | 97.86 |
| 20 | Round 2 | even | 0.50 | 1367 | 23 | 0.00 | 0.54 | 112 | 96.84 | 98.96 |
| 21 | Round 2 | even | 1.00 | 2734 | 23 | 0.11 | 0.93 | 112 | 96.84 | 100.00 |
| 22 | Round 2 | december | 0.25 | 7862 | 2 | 0.03 | 0.00 | 112 | 95.89 | 97.50 |
| 23 | Round 2 | december | 0.50 | 15725 | 2 | 0.02 | 0.01 | 112 | 95.89 | 98.36 |
| 24 | Round 2 | december | 1.00 | 31450 | 2 | 0.02 | 0.12 | 112 | 95.89 | 99.11 |
9. Direction on the screen¶
Open interest says contracts were opened, not who bought. The screen trades say which side was in a hurry. For each round, the share of screen volume in the return-window contracts that was initiated by buyers is compared with the same share on the 2026 placebo dates. Only contracts delivering in 2027 or later are covered, so round 1's return months from September to December 2026 are not in this test. Spread trades are split into their legs: buying the spread "A minus B" counts as buying A and selling B. A hedger who sells a near month and buys a return month through a spread is therefore counted as buying the return month. A spread with both legs inside the return windows adds equal buying and selling and pulls the share towards 50%, so the test is run twice, with and without such spreads. 16.4% of screen volume carries no aggressor side and is left out. I believe these are mostly opening-auction trades and implied fills between the spread and single-contract books, but did not check.
The buyer-initiated share in the return-window contracts was 50.1%, 50.5%, 49.0%, 49.2% and 49.9% in the five rounds counting all screen volume. Leaving out spreads inside the windows, it was 50.3%, 50.9%, 48.6%, 48.1% and 49.8%. None of these is high for 2026: the placebo p-values, the share of placebo dates with a share at least as high, are 0.54 or more in both versions. In round 2 the return-window contracts were net sold on the screen by about 20,000 contracts during the window. A hedger who rests limit orders and lets others cross the spread is counted as a seller-initiated fill, so this test misses patient buying; it would catch a hedger in a hurry.
fl = flow.copy()
fl["date"] = pd.to_datetime(fl["date"])
fl = fl[fl.date.isin(SESSIONS)]
fl["net"], fl["gross"] = fl["buy"] - fl["sell"], fl["buy"] + fl["sell"]
recent = meta[meta.expiration >= fl.date.min()].reset_index()
assert recent.raw_symbol.is_unique
sym2iid = recent.set_index("raw_symbol")["instrument_id"]
FLOW_CUT = pd.Timestamp("2027-01-01") # the trade cache holds contracts delivering 2027 or later
outr_f = fl[fl.symbol.str.fullmatch(r"CL[FGHJKMNQUVXZ]\d{1,2}")].assign(i1=lambda d: d.symbol.map(sym2iid))
spr_f = fl[fl.symbol.str.fullmatch(r"CL[FGHJKMNQUVXZ]\d{1,2}-CL[FGHJKMNQUVXZ]\d{1,2}")].copy()
spr_f[["l1", "l2"]] = spr_f.symbol.str.split("-", expand=True)
spr_f["i1"], spr_f["i2"] = spr_f.l1.map(sym2iid), spr_f.l2.map(sym2iid)
print(f"screen trades whose contracts all resolve: outrights {outr_f.i1.notna().mean():.1%}, "
f"spreads {(spr_f.i1.notna() & spr_f.i2.notna()).mean():.1%}; unsigned share of screen volume "
f"{fl['none'].sum() / (fl['none'].sum() + fl.gross.sum()):.1%}")
ON = outr_f.pivot_table(index="date", columns="i1", values="net", aggfunc="sum").reindex(SESSIONS).fillna(0)
OG = outr_f.pivot_table(index="date", columns="i1", values="gross", aggfunc="sum").reindex(SESSIONS).fillna(0)
SN = spr_f.pivot_table(index="date", columns="symbol", values="net", aggfunc="sum").reindex(SESSIONS).fillna(0)
SG = spr_f.pivot_table(index="date", columns="symbol", values="gross", aggfunc="sum").reindex(SESSIONS).fillna(0)
legs_of = spr_f.drop_duplicates("symbol").set_index("symbol")[["i1", "i2"]].reindex(SN.columns)
FLOW_JS = [j for j in STUDY_JS if SESSIONS[j - L_MAIN] >= fl.date.min()]
def buy_share(j, ids, crossing_only=False):
"""Buyer-initiated share of screen volume in contracts `ids` over the window ending at j.
Outright trades count directly. Buying the spread A-B buys A and sells B. With
crossing_only, spreads with both legs inside `ids` (which net to zero) are left out."""
ids = set(ids)
sl = slice(j - L_MAIN + 1, j + 1)
cols = [i for i in ids if i in ON.columns]
n, g = ON.iloc[sl][cols].sum().sum(), OG.iloc[sl][cols].sum().sum()
a1, a2 = legs_of.i1.isin(ids).to_numpy(), legs_of.i2.isin(ids).to_numpy()
sn, sg = SN.iloc[sl].sum().to_numpy(), SG.iloc[sl].sum().to_numpy()
both = a1 & a2
if crossing_only:
n += sn[a1 & ~a2].sum() - sn[a2 & ~a1].sum()
g += sg[a1 & ~a2].sum() + sg[a2 & ~a1].sum()
else:
n += sn[a1].sum() - sn[a2].sum()
g += sg[a1].sum() + sg[a2].sum()
return (0.5 + 0.5 * n / g if g > 0 else np.nan), n, g
blocked = set().union(*[set(range(pos[e] - L_MAIN, pos[e] + L_MAIN + 1)) for e in RES.end])
frows, FPL = [], {}
for r, rr in zip(RES.itertuples(), ROUNDS.itertuples()):
w = EV[(EV["round"] == r.round) & (EV.delivery >= FLOW_CUT)]
row = dict(round=r.round)
pls = []
for crossing in (False, True):
sa, na, ga = buy_share(pos[r.end], w[w.in_window].iid, crossing)
sb, _, _ = buy_share(pos[r.end], w[~w.in_window].iid, crossing)
pl = []
for j in FLOW_JS:
if j in blocked or SESSIONS[j] not in by_end_all:
continue
h = by_end_all[SESSIONS[j]]
h = h[meta.delivery.reindex(h.iid).to_numpy() >= FLOW_CUT]
mm = in_windows(meta.delivery.reindex(h.iid), rr.windows)
x, _, _ = buy_share(j, h.iid[mm], crossing)
pl.append(x)
pl = pd.Series(pl).dropna()
tag = "_crossing" if crossing else ""
row.update({f"screen_contracts{tag}": int(ga), f"net_bought{tag}": int(na), f"buyer_share{tag}": sa,
f"buyer_share_outside{tag}": sb, f"p{tag}": (1 + (pl >= sa).sum()) / (1 + len(pl))})
pls.append(pl)
row["placebo_windows"] = len(pls[0])
FPL[r.round] = pls
frows.append(row)
FRES = pd.DataFrame(frows)
FRES.round(3)
screen trades whose contracts all resolve: outrights 100.0%, spreads 97.8%; unsigned share of screen volume 16.4%
| round | screen_contracts | net_bought | buyer_share | buyer_share_outside | p | screen_contracts_crossing | net_bought_crossing | buyer_share_crossing | buyer_share_outside_crossing | p_crossing | placebo_windows | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | Round 1 | 1225717 | 3659 | 0.501 | 0.512 | 0.628 | 607687 | 3659 | 0.503 | 0.512 | 0.637 | 112 |
| 1 | Round 1.a | 588623 | 5715 | 0.505 | 0.489 | 0.566 | 303951 | 5715 | 0.509 | 0.482 | 0.540 | 112 |
| 2 | Round 1.b | 135828 | -2712 | 0.490 | 0.501 | 0.912 | 97940 | -2712 | 0.486 | 0.501 | 0.912 | 112 |
| 3 | Round 2 | 1259457 | -19957 | 0.492 | 0.468 | 0.903 | 538841 | -19957 | 0.481 | 0.467 | 0.912 | 112 |
| 4 | Round 3 | 862275 | -1877 | 0.499 | 0.500 | 0.726 | 395921 | -1877 | 0.498 | 0.499 | 0.752 | 112 |
The buyer-initiated share in each round's return-window contracts (blue) against its distribution on the placebo dates (grey), counting all screen volume (top) and leaving out spreads with both legs inside the windows (bottom).
fig, axes = plt.subplots(2, len(ROUNDS), figsize=(12, 4.6))
for k, r in enumerate(FRES.itertuples()):
for row, (tag, lab) in enumerate([("", "all screen volume"), ("_crossing", "spreads inside the windows left out")]):
ax = axes[row, k]
ax.hist(100 * FPL[r.round][row], bins=25, color=GREY)
ax.axvline(100 * getattr(r, "buyer_share" + tag), color=BLUE, linewidth=2)
ax.set_title((f"{r.round}\n" if row == 0 else "") +
f"{100 * getattr(r, 'buyer_share' + tag):.1f}% buyers, p = {getattr(r, 'p' + tag):.2f}", fontsize=10)
ax.set_yticks([]); ax.grid(axis="x", visible=False)
axes[0, 0].set_ylabel("all screen\nvolume")
axes[1, 0].set_ylabel("in-window spreads\nleft out")
fig.supxlabel("percent of screen volume in the return-window contracts (2027 and later) initiated by buyers; "
"grey: 2026 placebo dates", fontsize=10, color=MUTED)
plt.tight_layout(); plt.show()
10. Robustness¶
The main test was repeated with five other choices: the window ending on the award date instead of three trade dates after it; ending eight trade dates after it; treating the month after each return month as the hedge month (some physical crude for delivery in a month is priced off the following futures contract); leaving out the December and June contracts; and a six-day window starting on the bid deadline, with the detectors refitted on six-day history. The first table gives the autoencoder's count of buying anomalies inside the return windows and its placebo p-value, the second the same for the plain build score.
Across the 30 combinations for each score, the smallest autoencoder p-value is 0.08 (round 1, window ending on the award date, two anomalies). The smallest plain-score p-value is 0.07 (round 1.b, window ending eight trade dates after the award). Among 60 coarse p-values, a few that small are what chance alone produces, and no round is small in more than one variant.
variants = []
def summarise(res, label):
for r in res.itertuples():
variants.append(dict(variant=label, round=r.round, ae_flags=r.ae_flags_in, p_ae_flags=r.p_ae_flags_in,
iso_flags=r.iso_flags_in, p_iso_flags=r.p_iso_flags_in,
build_flags=r.build_flags_in, p_build_flags=r.p_build_flags_in,
p_build_gap=r.p_build_gap))
summarise(RES, "main: 10 trade dates ending award + 3")
summarise(evaluate(study, L_MAIN, offset=8)[0], "window ends award + 8")
summarise(evaluate(study, L_MAIN, offset=0)[0], "window ends on the award date")
summarise(evaluate(study, L_MAIN, shift=1)[0], "hedge month = return month + 1")
summarise(evaluate(study, L_MAIN, keep=lambda d: d.kind == "other")[0], "December and June contracts left out")
_r6, _s6, _n6 = build(6)
ref6, study6, _m6 = score(_r6, _s6)
summarise(evaluate(study6, 6, anchor="bid")[0], "6 trade dates from the bid deadline (detectors refitted)")
VAR = pd.DataFrame(variants)
VAR.assign(cell=lambda d: d.ae_flags.astype(str) + " (p " + d.p_ae_flags.round(2).astype(str) + ")") \
.pivot_table(index="variant", columns="round", values="cell", aggfunc="first", sort=False)
| round | Round 1 | Round 1.a | Round 1.b | Round 2 | Round 3 |
|---|---|---|---|---|---|
| variant | |||||
| main: 10 trade dates ending award + 3 | 1 (p 0.34) | 0 (p 1.0) | 1 (p 0.34) | 1 (p 0.67) | 0 (p 1.0) |
| window ends award + 8 | 1 (p 0.34) | 0 (p 1.0) | 2 (p 0.12) | 0 (p 1.0) | 0 (p 1.0) |
| window ends on the award date | 2 (p 0.08) | 0 (p 1.0) | 0 (p 1.0) | 1 (p 0.65) | 0 (p 1.0) |
| hedge month = return month + 1 | 2 (p 0.15) | 0 (p 1.0) | 2 (p 0.2) | 1 (p 0.63) | 0 (p 1.0) |
| December and June contracts left out | 0 (p 1.0) | 0 (p 1.0) | 1 (p 0.34) | 1 (p 0.67) | 0 (p 1.0) |
| 6 trade dates from the bid deadline (detectors refitted) | 1 (p 0.17) | 0 (p 1.0) | 0 (p 1.0) | 0 (p 1.0) | 0 (p 1.0) |
The plain build score under the same six variants.
VAR.assign(cell=lambda d: d.build_flags.astype(str) + " (p " + d.p_build_flags.round(2).astype(str) + ")") \
.pivot_table(index="variant", columns="round", values="cell", aggfunc="first", sort=False)
| round | Round 1 | Round 1.a | Round 1.b | Round 2 | Round 3 |
|---|---|---|---|---|---|
| variant | |||||
| main: 10 trade dates ending award + 3 | 1 (p 0.42) | 0 (p 1.0) | 2 (p 0.31) | 4 (p 0.29) | 0 (p 1.0) |
| window ends award + 8 | 1 (p 0.41) | 0 (p 1.0) | 4 (p 0.07) | 1 (p 0.85) | 0 (p 1.0) |
| window ends on the award date | 0 (p 1.0) | 0 (p 1.0) | 2 (p 0.29) | 4 (p 0.27) | 0 (p 1.0) |
| hedge month = return month + 1 | 2 (p 0.35) | 0 (p 1.0) | 3 (p 0.16) | 4 (p 0.27) | 1 (p 0.33) |
| December and June contracts left out | 1 (p 0.42) | 0 (p 1.0) | 2 (p 0.31) | 2 (p 0.66) | 0 (p 1.0) |
| 6 trade dates from the bid deadline (detectors refitted) | 2 (p 0.24) | 0 (p 1.0) | 0 (p 1.0) | 3 (p 0.44) | 1 (p 0.2) |
Conclusion¶
In the five 2026 strategic-reserve loan rounds awarded so far, the CL futures for the return months show no anomalous buying around the bid deadlines and award dates. The result is the same by two machine-learning detectors and by a plain score, each judged against ten years of contracts at the same distance from delivery and against the other dates of 2026. It holds for the direction of screen trading over the past year, and under five changes of window and month mapping. Round 3, which lent almost nothing, looks the same as the others.
The absence is weak evidence, and section 8 says how weak. It covers CL futures only (not Brent and not private swaps), contracts 6 to 42 months from delivery with at least 500 open contracts, buying concentrated in the ten trade dates ending three days after each award, and the measures defined in section 3. Within that box, the plain score would have caught a hedge of the full owed amount in round 2 on 93% of dates. It would have caught a quarter to a half of the owed amount, in rounds 1, 1.b and 2, on only 21 to 54% of dates. It would rarely have caught a hedge placed in December contracts, or anything in round 1.a. Outside the box the tests are blind: a hedge bought over weeks rather than days, a hedge in Brent or swaps, and no hedge at all look alike here. The one quantitative statement the data support is that the borrowers in round 2 did not hedge most of their owed barrels evenly across the return months during the ten days around the award. If they hedged, they did it more slowly, in other contracts, or elsewhere.
Two contracts remain worth watching, both in round 2's return windows: December 2028, which added 19,900 contracts starting on the bid deadline, and June 2029, whose open interest doubled in the three weeks after it while March 2029 and December 2029 stayed flat. Neither is unusual enough for 2026 to count as evidence on its own.
What would change this. Round 4 closes for bids on 6 October 2026. Its return windows include thin months (October 2028 to June 2029, and August to December 2029) alongside April to September 2027. Rerunning the cache builder and this notebook about two weeks after its award is a clean forward test. A position built in the thin 2029 months would stand out more than in rounds 1 and 2, whose windows covered almost the whole liquid curve.
Rebuilding¶
From the repository root, in soroban's environment, run soroban/.venv/bin/python soroban/notebooks/build_spr_loan_caches.py to refresh the caches. It re-pulls only the current year of statistics and re-reads the Energy Department's postings. Every Databento call is cost-checked, and all are inside the CME bundle at $0. Then re-execute this notebook. A new round needs one new row in the rounds table in section 1, with its dates, volume, premium and return windows taken from its request for proposal and award notice (both saved under data/databento/spr_loans/doe/).