Southeast Colorado alfalfa price forecasts, graded in public.

The Lessons Ledger

Every rule below was bought with a graded miss. They are binding on future reports: a forecast that violates one must say so explicitly and justify it. Newest first. This page only grows.

These rules are enforced in code, not memory (added 2026-08-18, same day as L4–L7): the graded record lives in a machine-readable ledger (prediction_grades.csv); a calibration engine (calibration.py) recomputes coverage and derives mandatory minimum range widths from measured error every weekly run; a survey-anchored estimator (production_estimate.py) fails the pipeline if the operative production estimate leaves the print-±-max-revision window; and the demand indicators we under-weighted (7-state drought index, weekly pasture condition, a cash-price benchmark ledger) are now collected as first-class data series. The test suite fails the build on violations. Current calibration floors, straight from the misses: price predictions ±$65, production ±30%, volume ±57% — until measured coverage earns narrower.


L10 — Indicators must beat the same release-aware baseline before they move the midpoint. (2026-08-28, whole-project retro)

The miss: rich water, weather, drought, ENSO, fuel, and production narratives were repeatedly converted into precise price adjustments without proving that those indicators improved an out-of-sample price forecast. The full model added only ~0.004 in-sample R² beyond price persistence and seasonality, while its live forecasts were low by $34.1/ton on average. The rule: a historical indicator needs at least 48 out-of-sample observations, at least $2/ton MAE improvement over the same release-aware baseline, and improvement in at least four of five evaluation years before it may alter the numeric midpoint. Failed indicators remain scenario or monitoring variables. Comparable AMS cash may anchor the current-market nowcast because it observes the target directly; it does not receive unearned long-horizon predictive status.

L9 — Confidence is earned by comparable forecast outcomes, not by research thoroughness. (2026-08-28, whole-project retro)

The miss: reports labeled near-term prices HIGH confidence because supply facts were well researched, while the local price mapping had no comparable track record. The first frozen scorecard contained only 2 of 10 two-sided ranges; the live model put 0 of 3 outcomes inside its bands. The rule: confidence comes only from resolved directly comparable sample size, point error, bias, and interval coverage. Until there are at least 20 comparable local outcomes, 70–90% interval coverage across at least two cycles, and no material same-signed bias, the maximum forecast confidence is LOW. Source count, citations, in-sample R², narrative agreement, and agent certainty cannot upgrade it.

L8 — Scores are computed, not typed; metrics are specified at freeze time, not grade time. (2026-08-18, the erratum)

The miss: the same day we published our first pre-registered scorecard, we published it wrong. A grader scored "weekly Niño-3.4" against the monthly index (a different number that happened to sit inside our range) and hand-typed "3 of 10" into four documents; the true score was 2 of 10. Corrected in public within hours — but the error class is structural: any number that travels from a ledger to prose by human hands can drift, and any prediction whose metric lives in prose can be graded against the wrong thing. The rule: every frozen prediction ships with a machine-readable source spec (exact release, table, row, column — prediction_sources.json); grading is done by script (grade_predictions.py) against that spec; every published score line is generated from the ledger (scorecard.py). Tests fail the build if a pending row lacks a spec, if a ledger outcome disagrees with the frozen scoring math, or if a published "X of Y" claim disagrees with the ledger.

L7 — Calibrate interval width to measured coverage, not to confidence. (2026-08-18)

The miss: our pre-registered ranges contained the outcome 2 times out of 10 (initially misreported as 3/10 — one row was graded against the wrong metric, corrected same day). Ranges that miss 80% of the time are roughly half as wide as honesty requires. The rule: prediction ranges are set from measured error (the vintage backtest: ±$25–30 on winter months; the release-grading record), not from how sure the narrative feels. If the last cycle's coverage was below ~70%, the next cycle's ranges widen mechanically — no discretion.

L6 — When misses share a sign across cycles, it's bias, not bad luck. (2026-08-18)

The miss: three consecutive reports where every price surprise broke bullish (June $225 contract → July cash rally → August $275 delivered / $350–380 stable lots). Each time we treated it as a one-off. The rule: two consecutive same-signed misses on a series triggers a named bias review in the next report. Current standing bias: we under-weight demand. Demand-side indicators (pasture condition %, neighbor-state DSCI, feeder cattle margins, KS–CO spread) are now first-class forecast inputs, not color.

L5 — Drought indices are demand proxies for irrigated crops, not supply proxies. (2026-08-18)

The miss: we read "worst water year on record" as "the alfalfa crop fails." USDA's survey says CO alfalfa −4.8% while winter wheat went −67% and corn −24%. Drought kills dryland; wells and senior rights keep irrigated alfalfa alive — what drought actually does to the alfalfa price is destroy pasture and grass hay, which turns every ruminant into an alfalfa buyer. The rule: supply estimates for irrigated crops must be built from water deliveries by region weighted by production share, with explicit credit for groundwater. DSCI/USDM feed the demand model.

L4 — News coverage samples the disaster, not the state. (2026-08-18 — the big one)

The miss: our 1.5M-ton (−36%) production estimate was assembled from canal shutdowns, fallowing headlines, and worst-on-record stories in the regions where the story was dramatic. USDA's farmer survey printed 2,244k (−4.8%). For our number to be right, the quiet 57% of the state would have had to fail too — and nobody writes articles about fields that got water. The rule: statewide estimates anchor on the survey base rate and the maximum historical revision. The plausible-final window is [Aug print − 12%, Aug print] (2022 is the record revision). An estimate outside that window requires extraordinary, quantified evidence — not vibes, not clippings. Our own estimate is now graded against this: 1.5M (frozen) vs ~1.95M (rev1), settled Jan 12, 2027.

L3 — Load-bearing negative claims need two independent verification methods. (2026-08-08)

The miss: a broken shell tool returned empty search results and we published "USDA dropped hay from the August report." The files contained the tables all along. Retracted same day. The rule: any "X is absent / X stopped happening" claim ships only after two independent checks (different tool, different method).

L2 — Measure markets from cash prints, never from lagging state averages. (2026-08-08)

The miss: July's "Kansas caps SE CO at ~$155–170 landed" used a stale NASS state average ($123) while Kansas cash grinding hay traded $200–210. The cap was off by ~$60–65 landed. The rule: floors and ceilings are computed from AMS cash trades plus current freight (the market_parity.py table), refreshed every cycle. NASS monthly averages are trend confirmation only — they lag 4–8 weeks.

L1 — Don't forecast fuel. Mark it to market. (2026-07-11)

The miss: two consecutive directional misses on diesel (April: missed the fall; July: called it "falling" at $4.58, it went to $5.35). The rule: every report uses the current posted EIA weekly price. No diesel opinions.


Why this page exists: on 2026-08-18 the project's pre-registered predictions were graded against the August USDA releases and scored 2/10 on intervals (initially published as 3/10 — one ENSO row was graded against the monthly index when the frozen prediction specified the weekly one; corrected the same day, in public), with the central production thesis missing by 28%. The full grading record is here; the reports that own each miss are in the reports archive. We keep the misses public because the alternative — quietly getting better without admitting what was wrong — is how forecasting products lie.