Gold Standard — Kalshi event-market forecaster

A pre-registered, paper-only test of whether an LLM can out-forecast the market's own price. Nothing here places real orders.

Headline scorecard

Resolved
1016
of 150 gate threshold
Brier (ours)
0.230
Brier (market)
0.138
Edge
-0.091
positive = we beat the market
BOTH gates FAIL -- kill, write the post-mortem, keep the track record

By category

CategorynBrier oursBrier mktedgehit
Climate and Weather3420.2620.142-0.11959%
Commodities2830.2190.183-0.03663%
Economics1920.2290.102-0.12868%
Financials640.2150.196-0.01859%
Politics980.1800.060-0.12073%
Science and Technology350.1830.061-0.12266%
World20.0410.000-0.041100%

By context kind

Context kindnBrier oursBrier mktedgehit
fred850.1970.123-0.07471%
news1670.2000.080-0.12071%
none1660.2390.139-0.10158%
nws3190.2590.136-0.12361%
prices2790.2180.180-0.03864%

Calibration

BandnSaid %Actual %
0.5-0.620153%47%
0.6-0.720663%57%
0.7-0.817074%66%
0.8-0.921284%62%
0.9-1.022794%84%

Paper P&L

n_trades=521   total_pnl=$-2492.30   total_fees=$1149.47   win_rate=27%

Open forecasts

TickerCategoryOur pMkt pClose timeRationale
KXGOLDD-26SEP2417-T4258Commodities0.820.482026-09-24T21:00:00ZGold closed 4290, about 32pts above threshold; likely holds by 5pm given moderate vol.
KXGOLDD-26SEP2417-T4268Commodities0.740.392026-09-24T21:00:00ZGold closed near 4290, only ~22pts buffer above 4268 threshold.
KXBRENTD-26SEP2417-T100Commodities0.400.492026-09-24T21:00:00ZWTI at 93.78; typical Brent premium of $5-6 puts Brent near threshold, uncertain.
KXBRENTD-26SEP2417-T100.50Commodities0.350.402026-09-24T21:00:00ZEstimated Brent near 99-100 from WTI plus typical spread; slightly below this threshold.
KX30YMORTW-26SEP24-T6.98Economics0.550.892026-09-24T15:59:00ZRates described as nearing 7%, plausibly just above or below 6.98%.
KXWTI-26SEP2414-T90.99Commodities0.920.942026-09-24T18:30:00ZWTI closed 93.80 on Sep24, comfortably above 90.99 settlement threshold.
KXWTI-26SEP2414-T88.99Commodities0.960.982026-09-24T18:30:00ZWTI closed 93.80 on Sep24, well above 88.99 settlement threshold.
KXCBDECISIONMEXICO-26SEP24-H25PEconomics0.050.122026-09-24T18:59:00ZNo Banxico-specific signal; hikes are uncommon in current global easing environment.
KXCBDECISIONMEXICO-26SEP24-C50PEconomics0.050.032026-09-24T18:59:00ZLarge 50bp+ cuts are historically rare single-meeting moves without crisis signal.
KXTEMPMIAH-26SEP2406-T70.99Climate and Weather0.970.992026-09-24T10:00:00ZOvernight low forecast 79F, far above 71F threshold at 6am.
KXTEMPCHIHS-26SEP2406-T58.99Climate and Weather0.620.902026-09-24T10:00:00ZOvernight low forecast 60F, just above 59F threshold; some downside risk.
KXJOBLESSCLAIMS-26SEP24-215000Economics0.080.082026-09-24T12:25:00ZClaims trending down, last reading 196k, well below 215k threshold.
KXJOBLESSCLAIMS-26SEP24-205000Economics0.320.172026-09-24T12:25:00ZRecent claims 196k-212k, downtrend makes reaching 205k less likely than not.
KX30YMORTW-26SEP24-T7.08Economics0.150.012026-09-24T15:59:00ZHeadlines say mortgage rates are 'nearing 7%', suggesting current rate below 7.08%.

Recent resolutions

TickerOur pMkt pOutcome
KXBIGGESTQUAKE-23SEP26-5.20.800.98Yes
KXGOLDD-26SEP2317-T43150.850.46No
KXGOLDD-26SEP2317-T43250.800.35No
KXBRENTD-26SEP2317-T970.350.32Yes
KXBRENTD-26SEP2317-T97.500.280.23Yes
KXTEMPMIAH-26SEP2306-T71.990.970.96Yes
KXTEMPLAXHS-26SEP2306-T67.990.250.04No
KXWTI-26SEP2314-T90.990.300.38Yes
KXWTI-26SEP2314-T91.990.170.20Yes
KXTEMPMIAH-26SEP2306-T74.990.950.43No
KXSILVERD-26SEP2217-T65.250.750.57Yes
KXSILVERD-26SEP2217-T65.750.480.34Yes
KXGOLDD-26SEP2217-T43150.880.57Yes
KXGOLDD-26SEP2217-T43250.780.46Yes
KXTEMPMIAH-26SEP2206-T77.990.780.08No
KXTEMPCHIHS-26SEP2206-T56.990.550.36No
KXGOLD15M-26SEP220600-000.500.69Yes
KXWTI-26SEP2214-T90.990.480.41No
KXWTI-26SEP2214-T90.490.850.54Yes
KXWTI15M-26SEP220600-000.500.15No

Pre-registration (locked before any forecast)

# PREREG.md — pre-registration of the gold-standard Kalshi forecaster

**Registered 2026-07-28 (America/Chicago), before any `source='ai'` forecast row exists.**
The git history is the proof: this file is committed at Phase 5, and the first production
(`source='ai'`) forecast is not created until Phase 7, after this commit. Any change to the rules
below requires a dated amendment appended to the "Amendments" section AND the owner's sign-off — the
original rules above the amendment line are never edited in place.

This is a paper-only research instrument. Nothing here places real orders. Educational / hypothetical.

---

## The pre-registered rules (copied verbatim from PLAN.md "LOCKED decisions")

1. **The model never sees the market price** in its prompt. We are testing independent skill;
   showing the price lets the model echo the crowd and fakes calibration. Log `mkt_p` separately
   at call time. (A later "sees-price" variant may be added as a separate scored segment. Not v1.)
2. **Question selection:** status open, `close_time` ≤ 30 days out, `volume_fp` ≥ 500, bid-ask
   spread ≤ $0.10, category in: Economics, Financials, Commodities, Climate and Weather, Politics,
   Science and Technology, World. Excluded: Sports, Entertainment, Mentions, Crypto, Elections,
   Companies, Social, Health. Max **2 markets per event** (for ladders pick the 2 strikes nearest
   $0.50 mid = nearest the money). Max **40 new forecasts per night**.
3. **One forecast per market**, made at first qualifying sighting, never revised. The model may
   **skip** any question (no informational basis); skips are logged and cost nothing.
4. **Paper trade rule:** fills at the **ask** (not mid). Buy YES if `our_p − yes_ask − fee(yes_ask)
   > 0.02`; buy NO if `(1 − our_p) − no_ask − fee(no_ask) > 0.02`. Kalshi fee per $1 contract:
   `fee(p) = ceil_to_cent(0.07 · p · (1 − p))`. Flat $10 notional per trade, max 1 trade per event.
5. **Scoring:** Brier of `our_p` vs Brier of `mkt_p` (mid at call time) on identical resolved sets,
   segmented by category and by context kind (fred / nws / prices / news / none), plus calibration
   buckets. Voided/unresolvable markets are excluded. Only rows with `source='ai'` count
   (`source='dev'` = build-time tests, never scored).
6. **Gates (evaluate at ≥150 resolved AI forecasts):** skill gate = our Brier < market Brier;
   money gate = paper P&L > 0 net of fees over ≥100 trades. Both pass → owner may fund a real
   bankroll ($500–2k, allowed to go to zero; separate authorization, separate build). Either fails
   → kill, write the post-mortem, keep the track record. Interim peeks fine; no action before n.
7. **Prompt discipline:** `PROMPT_VER` constant logged on every forecast; any prompt change bumps
   it. LOUD failures only: a failed/malformed LLM call writes a `runs` row and stderr; **never**
   silently fall back (market-radar's silent-fallback bug is the cautionary tale).

---

## Exact prompt text (PROMPT_VER = "v1")

The nightly batched call (`brain.py`, `--model sonnet`) emits ONE prompt built from a fixed
preamble plus one block per question. The preamble is **verbatim**:

> You are a calibrated event forecaster. For each question output a probability 0-1 that it
> resolves YES. You have NO market price and must not guess one - reason only from the evidence
> given. If you have no informational basis, set skip=true (this costs nothing and is better than
> a blind guess). Today is {UTC date} (UTC). Output STRICT JSON ONLY: an object mapping each
> ticker to {"p":<0..1>,"why":"<=20 words","skip":<bool>}. No prose outside the JSON.

Each per-question block that follows the preamble contains ONLY: the ticker, the question title,
its sub-title, its resolution rules, and the Phase-2 context pack (external evidence — FRED / NWS /
prices / news — which is itself price-free by construction; `context.py` contains zero Kalshi price
fields). **No market price, bid, ask, or implied probability is ever placed in the prompt.** This is
enforced structurally: `build_prompt()` is a separate pure function and the price fetch lives in
`_fetch_mid()`; a grep proves the string `_dollars` never appears in the prompt path.

If v1 is ever changed, `PROMPT_VER` is bumped (e.g. "v2") and forecasts made under different
versions are scored as separate segments — v1's track record is never retroactively altered.

---

## Gate arithmetic (exactly what `score.py` computes)

Scored set = resolved `forecasts` rows with `source='ai'`, `outcome ∈ {0,1}` (void `-1` and
unresolved `NULL` excluded), and both `our_p` and `mkt_p` populated.

- `brier_ours = mean((our_p − outcome)²)` over the scored set.
- `brier_mkt  = mean((mkt_p − outcome)²)` over the identical set (the crowd's mid at call time).
- `edge = brier_mkt − brier_ours` (positive ⇒ we beat the market; lower Brier is better).
- Paper P&L: over settled `paper_trades` (`pnl` not NULL), `total_pnl = Σ pnl`, `n_trades = count`.
  Per-trade P&L (LOCKED #4 `$1`-contract math, `contracts = $10 / entry_p`,
  `total_fee = contracts · fee(entry_p)`): WIN `pnl = 10·(1/entry_p − 1) − total_fee`;
  LOSS `pnl = −10 − total_fee`; VOID `pnl = 0`.

**Decision (only at `n_resolved ≥ 150`):**
- **skill_gate_pass** ⟺ `n_resolved ≥ 150` AND `brier_ours < brier_mkt`.
- **money_gate_pass** ⟺ `n_trades ≥ 100` AND `total_pnl > 0`.
- **Both pass** → the forecaster has earned a real-bankroll funding decision ($500–2k, may go to
  zero; separate authorization + separate build). **Either fails** → kill it, write the post-mortem,
  keep the documented track record as the durable asset.
- Before `n_resolved = 150`: interim peeks are allowed but **no action is taken** — `score.py`'s
  verdict reads "accruing — N of 150 resolved; no decision yet".

---

## Note on the fee example in PLAN.md

PLAN.md's Phase-4 acceptance line ("fee(0.50) on 20 contracts = $0.70") is an arithmetic slip that
contradicts the LOCKED 0.07 coefficient. The **coefficient is unchanged** (`fee(p) =
ceil_to_cent(0.07·p·(1−p))`, i.e. $0.02/contract at p=0.50). The implemented order fee is
`contracts × fee(entry_p)` = 20 × $0.02 = **$0.40** at the example point. This clarification does not
alter any pre-registered rule; it records the correct arithmetic under the LOCKED formula.

---

## Amendments

(none — the rules above are the original registration. Append dated, owner-signed amendments below
this line; never edit the original rules in place.)