Gold Standard — Kalshi event-market forecaster

A pre-registered, paper-only test of whether an LLM can out-forecast the market's own price. Nothing here places real orders.

Headline scorecard

Resolved
1080
of 150 gate threshold
Brier (ours)
0.227
Brier (market)
0.135
Edge
-0.092
positive = we beat the market
BOTH gates FAIL -- kill, write the post-mortem, keep the track record

By category

CategorynBrier oursBrier mktedgehit
Climate and Weather3790.2560.134-0.12160%
Commodities2930.2150.183-0.03364%
Economics2000.2220.098-0.12469%
Financials650.2110.197-0.01460%
Politics1060.1930.060-0.13472%
Science and Technology350.1830.061-0.12266%
World20.0410.000-0.041100%

By context kind

Context kindnBrier oursBrier mktedgehit
fred870.1940.121-0.07471%
news1720.1950.079-0.11672%
none1800.2370.130-0.10759%
nws3520.2550.130-0.12561%
prices2890.2140.180-0.03564%

Calibration

BandnSaid %Actual %
0.5-0.620853%48%
0.6-0.722063%58%
0.7-0.817974%65%
0.8-0.922184%63%
0.9-1.025294%84%

Paper P&L

n_trades=551   total_pnl=$-2861.17   total_fees=$1271.51   win_rate=26%

Open forecasts

TickerCategoryOur pMkt pClose timeRationale
KXWTI-26SEP2814-T94.49Commodities0.780.742026-09-28T18:30:00ZLast close 95.52 above 94.49; hours to settlement limit further drift.
KXWTI-26SEP2814-T93.99Commodities0.850.842026-09-28T18:30:00ZLast close 95.52 comfortably above 93.99 threshold near settlement.
KXTEMPMIAH-26SEP2806-T72.99Climate and Weather0.930.892026-09-28T10:00:00ZOvernight forecast low 76F, well above 73F threshold for 6am reading.
KXTEMPMIAH-26SEP2806-T73.99Climate and Weather0.800.032026-09-28T10:00:00ZOvernight forecast low 76F gives some margin above 74F threshold.
KXDEEPSHARE-26SEP28-25.2Science and Technology0.500.022026-09-28T03:59:00ZDeepSeek's share plausibly near this level given past dominance in cheap inference.
KXBABASHARE-26SEP28-6.2Science and Technology0.500.892026-09-28T03:59:00ZQwen's OpenRouter share plausibly near this level.
KXBABASHARE-26SEP28-6.5Science and Technology0.450.392026-09-28T03:59:00ZQwen's OpenRouter share plausibly near this level.
KXTOPUSAGEAI-26SEP21-ANTHScience and Technology0.040.012026-09-28T03:59:00ZAnthropic rarely leads OpenRouter weekly usage vs DeepSeek/Google.
KXXIAOMISHARE-26SEP28-1.9Science and Technology0.500.982026-09-28T03:59:00ZXiaomi is a newer OpenRouter entrant; share plausibly near this level.
KXXIAOMISHARE-26SEP28-2.5Science and Technology0.280.352026-09-28T03:59:00ZXiaomi's share reaching 2.5%+ is less likely for a newer entrant.
KXTENCENTSHARE-26SEP28-4.9Science and Technology0.250.012026-09-28T03:59:00ZTencent reaching near 5% OpenRouter share is a higher bar.
KXOPENSHARE-26SEP28-19.3Science and Technology0.320.012026-09-28T03:59:00ZOpenAI exceeding ~19% OpenRouter share is less typical.
KXGOOGSHARE-26SEP28-18.3Science and Technology0.350.992026-09-28T03:59:00ZGoogle's share growing but 18.3% is on the higher end of typical range.
KXGOOGSHARE-26SEP28-16.7Science and Technology0.550.992026-09-28T03:59:00ZGoogle's OpenRouter share plausibly near mid-teens to high-teens.
KXDEEPSHARE-26SEP28-28.1Science and Technology0.350.012026-09-28T03:59:00ZDeepSeek historically a top OpenRouter provider but 28%+ is a high bar.
KXANTHSHARE-26SEP28-2.5Science and Technology0.550.992026-09-28T03:59:00ZAnthropic's OpenRouter share is typically low single digits.
KXANTHSHARE-26SEP28-3.1Science and Technology0.300.012026-09-28T03:59:00ZAnthropic's OpenRouter share rarely exceeds low-3% range.
KXTENCENTSHARE-26SEP28-3.5Science and Technology0.500.012026-09-28T03:59:00ZTencent's OpenRouter share plausibly near this level; no exact figure available.
KXOPENSHARE-26SEP28-14.0Science and Technology0.650.992026-09-28T03:59:00ZOpenAI historically holds a mid-teens OpenRouter share.

Recent resolutions

TickerOur pMkt pOutcome
KXIPHONERELEASE-IPHONE18PRO-26SEP280.960.50Yes
KXTRUMPAPPROVE-26SEP27-E38.60.050.24No
KXTRUMPACT-26SEP20-T40.700.01No
KXTRUMPACT-26SEP20-T100.300.01No
KXTRUMPNOMNUM-26SEP20-T10.850.01No
KXTRUMPAPPROVE-26SEP27-E38.70.050.64Yes
KXTEMPMIAH-26SEP2706-T73.990.920.50No
KXTEMPMIAH-26SEP2706-T74.990.600.03No
KXTEMPCHIHS-26SEP2706-T48.990.950.14No
KXTRUMPNOMNUM-26SEP20-T100.350.01No
KXLOWTAUS-26SEP26-T660.010.02No
KXHIGHTSATX-26SEP26-B94.50.280.59Yes
KXLOWTMIN-26SEP26-T580.150.01No
KXLOWTDAL-26SEP26-T730.400.01No
KXLOWTCHI-26SEP26-T550.250.01No
KXLOWTAUS-26SEP26-B66.50.020.03No
KXHIGHNY-26SEP26-T620.350.74Yes
KXHIGHMIA-26SEP26-B89.50.150.54No
KXHIGHMIA-26SEP26-B87.50.400.24Yes
KXLOWTOKC-26SEP26-T700.450.01No

Pre-registration (locked before any forecast)

# PREREG.md — pre-registration of the gold-standard Kalshi forecaster

**Registered 2026-07-28 (America/Chicago), before any `source='ai'` forecast row exists.**
The git history is the proof: this file is committed at Phase 5, and the first production
(`source='ai'`) forecast is not created until Phase 7, after this commit. Any change to the rules
below requires a dated amendment appended to the "Amendments" section AND the owner's sign-off — the
original rules above the amendment line are never edited in place.

This is a paper-only research instrument. Nothing here places real orders. Educational / hypothetical.

---

## The pre-registered rules (copied verbatim from PLAN.md "LOCKED decisions")

1. **The model never sees the market price** in its prompt. We are testing independent skill;
   showing the price lets the model echo the crowd and fakes calibration. Log `mkt_p` separately
   at call time. (A later "sees-price" variant may be added as a separate scored segment. Not v1.)
2. **Question selection:** status open, `close_time` ≤ 30 days out, `volume_fp` ≥ 500, bid-ask
   spread ≤ $0.10, category in: Economics, Financials, Commodities, Climate and Weather, Politics,
   Science and Technology, World. Excluded: Sports, Entertainment, Mentions, Crypto, Elections,
   Companies, Social, Health. Max **2 markets per event** (for ladders pick the 2 strikes nearest
   $0.50 mid = nearest the money). Max **40 new forecasts per night**.
3. **One forecast per market**, made at first qualifying sighting, never revised. The model may
   **skip** any question (no informational basis); skips are logged and cost nothing.
4. **Paper trade rule:** fills at the **ask** (not mid). Buy YES if `our_p − yes_ask − fee(yes_ask)
   > 0.02`; buy NO if `(1 − our_p) − no_ask − fee(no_ask) > 0.02`. Kalshi fee per $1 contract:
   `fee(p) = ceil_to_cent(0.07 · p · (1 − p))`. Flat $10 notional per trade, max 1 trade per event.
5. **Scoring:** Brier of `our_p` vs Brier of `mkt_p` (mid at call time) on identical resolved sets,
   segmented by category and by context kind (fred / nws / prices / news / none), plus calibration
   buckets. Voided/unresolvable markets are excluded. Only rows with `source='ai'` count
   (`source='dev'` = build-time tests, never scored).
6. **Gates (evaluate at ≥150 resolved AI forecasts):** skill gate = our Brier < market Brier;
   money gate = paper P&L > 0 net of fees over ≥100 trades. Both pass → owner may fund a real
   bankroll ($500–2k, allowed to go to zero; separate authorization, separate build). Either fails
   → kill, write the post-mortem, keep the track record. Interim peeks fine; no action before n.
7. **Prompt discipline:** `PROMPT_VER` constant logged on every forecast; any prompt change bumps
   it. LOUD failures only: a failed/malformed LLM call writes a `runs` row and stderr; **never**
   silently fall back (market-radar's silent-fallback bug is the cautionary tale).

---

## Exact prompt text (PROMPT_VER = "v1")

The nightly batched call (`brain.py`, `--model sonnet`) emits ONE prompt built from a fixed
preamble plus one block per question. The preamble is **verbatim**:

> You are a calibrated event forecaster. For each question output a probability 0-1 that it
> resolves YES. You have NO market price and must not guess one - reason only from the evidence
> given. If you have no informational basis, set skip=true (this costs nothing and is better than
> a blind guess). Today is {UTC date} (UTC). Output STRICT JSON ONLY: an object mapping each
> ticker to {"p":<0..1>,"why":"<=20 words","skip":<bool>}. No prose outside the JSON.

Each per-question block that follows the preamble contains ONLY: the ticker, the question title,
its sub-title, its resolution rules, and the Phase-2 context pack (external evidence — FRED / NWS /
prices / news — which is itself price-free by construction; `context.py` contains zero Kalshi price
fields). **No market price, bid, ask, or implied probability is ever placed in the prompt.** This is
enforced structurally: `build_prompt()` is a separate pure function and the price fetch lives in
`_fetch_mid()`; a grep proves the string `_dollars` never appears in the prompt path.

If v1 is ever changed, `PROMPT_VER` is bumped (e.g. "v2") and forecasts made under different
versions are scored as separate segments — v1's track record is never retroactively altered.

---

## Gate arithmetic (exactly what `score.py` computes)

Scored set = resolved `forecasts` rows with `source='ai'`, `outcome ∈ {0,1}` (void `-1` and
unresolved `NULL` excluded), and both `our_p` and `mkt_p` populated.

- `brier_ours = mean((our_p − outcome)²)` over the scored set.
- `brier_mkt  = mean((mkt_p − outcome)²)` over the identical set (the crowd's mid at call time).
- `edge = brier_mkt − brier_ours` (positive ⇒ we beat the market; lower Brier is better).
- Paper P&L: over settled `paper_trades` (`pnl` not NULL), `total_pnl = Σ pnl`, `n_trades = count`.
  Per-trade P&L (LOCKED #4 `$1`-contract math, `contracts = $10 / entry_p`,
  `total_fee = contracts · fee(entry_p)`): WIN `pnl = 10·(1/entry_p − 1) − total_fee`;
  LOSS `pnl = −10 − total_fee`; VOID `pnl = 0`.

**Decision (only at `n_resolved ≥ 150`):**
- **skill_gate_pass** ⟺ `n_resolved ≥ 150` AND `brier_ours < brier_mkt`.
- **money_gate_pass** ⟺ `n_trades ≥ 100` AND `total_pnl > 0`.
- **Both pass** → the forecaster has earned a real-bankroll funding decision ($500–2k, may go to
  zero; separate authorization + separate build). **Either fails** → kill it, write the post-mortem,
  keep the documented track record as the durable asset.
- Before `n_resolved = 150`: interim peeks are allowed but **no action is taken** — `score.py`'s
  verdict reads "accruing — N of 150 resolved; no decision yet".

---

## Note on the fee example in PLAN.md

PLAN.md's Phase-4 acceptance line ("fee(0.50) on 20 contracts = $0.70") is an arithmetic slip that
contradicts the LOCKED 0.07 coefficient. The **coefficient is unchanged** (`fee(p) =
ceil_to_cent(0.07·p·(1−p))`, i.e. $0.02/contract at p=0.50). The implemented order fee is
`contracts × fee(entry_p)` = 20 × $0.02 = **$0.40** at the example point. This clarification does not
alter any pre-registered rule; it records the correct arithmetic under the LOCKED formula.

---

## Amendments

(none — the rules above are the original registration. Append dated, owner-signed amendments below
this line; never edit the original rules in place.)