Everything behind the number, written down.
Lighthouse makes one claim: that it can estimate where a tokenized US equity will reopen better than the quote Bitget is showing while the US market is shut. This page is the whole method, the whole validation, and the places it fails — so the claim can be argued with rather than believed.
What this is
Bitget lists rTokens: tokenized US equities backed 1:1 through Reality Protocol, with roughly 1,650 symbols live. During US trading hours an order routes through to NYSE or Nasdaq and the price you see is a real transaction price. Outside those hours it is not. Bitget says so plainly in its own documentation: the price shown is an indicative quote from market makers, not a transaction price.
That distinction is doing a lot of work, because those quotes are not decorative. They mark your collateral in the unified account. They set the number you see when you decide whether to hold through the night. Lighthouse exists to answer one question about them: how much of tonight's quote will still be true when the market actually reopens?
The dark window
The obvious framing is wrong, and getting it wrong roughly triples every error figure on this site. US equities do not stop trading at 16:00 ET. Extended hours run 04:00 to 20:00 ET, so the stretch from the closing bell to the next morning is mostly still quoted by real US venues.
The genuinely dark window — no US venue of any kind quoting, only Bitget's market makers — is 20:00 to 04:00 ET. That is the eight-hour span this desk models. The 19:00 ET hourly bar closes at 20:00 and gives us the anchor; the 04:00 ET bar is the reopen we grade against. Friday 20:00 to Monday 04:00 is folded into one long weekend window, which behaves very differently.
How fair value is built
Every rToken's overnight move splits into two parts. Some of it is the whole tape moving together — a futures drift, a macro print, general risk appetite. The rest is specific to the name. We separate them, because they survive to the reopen at very different rates.
- 01The common move
Take the median return across every name on the board since its 20:00 close. The median, not the mean, so a single stub quote cannot drag the tape.
- 02Each name's share of it
Multiply by that name's beta, fitted on closed windows rather than assumed to be 1.0.
- 03What's left is idiosyncratic
The quote's move minus its share of the common move.
- 04Shrink both, separately
Each component is multiplied by its own weight — around 0.85 on the common move and 0.70 on the name-specific one at midnight. A weight below 1.0 means the quote is overshooting and we pull it back toward the close.
- 05Fair value
The 20:00 close, moved by the shrunk components. Every figure on the board is that number against the live quote.
The weights are not constant across the night. They are fitted per elapsed bucket, because a quote two hours into the window and a quote ten minutes before the reopen deserve very different amounts of trust. Fitting minimises median absolute error by coordinate descent, not squared error by least squares — overnight returns have fat tails, and squared error would let a handful of gap nights dictate the weights for every ordinary one.
How it was validated
The single guarantee this project rests on: no forecast is ever scored by a model that was allowed to see it. For each window we refit on the trailing 40 windows that had already closed, then score the next one. Walk-forward, never in-sample. There is a test in the repository whose only job is to assert that the window being scored never appears in its own training set.
Across 20,934 graded forecasts on — scored windows, Lighthouse beats the venue's own quote on roughly 54 to 56% of individual forecasts. That is a real edge and a modest one, and we would rather publish it at that size than dress it up.
What we got wrong
Three findings contradicted what we expected going in. All three are in the ledger, and none were removed for being inconvenient.
- 01The venue's quote starts out worse than doing nothing
Early in the window the quote is worse than assuming the close held.
- 02The first version of the model learned to do nothing
Fitted on a fixed train/test split it converged on weights of almost exactly 1.0 — it had learned the identity function. The cause was assuming stationarity across months of data. Rolling recalibration fixed it, and the failure is why the walk-forward test exists.
- 03A seventh of the universe has a stale book
Thirteen of sixty names showed identical prices at the close and the reopen, which is not a forecast being right, it is nobody trading. They are excluded by a minimum-realised-move filter and named in the calibration file.
The collateral arithmetic
rTokens are accepted as collateral in Bitget's unified account at up to 95%. Between 20:00 and 04:00 ET that collateral is marked at the indicative quote — so the margin ratio and liquidation distance the venue shows you are computed on a price nobody traded on.
The collateral page prices the gap. Collateral is the mark times the haircut; liquidation is where collateral stops covering the debt. We compute your distance to that point twice — once at the venue's mark, once at fair value — and then read the probability of crossing it off the measured error tails for the names you actually hold, blended by position size.
The analyst
The research analyst is an OpenAI-compatible model with no browser, no memory and no prices of its own. It receives one message: the desk state — tonight's full board, each name's historical accuracy and overshoot, the shrinkage weights in force, and the walk-forward error table. That context is published verbatim next to every answer.
It answers in four parts — a read, the reasoning, the figures it leaned on, and what would change its mind. Then every figure it cites is matched back against the context before rendering. Anything the model could not have read there is stripped and the count of what was stripped is shown. It is instructed never to say buy or sell, and to say plainly when the context does not contain what the question needs.
Data and reproduction
Everything comes from Bitget's public market endpoints — tickers and hourly candles. No API key, no account, no signing. The history endpoint requires an explicit endTime, so history is paged backwards from now and de-duplicated.
| artifact | what it holds |
|---|---|
| ledger.csv | 20,934 rows: anchor close, venue quote, fair value, realised reopen, and the weights in force for each |
| calibration.json | the frozen model — edges, lambdas, per-name betas, error bands, weekend statistics |
| tails.json | measured survival curves per name and per horizon, behind the collateral probability |
| tests/ | including the one asserting no window is ever scored by a model that saw it |
One thing does not come from the price endpoints: the earnings calendar. That is read live from Bitget's own agent server at agent.bitget.com/mcp over MCP — keyless, with the tool catalogue discovered at runtime rather than hardcoded, so the desk keeps working if the catalogue moves and simply says the calendar is unavailable if the server does not answer. It matters because a quote sitting far from fair value the night before a report is news arriving, not a stale mark, and the desk should not call those the same thing.
The Python side fetches, builds windows, fits and grades. The web side re-implements only the forward pass — market factor, shrinkage, band — so the browser can price the live board without a server. The numbers on this site are produced by the same arithmetic that produced the ledger.
Limits
- 01The weekend model is not fitted
Thirty-one weekend windows is too few to fit on without overfitting, so weekend behaviour is reported and not modelled. The board applies overnight weights on a weekend and says so.
- 02The edge is modest
Beating the venue on 54 to 56% of forecasts is an edge, not an oracle. On some names the venue's quote is simply better than ours, and those names are published next to the ones we win on.
- 03Hourly resolution
Everything is built on hourly bars. A quote that moves and reverts inside an hour is invisible to this desk.
- 04Thin books distort everything
Names that barely trade overnight produce error figures that look excellent because nothing happened. They are filtered out, not flattered.
- 05It cannot see news
Fair value is built from price alone. An earnings release at 21:00 ET is exactly the situation where the venue's quote deserves more trust than our shrinkage gives it.
The desk prices the live board every 45 seconds against the calibration described above.