- Insights
- 2025
- Building a reference-data spine for markets the vendors do not cover
Insights · Data Analytics
Building a reference-data spine for markets the vendors do not cover
A backtest run on a restated vendor file measures the file. This note explains what a point-in-time data spine is, why a value is never overwritten, what the discipline costs each year, and the four strategies it stopped before they held capital.

Key points
A backtest on a restated vendor file measures the file, and it will not survive contact with capital.
Append every correction and overwrite nothing, so a result can be reproduced against what was known that day.
Four strategies rejected on point-in-time testing paid for a decade of the data spine.
The Quantitative Strategies division trades instruments that commercial data vendors cover thinly or not at all: small-issue corporate credit, regional interest-rate products, listed equities in markets holding a few hundred names, and the futures used to hedge them. Coverage gaps are not an inconvenience. A backtest built on a vendor file that has been restated since is a result about the file rather than about the market. The division therefore builds and keeps its own reference-data spine.
A spine is three things. An identifier map ties every instrument to every code it has ever carried. A point-in-time store records what was known on each date, including what was later found to be wrong. A corporate-action history explains every discontinuity in a price series. The first is tedious. The second decides whether the research is honest. The third is where errors hide. Together they occupy 340 gigabytes, which is small, and five people, which is not.
The operating rule is that a value is never overwritten. A correction is appended with its own timestamp and the original is retained. That rule is unpopular during a busy week and it is not negotiable. Without it, a backtest run in March and repeated in September produces two answers and no explanation for the difference. With it, any result can be reproduced against the data as it stood on the day it was produced. The Model Risk Management policy requires exactly that lineage before a model may hold capital.
Vendor feeds are inputs to the spine and not the spine itself. Humber Ridge Data supplies North American pricing and Sumida Information Systems supplies Japanese instrument and corporate-action data. Both are held to the Ethical Technology Charter through the Third-Party Risk and Outsourcing policy, and both are obliged by contract to notify restatements rather than correct them silently. Where a supplier will not accept that obligation, the division takes the raw feed and reconstructs the history itself. Two of its markets are covered that way.
The spine cost US$2.1 million in 2024, of which US$1.4 million was people. Against a quantitative book capped by the Risk & Valuation Committee at 18 per cent of the proprietary balance sheet, that is a heavy overhead for a small division. It is defended on one ground. Between 2022 and 2024 the division rejected four strategies whose apparent performance disappeared when they were tested against point-in-time data rather than the current file. Any one of them would have cost more than a decade of the spine.
The data is not sold. The firm does not sell client data and it does not sell this either, although a market exists for both. The reason is not commercial. A dataset that is also a product acquires customers, and customers acquire influence over what is collected and how it is corrected. The division needs a record that answers only to research. Transparency to clients is met another way: any client whose mandate is model-assisted may see which models were used and what data they read.
Published 2025-02-18 by the Quantitative Strategies division. Research is prepared for eligible counterparties and does not constitute advice.
More on the sector