Historical Data on Repurchase Agreements from the Canadian Depository for Securities
Bibliographic record
Abstract
We develop an algorithm that extracts information about sale and repurchase agreements (repos) from disaggregated settlement data in order to generate a new historical dataset for research. Data from Canada’s fixed-income settlement authority, the Canadian Depository for Securities (CDS), is a valuable source of historical information on Canada’s fixed-income markets, especially from 2003 to 2016 when few other data sources were available. However, the CDS does not contain details on the terms of trade for repos, such as the repo rate, term or haircut. In the data, each repo is recorded as two distinct settlements but, critically, the sale and repurchase legs of a repo are not explicitly associated. We use a variant of the Gale-Shapley algorithm to solve a “stable roommates” problem to link repos’ sale and repurchase transactions and compute their terms of trade. We verify our algorithm by running it on a separate dataset that explicitly associates the sale and repurchase legs of a repo. In addition, we verify the computed repo terms of trade by comparing a subsample of the CDS data with a third dataset that reports terms of trade directly. The derived data are useful for researchers to study the evolution of fixed-income market structure and market conditions
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.030 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.005 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".