A study of radio frequency interference for the Canadian Hydrogen Intensity Mapping Experiment (CHIME) Pathfinder
Bibliographic record
Abstract
The following document describes pursued studies to understand the properties of radio frequency interference (RFI) which affects the quality of the data of the Canadian Hydrogen Intensity Mapping Experiment Pathfinder at the Dominion Radio Astronomy Observatory in Penticton, British Columbia. The Canadian Hydrogen Intensity Mapping Experiment is a challenging project aimed to trace large scale structure by observing the 21cm emission line of neutral hydrogen in the frequency spectrum 400-800MHz to research the nature of Dark Energy. RFI is terrestrial signal caused by radio bands, TV stations, satellites etc. that produces unwanted disturbances in the frequency spectrum which adds power to the data. It represents a challenge to measure faint sources in the sky and we seek ways to identify it based on its statistical properties such as non-Gaussianity. We have designed algorithms that aim to identify and flag RFI in our data. Digital TV bands cause permanent corruption in the affected frequency bins and account for a 19% loss of bandwidth. The 5 sigma threshold cut searches for time-varying RFI in each frequency bin. Outliers above 5 standard deviations are iteratively flagged but not all of the occurring RFI were recognized due to non-Gaussianity. The median absolute deviation cut is a robust statistical method that uses sky data only. Identification of short-lived and long-lived RFI occurrences originating mainly from the sky has been successful. A correlation coefficient algorithm uses a combination of a reference RFI antenna sensitive to the horizon and a sky antenna to find correlated signals that are significantly above expected thermal noise of the radiometer while disregarding correlation due to sky signal. RFI at the horizon is well recognized by this method.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".