ORDER STATISTIC FILTER (OSF): A NOVEL APPROACH TO DOCUMENT ANALYSIS
Bibliographic record
Abstract
Page segmentation is one of the important and basic research subjects of document analysis. There are two major kinds of page segmentation methods, i.e. hierarchical and no-hierarchical ones. Most traditional techniques such as top–down and bottom–up approaches belong to the hierarchical method. Though these two approaches have been used till now, they are not effective for processing documents with high geometric complexity and the process of splitting document needs iterative operations which is time consuming. A non-hierarchical method called the modified fractal signature (MFS) was presented in recent years. It can overcome the above weaknesses, however the MFS needs to calculate modified fractal signature which makes the theory very complex. In this thesis, we present a new page segmentation approach: Median Order Statistic Filter (MedOSF) — Maximum Order Statistic Filter (MaxOSF) approach which is more direct and much simpler. We use the MedOSF to remove the salt–pepper noise of the document and use the MaxOSF to do the page segmentation. In practice, they not only can adaptively process the documents with high geometrical complexity, but also save a lot of computing time.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".