Grouping Data On Infrastructure Development In Langkat District Using The Clustering Method (Case Study: PUPR, Langkat Regency)
Bibliographic record
Abstract
A building is a man-made structure consisting of walls and a roof permanently erected in a place. Buildings can also be called houses and buildings, namely all facilities, infrastructure or infrastructure in culture as well as human life in building their civilization. Public Works and Public Housing (PUPR) play an important role in increasing the development of national infrastructure in Indonesia so that PUPR can assist in clustering research in infrastructure development in Langkat Regency which is very large every year by grouping the data based on activity names, company names, sub-districts development, and look at the last four years.To classify existing development infrastructure in Langkat Regency with the previous system used by the PUPR Service which is still running by recording in a ledger and hindering reporting performance in grouping PUPR service infrastructure development in road construction, bridge construction and others. So that the existence of grouping using the clustering method helps the PUPR service in clustering infrastructure development data in Langkat Regency to be more effective and efficient.The clustering method is one of the methods that can be applied in classifying infrastructure development data taken from the analysis of Langkat Regency PUPR data regarding developments that have taken place in several sub-districts in Langkat Regency. This clustering method has been widely used by previous studies to group data
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.006 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".