Application of the Clustering Algorithm for the Classification of General Criminal Cases at the Binjai District Attorney's Office
Bibliographic record
Abstract
The Binjai District Attorney's Office in carrying out its duties and functions, one of which handles general crimes, where so far the SPDP (Warranty to Commence Investigation) from the police that has entered the Binjai District Prosecutor's Office amounted to approximately 50 (fifty) cases each month. This amount consists of several types of general criminal cases. It is known that the types of general criminal cases amount to approximately 215 (two hundred and fifteen) types of cases, from this data, a method of classifying/clustering is needed from the types of cases that exist each month so that the data can be processed so as to produce the highest, moderate and highest scores. the lowest value of a type of case. The Binjai District Attorney's Office often receives requests for data from other ministries or agencies such as the BPS (Central Statistics Agency), the National Commission on Women and the National Commission on Children in the form of data recapitulation of crimes against women and children as perpetrators of crimes. The Binjai District Attorney's Office has a case handling system where the recapitulation cannot be taken directly but instead collects data manually, because the existing case handling system does not have the recapitulation as requested.The application of clustering has been carried out by many previous researchers. Among them, the K-Means Clustering Algorithm Analysis Mapping the Number of Crimes. The research was carried out using a data mining model in classifying illegal fishing with the K-Means algorithm analysis by determining the shortest distance using the eulclidean distance, more optimal than using the mahattan distance and chbchep distance in classifying student achievement, determining the centroid (central point) in the early stages of the algorithm K-Means is very influential on cluster results as the results of tests carried out using 267 records with different centroids produce different cluster results as well, a clustering model is obtained that can be used for illegal fishing in decision making for illegal fishing crimes high, medium, moderate .
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.006 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".