Validity of Teacher-Made Assessment: A Table of Specification Approach
Bibliographic record
Abstract
The validity of teacher-made assessment remains debatable in the educational assessment process. This studyinvestigates the content validity of teacher-made assessment in three Chinese Elementary Schools in Johor,Malaysia. It also examines teacher understanding of table of specification in the sampled schools. Aquestionnaire with 10 items was distributed to 30 teachers in order to collect the data on table of specification.Items 1 to 4 examine teacher understanding of the table of specification while items 5 to 10 test the contentvalidity of teacher-made assessment. The results showed that teachers exhibited a low understanding of the tableof specification. The analysis revealed that the majority of them never attended courses concerning table ofspecification and were unable to build a comprehensive table of specification for the subjects they teach. Thefindings also demonstrated that teacher-made assessment was valid in terms of content validity. However, mostof the teachers did not refer to the table of specification while building instruments for assessment. This indicatesthat teachers lack basic knowledge in designing a standard table of specification and they lack awareness on theimportance of the table of specification. Recommendations of the study for teacher-made assessmentimprovements were also addressed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.168 | 0.373 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.008 | 0.007 |
| Science and technology studies | 0.003 | 0.006 |
| Scholarly communication | 0.006 | 0.007 |
| Open science | 0.003 | 0.005 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".