MétaCan
Menu
Back to cohort
Record W2296655817 · doi:10.14288/1.0077401

The construction of a criterion-referenced physical education knowledge test

2010· article· en· W2296655817 on OpenAlexaff
Gail E. Wilson

Bibliographic record

VenuecIRcle (University of British Columbia) · 2010
Typearticle
Languageen
FieldHealth Professions
TopicPhysical Education and Pedagogy
Canadian institutionsUniversity of British Columbia
Fundersnot available
KeywordsTest (biology)Criterion-referenced testMathematicsComputer scienceStatisticsGeologyStandardized test

Abstract

fetched live from OpenAlex

Throughout the last two decades, physical educators have worked to develop a specific body of knowledge. Associated with the formation of this body of knowledge has been a trend by most physical educators to include a cognitive objective as one of the stated aims in their physical education, curricula. As a result, the need for adequate knowledge assessment instruments has become apparent. Although some assessment of knowledge in physical and health education has occurred since the late 1920's, the majority of tests which have been developed to date are directed towards the evaluation of knowledge in specific sports or activities. Relatively few tests are available that assess general knowledge concepts in physical education. As well, all of the knowledge tests that have been produced are norm-referenced' instruments. That is, they have been constructed for the purpose of ranking individuals and comparing differences among them. The purpose of this study was to design a criterion-referenced test which would assess the physical education knowledge of grade eleven high school students in British Columbia and which could function as a measurement instrument for the evaluation of groups or classes. As a criterion-referenced assessment tool, the knowledge test assesses the performance of individuals based on' objectives which had been previously formulated by the Learning Assessment Branch of the Ministry of Education in British Columbia. In order to prepare a table of specifications for the design of the test, the specific objectives to be measured were grouped into six subtest areas. Multiple-choice items were then constructed according to the requirements of the table of specifications. For the initial pilot administration of the test, two test forms, of 48 items each, were developed. Each of these forms included three of the six sub-test areas. One half of the 288 students to whom the first pilot was administered answered Form A while the remaining students answered Form B. Following the administration of pilot test 1, the results obtained were analysed by the Laboratory of Educational Research Test Analysis Package (LERTAP), and were subjectively reviewed by an advisory panel. As a result of these procedures, 70 items were retained for use on the second pilot test. This test was administered to 133 students and the results were again analysed subjectively and psychometrically. Thirty-eight items from pilot test 2 were considered acceptable for use on the final pilot test. In order to maintain adherence to the table of specifications, nine new items were developed and after approval by the advisory panel, were included on the third test form. This form was given to 800 grade eleven students and the responses of 250 randomly selected students were analysed by the LERTAP procedure. The analysis indicated that all items were psychometrically sound and the reliability of this form was estimated at .71. Thus, the items utilized during the third pilot administration constituted the final form of the knowledge test. The test is suitable for evaluating groups and the six sub-tests, as well as the total test, can be used to identify strengths and weaknesses within programs.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.953
Threshold uncertainty score0.974

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0010.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.025
GPT teacher head0.333
Teacher spread0.308 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2010
Admission routes1
Has abstractyes

Explore more

Same venuecIRcle (University of British Columbia)Same topicPhysical Education and PedagogyFrench-language works237,207