MétaCan
Menu
Back to cohort
Record W3208387531 · doi:10.5281/zenodo.3406508

AWARE characterization factor samples

2019· dataset· en· W3208387531 on OpenAlexaff
Pascal Lesage, Anne‐Marie Boulay, Stefan M. Pfister

Bibliographic record

VenueFigshare · 2019
Typedataset
Languageen
FieldEngineering
TopicIndustrial Vision Systems and Defect Detection
Canadian institutionsPolytechnique Montréal
Fundersnot available
KeywordsFactor (programming language)Characterization (materials science)Computer scienceMaterials scienceNanotechnologyProgramming language

Abstract

fetched live from OpenAlex

Files contain 4999 samples of AWARE characterization factors, as well as sampled independent data used in their calculations and selected intermediate results. AWARE is a consensus-based method development to assess water use in LCA. It was developed by the WULCA UNEP/SETAC working group. Its characterization factors represent the relative Available WAter REmaining per area in a watershed, after the demand of humans and aquatic ecosystems has been met. It assesses the potential of water deprivation, to either humans or ecosystems, building on the assumption that the less water remaining available per area, the more likely another user will be deprived. The code used to generate the samples can be found here: https://github.com/PascalLesage/aware_cf_calculator/ The following datasets are supplied: <strong>1) AWARE_characterization_factor_samples.zip</strong> Actual characterization factors resulting from the Monte Carlo Simulation. Contains 4 zip files: * monthly_cf.zip: contains 116,484 arrays of 4999 monthly characterization factor samples for each of 9707 watershed and for each month, in csv format. Names are cf_&lt;BAS34S_ID&gt;_&lt;MONTH&gt;.csv, where &lt;BAS34S_ID&gt; is the watershed id and &lt;MONTH&gt; is the first three letters of the month ('jan', 'feb', etc.). * average_agri_cf.zip: contains 9707 arrays of 4999 annual average, agricultural use, characterization factor samples for each watershed, in csv format. Names are cf_average_agri_&lt;BAS34S_ID&gt;.csv. * average_non_agri_cf.zip: contains 9707 arrays of 4999 annual average, non-agricultural use, characterization factor samples for each watershed, in csv format. Names are cf_average_non_agri_&lt;BAS34S_ID&gt;.csv. * average_unknown_cf.zip: contains 9707 arrays of 4999 annual average, unspecified use, characterization factor samples for each watershed, in csv format. Names are cf_average_unknown_&lt;BAS34S_ID&gt;.csv.. <strong>2) AWARE_base_data.xlsx</strong> Excel file with the deterministic data, per watershed and per month, for each of the independent variables used in the calculation of AWARE characterization factors. Specifically, it includes: Monthly irrigation<br> Description: irrigation water, per month, per basin<br> Unit: m3/month<br> Location in Excel doc: Irrigation<br> File name once imported: irrigation.pickle<br> table shape: (11050, 12) Non-irrigation hwc: electricity, domestic, livestock, manufacturing<br> Description: non-irrigation uses of water<br> Unit: m3/year<br> Location in Excel doc: hwc_non_irrigation<br> File name once imported: electricity.pickle, domestic.pickle,<br> livestock.pickle, manufacturing.pickle<br> table shape: 3 x (11050,) avail_delta<br> Description: Difference between "pristine" natural availability<br> reported in PastorXNatAvail and natural availability calculated<br> from "Actual availability as received from WaterGap - after<br> human consumption" (Avail!W:AH) plus HWC.<br> This should be added to calculated water availability to<br> get the water availability used for the calculation of EWR<br> Unit: m3/month<br> Location in Excel doc: avail_delta<br> File name once imported: avail_delta.pickle<br> table shape: (11050, 12) avail_net<br> Description: Actual availability as received from WaterGap - after human consumption<br> Unit: m3/month<br> Location in Excel doc: avail_net<br> File name once imported: avail_net.pickle<br> table shape: (11050, 12) pastor<br> Description: fraction of PRISTINE water availability that should be reserved for environment<br> Unit: unitless<br> Location in Excel doc: pastor<br> File name once imported: pastor.pickle<br> table shape: (11050, 12) area<br> Description: area<br> Unit: m2<br> Location in Excel doc: area<br> File name once imported: area.pickle<br> table shape: (11050,)<br> It also includes: * information on the distributions used for each variable (uncertainty tab) * two filters used to exclude watersheds that are either in Greenland (polar filter) or without data from the Pastor et al. (2014) method (122 cells), representing small coastal cells with no direct overlap (pastor filter). (filters tab) <strong>3) independent_variable_samples.zip</strong> Samples for each of the independent variables used in the calculation of characterization factors. Only random variables are contained. For all watershed or watershed-months without samples, the Monte Carlo simulation used the deterministic values found in the AWARE_base_data.xlsx file. The files are in csv format. The first column contains the watershed id (BAS34S_ID) if the data is annual or the (BAS34S_ID, month) for data with a monthly resolution. the other 4999 columns contain the sampled data. The names of the files are &lt;variable_name.csv&gt;. <strong>4) intermediate_variables.zip</strong> Contains results of intermediate calculations, used in the calculation of characterization factors. The zip file contains 3 zip files: * AMD_world_over_AMD_i.zip: contains 116,484 arrays (for each watershed-month) of 4999 calculated values of the ratio between the AMD (Availability Minus Demand) for the watershed-month and AMD_glo, the world weighted AMD average. Format is csv.<br> * AMD_world.zip: contains one array of 4999 calculated values of the world average AMD. Format is csv. * HWC.zip: contains 116,484 arrays (for each watershed-month) of 4999 calculated values of the total Human Water Consumption. Format is csv. <strong>5) watershedBAS34S_ID.zip</strong> Contains the GIS files to link the watershed ids (BAS34S_ID) to actual spatial data.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Insufficient payload (model declined to judge)
Consensus categoriesInsufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.188
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0010.000
Insufficient payload (model declined to judge)0.2110.024

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.080
GPT teacher head0.258
Teacher spread0.178 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2019
Admission routes1
Has abstractyes

Explore more

Same venueFigshareSame topicIndustrial Vision Systems and Defect DetectionFrench-language works237,207