Arctic Ocean Net Primary Production Model Code
Bibliographic record
Abstract
Contributors listed in alphabetical order Model Description The Takuvik Net Primary Production (TNPP) model is a light photosynthesis model that uses satellite data to estimate the net primary production (NPP) in the Arctic Ocean. The model was run on a Pan-Arctic scale (above 45°N), at 4 km resolution, and comes from the updated depth and wavelength resolved model from Belanger et al. (2013). The update includes improved resolution of atmospheric products and the addition of verticality for the chlorophyll profile according to Ardyna et al. (2013). The model was originally run using MODIS reflectances as input data. Currently the TNPP model was adapted to use reflectances at the same wavelengths as initially but derived from ESA's OC-CCI v6.0 product. Other inputs needed to run the model are atmospheric variables, bathymetry and chlorophyll-a concentration. The latter variable was estimated with a semi-analytical GSM algorithm, which, according to the work of Li et al., (2024), performed better than other ocean colour models to estimate chlorophyll-a in the Arctic Ocean. A description of the NPP model, including input data, methods, intermediate variables, constant values and photosynthesis models, is available in the articles by Belanger et al. (2013) and Li et al. (2024). Input files An input dataset for 2019-08-17 (day number 229 of the year 2019) has been uploaded to be used as an example to run the model including the bathymetry, reflectances, chlorophyll-a concentration and atmospheric data. Once you have downloaded the ppv1.zip file from this Zenodo repository and unzipped it on your computer, the input data for the example will be in the folder where you unzipped the file ( /ppv1/tutorial/inputs/) Atmospheric variables. The irradiance at the sea surface is retrieved using look-up-table of Ed (0+, lamda, t) at 5nm resolution each 3h time step, derived from the Santa Barbara DISORT Atmospheric Radiative Transfer Model (SBDART, Ricchiazi et al.,1998) following the methodology of Bélanger et al. (2013). The inputs data of the look-up-table are the daily values of solar zenith angle (theta_s), cloud fraction (CF) and total ozone concentration (O3) comes from MODIS (Platnick et al., 2015). Other atmospheric parameters such as the water vapor content and aerosol optical thickness were taken from climatological data (Zhang et al., 2004). Reflectances (Rrs). In this version of the TNPP model, daily reflectances from OC-CCI v6.0 data (Sathyendranath et al., 2023) have been used. This level 3 binned (L3b) product is based on a prior reflectance merging from multi-sensor time-series of satellite ocean-colour data. Chlorophyll-a concentration (CHL). The CHLA (mg.m-3) were calculated using the arctic-optimized version of the semi-analytical Garver-Siegel-Maritorena algorithm (GSM) from Li et al. (2024). Bathymetry. The International Bathymetric Chart of the Arctic Ocean (IBCAO) 5.0, which offers a 100 × 100 m grid cells resolution (Jakobsson et al., 2024) was used for the bathymetry. Docker image The Takuvik NPP model has been built in a Docker image. The first step is to install Docker on your machine. Please consult Docker documentation to install Docker on your computer. Download ppv1 docker image Once Docker is installed, you will need to pull the ppv1 image on your machine. At this time, the docker is private and you will need to log on Docker Hub first. Create an account if needed. You can use this link to view the PPv1 image. docker login docker pull takuvik/ppv1 docker login will prompt you for your credential. docker pull takuvik/ppv1 will download the PPv1 docker image on your computer. Modifying the code The code of the model can be modified. Only if you want to modify the code do you need to follow the steps below (build and push ppv1 docker image). Otherwise, you can skip to the Running the ppv1 example section. You will have to first download the ppv1.zip file from this Zenodo repository, and unzip the file in . The ppv1 is also available on GitHub (https://github.com/POMPTakuvik/ppv1.git). Once you have downloaded the ppv1 repository on your local folder, you can browse and modify the source code contained in the Source folder. /ppv1/trunk/Source/ Build ppv1 docker image After you have made modifications to the source code, you will have to rebuild the ppv1 docker: docker build -t ppv1 . Note: You need to be in ppv1/trunk/Source/ folder (where the Dockerfile is located) to run the build command. Push ppv1 docker image docker push takuvik/ppv1 Note: We push the build image to our docker takuvik/ppv1, for you will be /ppv1 Running the ppv1 example Once the ppv1 docker is installed and after a successful docker build you will be able to execute a test run as follows after making some quick changes to the configuration file (see later). It takes about 20 minutes to run a single day (depending on your PC configuration and the number of documented pixels). -v section Docker needs access to the configuration file, input and output paths, which are stored outside of Docker. The -v argument maps [host folder]:[container folder]. These paths will highly depend on your configuration file and your computer setup. docker run -i\-v /ppv1/tutorial/inputs/config:/takuvik/configuration_file \-v /ppv1/tutorial/inputs/:/data/ \takuvik/ppv1 configuration_file/ where / is the folder where is placed the ppv1 unzipped folder, and / is the file that provide information about the input data to use, running start and end dates and vertically (in the example we used the config_example.json) Modifying the ppv1 configuration file ./ppv1/tutorial/inputs/config/config_example.json The two first highlighted terms of the configuration file must be modified to specify the start and end calculation dates. For the "path format" for the input variables, the path to be followed to get the input variables has been defined in the previous step (the -v section). { "calculation_start_date":"2019-08-17T00:00:00.000Z", "calculation_end_date": "2019-08-17T00:00:00.000Z", "calculation_time_step_in_offset_alias": "1D", "input_file": { "takuvik_atmosphere": { "file_path_template_type": "year_and_day_of_year", "file_type": "netcdf_indexed_one_level", "path_format": "/data/MODISA/L3BIN/{}/{}/", "file_name_format": "A{}{}_061_.L3b_DAY_ATMOSPHERE_above_45n.nc", "details": { "index_variable": "bin_index" } }, "takuvik_rrs": { "file_path_template_type": "year_and_day_of_year", "file_type": "netcdf_indexed_one_level", "path_format": "/data/CCI/CCI_v6.0/formatted_for_ppv1/{}/{}/", "file_name_format": "CCI_{}{}_L3b_DAY_RRS_above_45n.nc", "details": { "index_variable": "bin_index" } }, "takuvik_chla": { "file_path_template_type": "year_and_day_of_year", "file_type": "netcdf_from_bloomstate2", "path_format":"/data/Bloomstate2/CCI_v6.0/{}/{}/", "file_name_format": "C{}{}_chlz_00_09.nc", "details": { "index_variable": "bin_index" } }, "takuvik_bathy": { "file_path_template_type": "constant_file", "file_type": "csv", "path_format": "/data/Bathymetry/", "file_name_format": "Province_Zbot_MODISA_L3binV2.csv>, "details": { "index_variable": "bin_index" } }, "other_values": { "file_path_template_type": "constant_file", "file_type": "flat_json", "path_format": "/data/other/", "file_name_format": "other_values_pp_integration.json" }, "downward_irradiance_table": { "file_path_template_type": "constant_file", "file_type": "base", "path_format": "/data/LUTS/", "file_name_format": "Ed0moins_LUT_5nm_v2.dat" } }, "variable": { "rrs_412": { "input_file_name": "takuvik_rrs", "column_name": "Rrs412", "type": "vector" }, "rrs_443": { "input_file_name": "takuvik_rrs", "column_name": "Rrs443", "type": "vector" }, "rrs_488": { "input_file_name": "takuvik_rrs", "column_name": "Rrs488", "type": "vector" }, "rrs_531": { "input_file_name": "takuvik_rrs", "column_name": "Rrs531", "type": "vector" }, "rrs_555": { "input_file_name": "takuvik_rrs", "column_name": "Rrs555", "type": "vector" }, "rrs_667": { "input_file_name": "takuvik_rrs", "column_name": "Rrs667", "type": "vector" }, "cloud_fraction": { "input_file_name": "takuvik_atmosphere", "column_name": "CF_mean", "type": "vector" }, "taucl": { "input_file_name": "takuvik_atmosphere", "column_name": "TauCld_mean", "type": "vector" }, "ozone": { "input_file_name": "takuvik_atmosphere", "column_name": "O3_mean", "type": "vector" }, "latitude": { "input_file_name": "takuvik_rrs", "column_name": "lat", "type": "vector" }, "longitude": { "input_file_name": "takuvik_rrs", "column_name": "lon", "type":
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.104 | 0.078 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".