Retrieving of particulate matter from optical measurements: A semiparametric approach
Bibliographic record
Abstract
The fine particle abundance, i.e., particle matter (PM) concentration, is one of the indicators of air quality and is therefore subject to ground‐based measurements. Complementary satellite aerosol remote sensing techniques provide one with maps of the aerosol optical thickness (AOT), which is sensitive to particle abundance. This paper investigates the problem of retrieving the PM concentration from the AOT, both on daily average values, on the basis of a large data set where data from the air quality networks are combined with ground‐based measurements of the AOTs. It is found that a linear model fails at explaining the data well but that the performance may be significantly improved when such a linear relationship is conditioned on auxiliary parameters, mainly meteorological variables. The proposed model is expressed as an additive varying coefficient model (AVCM), which is defined as a linear model where the coefficients are additive functions of the auxiliary parameters. The model is represented using penalized smoothing splines, allowing for a proper control of the overall number of degrees of freedom via multiple smoothness parameters selection. The methodology is applied to data collected around Lille (France). The PM10concentrations are retrieved with an average uncertainty of less than 20%, leading to a correlation coefficient of 0.87 between fitted and expected PM10.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".