Use of a (Quantitative) Structure–Activity Relationship [(Q)Sar] Model to Predict the Toxicity of Naphthenic Acids
Bibliographic record
Abstract
Naphthenic acids (NA) are a complex mixture of carboxylic acids that are natural constituents of oil sand found in north-eastern Alberta, Canada. NA are released and concentrated in the alkaline water used in the extraction of bitumen from oil sand sediment. NA have been identified as the principal toxic components of oil sands process-affected water (OSPW), and microbial degradation of lower molecular weight (MW) NA decreases the toxicity of NA mixtures in OSPW. Analysis by proton nuclear magnetic resonance spectroscopy indicated that larger, more cyclic NA contain greater carboxylic acid content, thereby decreasing their hydrophobicity and acute toxicity in comparison to lower MW NA. The relationship between the acute toxicity of NA and hydrophobicity suggests that narcosis is the probable mode of acute toxic action. The applicability of a (quantitative) structure-activity relationship [(Q)SAR] model to accurately predict the toxicity of NA-like surrogates was investigated. The U.S. Environmental Protection Agency (EPA) ECOSAR model predicted the toxicity of NA-like surrogates with acceptable accuracy in comparison to observed toxicity values from Vibrio fischeri and Daphnia magna assays, indicating that the model has potential to serve as a prioritization tool for identifying NA structures likely to produce an increased toxicity. Investigating NA of equal MW, the ECOSAR model predicted increased toxic potency for NA containing fewer carbon rings. Furthermore, NA structures with a linear grouping of carbon rings had a greater predicted toxic potency than structures containing carbon rings in a clustered grouping.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".