Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Extracting Evidence-linked Abundance Drivers from Species-account Text

UK Centre for Ecology & Hydrology

Challenge and Methodological Approach Summary

Species accounts contain useful ecological knowledge about habitat, management, biogeography and likely reasons for change. However, that knowledge is usually written as prose, which makes it difficult to reuse directly in structured workflows such as causal diagrams or expert-review tables.

This notebook demonstrates a constrained workflow for using a small local instruction-tuned language model to extract candidate ecological driver relationships from species-account text. Each sentence is processed with a short context window, and the model returns a small JSON object containing a driver phrase, driver type, target, effect direction, confidence score and evidence quote. Standard Python code then validates, groups and aggregates these rows into evidence-linked candidate edges for expert review.

Important:
The output is not an automatically generated causal model. The model proposes evidence-linked rows; the analyst decides whether the evidence is useful, whether the direction is plausible and whether the edge belongs in a causal diagram.

Introduction

Biodiversity records are rarely collected from a perfectly representative sample of places. Some sites are visited more often because they are accessible, well known, close to recorders, or more likely to contain species of interest. These patterns matter when inferring changes in species’ distributions or abundance from opportunistic records.

Causal diagrams are one way to reason about this problem. They make explicit which processes are thought to affect a species, which processes affect where records are collected, and where those processes overlap. Boyd et al. (2025) discuss this in the context of using causal diagrams and superpopulation models to correct geographic biases in biodiversity monitoring data.

In practice, much of the ecological knowledge needed to draft these diagrams is held as prose. Plant Atlas-style species accounts describe habitat, management, broad biogeography, recent trends and likely reasons for change, but those descriptions are not directly usable as candidate graph edges.

This notebook asks a narrow practical question: can a small local LLM help convert species-account prose into an evidence-linked table of candidate abundance drivers for expert review?

The included demo data are synthetic Plant Atlas-style examples. They mimic the structure of local account extracts, but they are not copied Plant Atlas prose and should not be quoted as Plant Atlas content.

Running the notebook

Clone the accompanying notebook repository, create the conda environment, and point MODEL_PATH to a locally available Hugging Face-compatible instruction model.

conda env create -f environment.yml
conda activate plant-atlas-llm-dag

A typical local-model setup is:

export MODEL_PATH=../models/qwen2.5-3b-instruct
export LLM_LOAD_IN_4BIT=1
export LLM_4BIT_QUANT_TYPE=nf4
export LLM_4BIT_USE_DOUBLE_QUANT=1
export LLM_4BIT_COMPUTE_DTYPE=float16
export TRANSFORMERS_OFFLINE=1
export HF_HUB_OFFLINE=1

# Recommended first run.
export LLM_MAX_SENTENCES=10
export LLM_BATCH_SIZE=1
export LLM_MAX_NEW_TOKENS=128

If the smoke test works, unset LLM_MAX_SENTENCES or increase it gradually. The notebook intentionally does not download model weights during execution, so text, prompts and model outputs remain inside the local analysis environment.

General use

The example is framed around Plant Atlas-style accounts, but the pattern is more general. It may also be useful for habitat management plans, protected-area reports, species recovery documents, invasive-species risk assessments or literature-screening outputs.

The key constraint is that the model task remains narrow and reviewable. The LLM handles sentence-level JSON extraction; ecological judgement remains with the analyst.

Why a small local LLM?

A small local model is used because the task has been deliberately narrowed. The model sees one sentence at a time and returns a compact schema: driver phrase, driver type, target, effect direction, confidence and evidence quote.

This is also a proportionate design choice. LLM workflows are not automatically environmentally benign, because training and inference have compute, energy and hardware costs. The aim here is to use a model that is sufficient for a constrained sentence-level extraction step, rather than starting with a much larger hosted model.

Importing required libraries

The workflow uses standard scientific Python packages plus a required local LLM route through Hugging Face transformers. If a local model is not available, the notebook fails early with a setup error rather than silently using a non-LLM substitute.

Warning and logging filters are applied so that the rendered notebook focuses on the workflow and results. Sentence-level parsing status is still written to an audit CSV for reproducibility and troubleshooting.

Notebook started: 2026-04-28T23:04:34

Configuration

This cell collects the main run settings. For the built-in synthetic example, leave USE_DEMO_DATA = True. For a real analysis, set USE_DEMO_DATA = False and either place local CSV/ZIP files in PLANT_ATLAS_DATA_DIR or add explicit paths to INPUT_PATHS.

The default MODEL_PATH is the path used during development. On another machine, set MODEL_PATH before starting Jupyter so that it points to a local Hugging Face-compatible instruction model directory.

Run configuration hash: ead3a98f083d
{'use_demo_data': True, 'data_dir': 'data', 'input_paths': [], 'output_dir': 'outputs_plant_atlas_llm_to_dag', 'context_window_sentences': 1, 'max_sentence_chars': 900, 'model_path': '../models/qwen2.5-3b-instruct', 'llm_local_files_only': True, 'llm_temperature': 0.0, 'llm_max_new_tokens': 192, 'llm_max_input_tokens': 1536, 'llm_batch_size': 1, 'llm_device_map': 'auto', 'llm_torch_dtype': 'auto', 'llm_load_in_4bit': True, 'llm_4bit_quant_type': 'nf4', 'llm_4bit_use_double_quant': True, 'llm_4bit_compute_dtype': 'float16', 'show_extraction_diagnostics': False, 'llm_max_sentences': None, 'phrase_similarity_threshold': 0.55}

Workflow overview

The local LLM is used only where language understanding is useful: converting short prose into a structured record. Everything around that step is explicit Python, which keeps the route from source sentence to proposed graph edge inspectable.

The broad workflow is:

<Figure size 1200x240 with 1 Axes>

Create or load species-account text

The built-in rows below are synthetic. They mimic the structure of local Plant Atlas-style CSV extracts, including fields such as canonical, atlasSpeciesDescription, atlasSpeciesBiogeography and atlasSpeciesTrends. They are not copied Plant Atlas account text and should not be quoted as Plant Atlas content.

For a real analysis, set USE_DEMO_DATA = False and point PLANT_ATLAS_DATA_DIR at local CSV or ZIP files containing the account text to be analysed.

Loading...
Loaded 5 species accounts

Data exploration

Before running the LLM, check that the input text looks sensible: how many species are present, how many sentences have been created, and which broad terms appear most often. These simple checks catch common ingestion problems such as selecting the wrong text column or duplicating accounts.

Sentence rows: 30
Loading...
<Figure size 900x500 with 1 Axes>
Loading...
<Figure size 900x500 with 1 Axes>

Rule-based sanity check

The next cell is a simple rule-based sanity check. It is not the main method and it is not used as a substitute for the LLM extraction. It provides a quick comparator: if a sentence contains an obvious phrase such as drainage or grazing, the keyword pass should usually find it.

This is useful when developing prompts because it separates two problems: whether the input text contains relevant ecological phrases at all, and whether the local model extracts them in the desired schema.

Baseline deterministic candidate edges: 8
Loading...

Local LLM evidence extraction

This is the main extraction step. Each sentence is sent to the local model with a short neighbouring context window and a constrained JSON schema. The model is asked for candidate driver relationships only when the sentence contains evidence for a factor that may affect habitat suitability, population performance or abundance.

The prompt is deliberately narrow. The model is not asked to write a causal diagram or summarise a species account. It only proposes sentence-grounded candidate rows that can be validated and reviewed later.

You are helping ecologists convert Plant Atlas species-account prose into
reviewable candidate causal-graph edges.

Species: Kickxia spuria
Sentence ID: 0

CURRENT SENTENCE:
A small annual of arable fields, open disturbed ground and sunny field edges, most often on light calcareous soils.

NEIGHBOURING CONTEXT FOR DISAMBIGUATION ONLY:
[CURRENT] A small annual of arable fields, open disturbed ground and sunny field edges, most often on light calcareous soils. [context_1] It can persist where the crop is open and where cultivation creates bare ground for germination.

Task:
Extract candidate ecological driver relationships from the CURRENT SENTENCE.
Use only evidence stated in the CURRENT SENTENCE. Do not infer from general
knowledge. Ignore dates, pure locations, taxonomic notes, recording history and
introduction-history statements unless they are directly linked to abundance,
habitat suitability or population performance.

Return one JSON object and nothing else. No markdown. No explanation. No trailing
text. Keep source_raw labels short so that they can be read on a graph. If the current sentence contains no relevant relationship, return exactly:
{"edges": []}

Definitions:
- source_raw: an exact, short driver phrase copied from the CURRENT SENTENCE. Prefer compact noun phrases of one to four words, for example "grazing", "bare ground", "calcareous soils" or "herbicide use". Do not use a whole clause as the source.
- source_type: one of climate, hydrology, soil_substrate, habitat_structure,
  disturbance, management, land_use, biotic_interaction, disease_pathogen,
  pollution, topography or other.
- target_canonical: one of abundance, habitat suitability or population performance. Use abundance when the sentence says the driver maintains, increases, reduces, threatens, explains losses of, or is associated with local populations, colonies, rarity, decline, persistence or abundance. Use habitat suitability only when the sentence is mainly about habitat condition. Use population performance for flowering, recruitment, seed set or vigour.
- effect: positive means the source tends to increase or maintain the target;
  negative means it tends to reduce the target; mixed means both are stated;
  unknown means the relationship is stated but direction is unclear.
- confidence: a number from 0 to 1.
- evidence_quote: a short quote copied from the CURRENT SENTENCE.

JSON schema example:
{
  "edges": [
    {
      "source_raw": "exact phrase copied from the current sentence",
      "source_type": "climate|hydrology|soil_substrate|habitat_structure|disturbance|management|land_use|biotic_interaction|disease_pathogen|pollution|topography|other",
      "target_canonical": "abundance|habitat suitability|population performance",
      "effect": "positive|negative|mixed|unknown",
      "confidence": 0.7,
      "evidence_quote": "short quote copied from the current sentence"
    }
  ]
}
Prepared 30 sentence-level LLM calls.
Batch size: 1; max_new_tokens: 192; max_input_tokens: 1536
Loading local LLM from: ../models/qwen2.5-3b-instruct
This one-off model-load step can take a few minutes on first run.
4-bit quantised loading: True
Loading...
Loading...
LLM-assisted candidate edges: 37
Sentences processed: 30
Sentences returning at least one candidate edge: 29
Loading...

Inspect local LLM extraction coverage

The first check is whether the model produced candidate edges, how many sentences failed parsing or generation, and how many candidate edges were returned per species. This is a diagnostic step, not evidence that the extracted edges are correct.

For a real analysis, inspect the error rate before moving on. A few failures may be acceptable if the review table still contains useful evidence, but a high failure rate usually means the prompt, model or token limits need attention.

<Figure size 800x350 with 1 Axes>
<Figure size 800x350 with 1 Axes>
Loading...

Validate candidate edges

The validation step checks that each proposed row has the required fields, allowed values and a quote that can be linked back to the source sentence. Rows that fail validation are retained with flags rather than hidden.

This keeps the workflow auditable: the review table shows both the proposed relationship and the reason it should be treated with caution.

Validated candidate edges: 37
Loading...
accepted_for_environmental_dag True 33 False 4 Name: count, dtype: int64

Group similar driver phrases

The LLM extracts phrases as they appear in the sentence, so the same ecological idea may appear in several wordings. For example, fertiliser inputs, high fertiliser inputs and fertilised grassland should often be reviewed together. This section uses a lightweight token-overlap approach to create provisional phrase clusters.

For a larger analysis, this step could be replaced by embedding-based clustering or an expert-curated vocabulary. A transparent string-based method is used here because it is easy to inspect.

Semantic clusters: 30
Loading...

Visualise extracted evidence

The plots below summarise the extracted candidate rows before they are aggregated into graph edges. They are intended as quick checks: which driver types dominate, which effect directions appear, and whether confidence scores look plausible.

These plots do not validate the ecology. They help identify issues worth checking in the review table.

<Figure size 900x500 with 1 Axes>
<Figure size 900x500 with 1 Axes>

Aggregate sentence-level evidence into graph edges

The model works at sentence level, but a graph edge should usually aggregate repeated evidence for the same species, source and target. This section combines compatible candidate rows, keeps examples of the raw phrases and quotes, and records the number of supporting sentences.

Aggregation is deliberately conservative. It should make review easier without hiding where the evidence came from.

Aggregated DAG edges: 29
Loading...

Add transparent bridge edges

Some extracted relationships point to intermediate concepts such as habitat suitability or population performance rather than directly to abundance. This section adds simple bridge edges from those intermediate nodes to abundance.

These bridge edges are marked as assumptions. They make the graph easier to inspect, but they should still be checked during expert review.

DAG edges after optional bridge edges: 35
Loading...

Build graph objects and check for cycles

This section turns the aggregated edge table into one graph per species and checks whether the graph is acyclic. The graph representation is mainly a way to organise review.

Acyclicity is useful because the intended downstream use is a candidate DAG. However, an automatically produced cycle is not necessarily an ecological failure. It may show that the schema needs refinement or that a bidirectional process has been flattened into directed edges.

Loading...

Visualise one species graph

The plot below is deliberately closer to a hand-drawn causal diagram than to a generic network plot. abundance is placed as the focal outcome and short driver labels are arranged around it. Positive and negative relationships are shown differently because the sign is one of the most useful review fields.

The full evidence phrases remain in the review table. The shortened node labels are only used to make the figure readable.

Saved: outputs_plant_atlas_llm_to_dag/visualisations/Lagurus_ovatus_candidate_dag.png
<Figure size 1000x650 with 1 Axes>
Saved 5 DAG plot(s) to outputs_plant_atlas_llm_to_dag/visualisations

Expert-review table

This table is the main hand-off from the automated workflow to ecological review. It keeps the species, source node, target node, effect direction, evidence quotes, confidence summaries and assumption flags in one place.

The LLM output is not treated as final. The table is designed so that a reviewer can accept, edit, merge or reject candidate edges while seeing the source evidence.

Loading...
Expert-review CSV: outputs_plant_atlas_llm_to_dag/expert_review_edges.csv

Run summary and metadata

The summary table gives one row per species, and the metadata JSON records the key settings used in the run. This is important for LLM-assisted workflows because changes in the model, prompt, token limit or batch settings can change the extracted records.

Loading...
Notebook elapsed time: 7.87 minutes
Run metadata: outputs_plant_atlas_llm_to_dag/run_metadata.json

Output files

The notebook writes the following files:

  • species_text_table.csv: species-level source text used in the run;

  • all_sentences.csv: sentence-level evidence units with neighbouring context;

  • baseline_deterministic_candidate_edges.csv: simple rule-based extraction used only as a sanity check;

  • all_raw_candidate_edges.csv: parsed LLM candidate edges;

  • llm_extraction_audit.csv: sentence-level LLM parsing audit;

  • all_candidate_edges_validated.csv: candidate rows with validation flags;

  • all_candidate_edges_clustered.csv: candidate rows with simple phrase-cluster identifiers;

  • driver_semantic_clusters.csv: one row per simple driver phrase cluster;

  • all_dag_edges.csv: aggregated sentence-level evidence as graph edges;

  • all_dag_edges_with_bridge_edges.csv: aggregated edges plus transparent bridge edges to abundance;

  • dag_diagnostics.csv: node and edge counts and cycle checks;

  • all_dag_nodes.csv: graph nodes and node types;

  • expert_review_edges.csv: the main review hand-off table;

  • run_summary.csv: one-row-per-species summary;

  • run_metadata.json: model path, generation settings, run hash and development hardware metadata;

  • visualisations/*.png: simple graph plots for species with extracted edges.

The review table is the main product. The graph plots and diagnostics help spot obvious problems before detailed ecological review.

Conclusions

This notebook shows a practical route from species-account prose to evidence-linked candidate graph edges using a small local LLM. The design keeps the model task narrow: one sentence, one prompt, one JSON schema and explicit validation afterwards.

For the synthetic Plant Atlas-style example, the output should be read as a candidate review table rather than accepted ecological knowledge. The model proposes rows; the analyst decides whether the evidence is useful, whether the direction is plausible and whether the proposed edge belongs in a causal diagram.

The workflow is intentionally modular. If a different local model is used, only the model-loading and prompt-tuning parts should need attention. If a different text source is used, the same sentence-grounded extraction and validation pattern may still be useful.

References

Boyd, R. J., Botham, M., Dennis, E., Fox, R., Harrower, C., Middlebrook, I., Roy, D. B., & Pescott, O. L. (2025). Using causal diagrams and superpopulation models to correct geographic biases in biodiversity monitoring data. Methods in Ecology and Evolution. Boyd et al. (2025)

Botanical Society of Britain and Ireland. (2023). Plant Atlas 2020. https://plantatlas2020.org/

Hugging Face. Transformers documentation: bitsandbytes quantization. https://huggingface.co/docs/transformers/quantization/bitsandbytes

Hugging Face. Transformers installation/offline mode documentation. https://huggingface.co/docs/transformers/installation#offline-mode

Torrance, A. W. (2024). The environmental impacts of large language models. Scientific Reports, 14, 25020. Ren et al. (2024)

References
  1. Boyd, R. J., Botham, M., Dennis, E., Fox, R., Harrower, C., Middlebrook, I., Roy, D. B., & Pescott, O. L. (2025). Using causal diagrams and superpopulation models to correct geographic biases in biodiversity monitoring data. Methods in Ecology and Evolution, 16(2), 332–344. 10.1111/2041-210x.14492
  2. Ren, S., Tomlinson, B., Black, R. W., & Torrance, A. W. (2024). Reconciling the contrasting narratives on the environmental impact of large language models. Scientific Reports, 14(1). 10.1038/s41598-024-76682-6