Cairngorm Club journals · 1893–2023

Building a text-dataset of Scotland’s mountain snow

A human-verified record of snow and ice observations recovered from historic mountaineering journals—made searchable by place, season and year.

2,381 curated observations
1618–2018 dated record span
16 mapped regions
96.3% with map coordinates

Cairngorms Western Massif · photograph by Craig Aitchison

Text is an under-utilised asset for Scottish climate change

Two centuries of observations, mined using AI and verified by human volunteers.

Mountaineers recorded snowfields, cornices, avalanches, frozen lochs in journals and diaries well before the deliberate action of the late 20th and 21st centuries. SNOSCOT brings those dispersed accounts together as structured observations while retaining the original language of the source material. We used AI to mine the text (wince), but hand verified all extracted fields against the reference text. This project would not have been possible without AI.

The collection centres on Cairngorm Club journals published from 1893 to 2023 and includes additional historical references reaching back to 1618. It is a record of what writers noticed and what survives in the archive—not a continuous instrument record.

Interactive atlas

Find the observation behind the point.

Search the full text, choose a region, or narrow the dated record. Map points represent broad regional groupings rather than exact observation sites.

Loading observations…

Exploratory data analysis

What is inside the archive?

A descriptive view of when observations appear, where they were assigned, and the language used to characterise snow and ice.

546

named places

Specific mountains, corries, glens and passes appear across the record.

6.36/10

mean snow score

The curation score describes how positive or substantial the snow description is. The "Snow Score" is itself defined and calculated by GPT4o.

144

years represented

Coverage is highly uneven and follows the availability and content of source journals.

April

most-recorded month

465 observations with an April date—the strongest monthly concentration.

01 · Temporal coverage

Observations by decade

count

The high early-20th-century count reflects journal coverage as well as what was observed.

02 · Seasonality

Observations by month

dated rows

Monthly analysis uses 1,997 valid month values.

03 · Geography

Most represented regions

top 10

04 · Curation score

Score distribution

0–10

A high score signals a strongly positive or substantial snow description, not data quality.

05 · Vocabulary

Frequent extracted terms

tag mentions

06 · Data completeness

Useful context, with visible gaps.

The original quotation and regional assignment are almost complete. Exact dates and point-level geography are less consistent in historical prose, so missingness is retained rather than guessed.

From page to observation

Machine-assisted discovery.
Human-verified evidence.

The project was designed to make a very large historical archive tractable without treating model output as ground truth. Language models identified candidate passages; people checked, corrected and contextualised every retained observation.

  1. 01

    Acquire & extract

    Journal PDFs are collected and their pages converted into machine-readable text. This involves a process known as document-chunking where overlapping chunks are generated for downstream extraction. This is because LLMs have a context window

  2. 02

    Identify candidates

    A GPT-4o-family model finds compact passages that explicitly mention snow, ice or related cryosphere conditions. Why GPT-4o? Fundamentally, cost. SimonF92 funded API calls with his own money. Price estimates to redo this with advanced models are ~£150-£250.

  3. 03

    Verify & repair

    Each quotation, entity, score, date and location is reviewed against the source text by a human being; hallucinations are removed, dates are refined and locations are narrowed down as much as possible.

  4. 04

    Map & publish

    Specific place names are reconciled into broad geographic groups and released with the curated output.

Dataset structure

Twelve fields keep the quote and its context together.

The quotation remains the central unit. Derived labels make the archive searchable, while comments preserve decisions that need qualification.

FieldWhat it contains
textThe retained verbatim passage from the source.
entityExtracted snow and ice terms, sometimes comma-separated.
scoreA 0–10 description score, from absent/melted to substantial or strongly positive snow.
dateThe best available original date string, including partial dates.
annotator_commentHuman notes about uncertainty, derivation or historical context.
general_locationA consistent broad mountain region used by the atlas.
specific_locationThe named mountain, corrie, loch, glen or pass.
year · season · month · dayParsed date components where the source supports them.
CoordinatesA representative coordinate for the broad mapped region.

Open research dataset

Collaboration is key

The mined and curated dataset outputs are released under Apache 2.0. Rights in the original journal PDFs remain with the Cairngorm Club; contact the relevant rights holder before reusing source documents themselves. If an academic institution wishes to adopt and refine this work with larger models for a PhD student project then please do get in touch.

Dataset DOI 10.57967/hf/7619

Ben Macdui · photograph by Peter Hudson