Contentinental Divide Trail Water Data Analysis

Good water data and uncertainty about upcoming water sources dictates how much you carry, and how dehydrated you might end up being. I built a model of available CDT water information to provide better statistics and probability distributions for water across the trail. This page explains what is in the model, how it works, and how well it works. The explorer and the app can be found on my CDT page.

Disclosure: Generative AI was utilized to assist in model fitting and parameterization of the model to best estimate water characteristics given the data sets that I obtained.

Model Flow

  1. Water Observations What hikers wrote about water at mapped sources
  2. Scoring Methods Using a language model to quantify abundance, quality, and sentiment from hiker descriptions
  3. Typical-year modelSeason, a fixed effect for the year, and pooling across neighboring water sources
  4. ProbabilitiesChance of dry or flowing for any source for a specific date
  5. Most likely to have Water WaypointsWater points I’ve put on a Garmin course, selected to be optimal for reducing water carry length and likelihood of finding water

Information about underlying data

The raw data a corpus of text, and date stamps, that hikers have written about water at a mapped sources along the CDT. Things like “Dry.”, “Flowing, clear and cold”, or “Cow pond, murky, but it is wet”. Most entries are a handful of words, and there is no real quantification about quality or quantity. This is the first problem that comes from establishing a numerical dataset to assess water quality: how do you convert this into a numerical evaluation?

Step 1: Turning words into numbers

I’ve given every entry three scores:

  • Quantity on a five-point likert scale, from dry to abundant (see below).
  • Quality as “poor”, “mixed”, or “good”, based on textual context in relevant posts (clear, cold, murky, cows, algae). Most entries do not mention water quality, so there is quite a lot of prediction going into this metric.
  • Sentiment ranges from very negative to very positive. This scoring is effectively an amalgamation of the other two. A flowing source with clean water reads as happy, a dry one reads as not — gives an overall vibe check on whether or not folks are using the source.
  • 0Dry
  • 1Stagnant or pools
  • 2Trickle
  • 3Flowing
  • 4Abundant

I used a set of text-based rules to handle obvious phrasing (“dry”, “no flow”, “trickle”, “plenty”), and employed a small language model to turn each sentence into numbers to better assess water quality and sentiment. I hand checked a limited subset of the sample to see how well text connotation was converted into meaningful quantified numbers. Quantity metrics were accurate about 85% of the time, which I think is good enough for averages and the purpose of this project!

Step 2: Building a typical year for every source

Each scored entry of text is used as one data point in the model: a source s, with a year y, a day of the year d, and a score. I wanted to create a smooth function of date relative to quantity and quality for every source, with the year-to-year weather effect pulled out so what remains is an expected value based on the season. I used this structure to fit four outcomes: expected quantity of water, chance of the source being dry, water quality, and overall sentiment.

Constructing each season

Day of year gets rescaled to a position in the general CDT hiking season, t, and entries outside March 16 to November 16 are dropped.

t = clip( (d − 75) / 245 , 0 , 1 )d = 75 is Mar 16, d = 320 is Nov 16

The linear predictor

Everything runs through one linear predictor, η, which is the same for all four outcomes. Only the link and the loss change.

ηs,y(t) = α + γylevel + year effect
    + Σk=14 ak cos(kπt)trail-wide season
    + Σk=13 bβ(s),k cos(kπt)50-mile bin departure
    + δs + Σk=13 cs,k cos(kπt)source offset and shape
  • α + γy: a global intercept plus a fixed effect for each year. Entries from 2016 and earlier are pooled into a single level because the early years are thin.
  • ak: the trail-wide seasonal curve, four cosine terms (effectively the average quantity of water on the trail over time).
  • bβ(s),k: how the 50-mile bin containing source s departs from the trail-wide curve (each 50-mile section of the trail may exhibit a difference in overall quality from the trail as a whole). β(s) = [mile / 50], using the mile snapped onto the route.
  • δs and cs,k: the source’s own offset and its own departure from its bin curve (how the source varies compared to the 50-mile segment of trail which pools overall effects of that section together)

Why the cosine terms? For the range [0, 1], cos(kπt) is a half-range cosine basis. It has zero slope at both ends and no periodic wraparound, which is the right behavior here because March and November are not the same season. The cosine basis gets a smooth seasonal shape out of very few parameters per block.

Outcomes, links, and loss

  • Quantity, quality, sentiment: Gaussian with identity link, so E[score] = η, fit by penalized least squares.
  • Chance of dry: binary outcome y = 1 when the quantity score is 0. Logit link, penalized maximum likelihood. This is the complement of the k = 1 threshold fit below, same data and same penalties.
  • Quantity distribution: four more logistic fits, one for each threshold k = 1…4, with outcome 1[q ≥ k]. Each gets its own η(k).
P(dry) = 1 / (1 + e−η)chance of dry
P(q ≥ k) = 1 / (1 + e−η(k))k = 1, 2, 3, 4
P(q = k) = P(q ≥ k) − P(q ≥ k + 1)the bars in the app

The four threshold fits are independent, so they can cross each other in places where the data are limited. I force monotonicity afterward by taking a running minimum across k, so P(q ≥ k) never increases with k.

Shrinkage: how the pooling works

The year and trail-wide terms are effectively unpenalized (λ = 0.01). The bin, source-offset, and source-shape terms get ridge penalties, which is where the pooling comes from. The objective is:

minθ   ℓ(y, η) + λbin Σ b² + λδ Σ δ² + λc Σ c² + 0.01 ( Σ γ² + Σ a² )

Here ℓ is squared error for the Gaussian outcomes and the negative log-likelihood for the logistic ones (the logistic version carries an extra factor of ½ on the penalty, a scikit-learn convention). I used λbin = 32, λδ = 2, λc = 10 for quantity and sentiment, and λc = 20 for quality. The logistic fits use λbin = 8, λδ = 0.25, λc = 2. The logistic scale is different, so the weights are smaller.

The Bayesian reading is that each ridge penalty is a zero-mean Gaussian prior with variance σ²/λ on that block, and the fit is the posterior mode. This is partial pooling, in the same spirit as a random-effects model. A source with few entries has its offset and shape shrunk toward its bin, the bin toward the trail, and a source with plenty of entries keeps its own shape. It is not a full mixed model. The variance components are not estimated by REML.

The typical year and the bands

To draw a typical year, I replace γy with the entry-weighted mean of the fitted year effects, then evaluate the curve on a weekly grid:

γ̄ = Σy ny γy / Σy nyny = entries in year y
p̄s(t) = 1 / (1 + e−η̄s(t))η̄s uses γ̄ in place of γy

The shaded bands come from a Laplace-style approximation on each source’s own block, θs = (δs, cs,1, cs,2, cs,3). Conditioning on everything else, its covariance is Σs = (Φ′WΦ + Λ)−1, scaled by the residual variance σ² for the Gaussian outcomes. Φ has rows φ(t) = (1, cos πt, cos 2πt, cos 3πt) for each of the source’s entries. Λ = diag(λδ, λc, λc, λc). For the logistic fits, W = diag(p(1 − p)). The band is η̄ ± 1.28 · SD, mapped through the logistic for probabilities.

Var[η̄s(t)] = φ(t)′ Σs φ(t)band = η̄ ± 1.28 √Var   (80%)

Two things make these bands narrower than they should be: they ignore uncertainty in the year, bin, and trail-wide terms, and they treat entries as independent given the model. Same-day entries at the same source are correlated, so the effective sample size is smaller than the raw count. The year-effect whiskers in the chart below come from a bootstrap, not from this approximation.

Summarizing the analysis

1. The year is a fixed effect. Without it, a big-snow year makes June look wet and a drought year makes it look dry, and the curve ends up learning the weather instead of the season. The model gives each year its own adjustment and then removes it, so what is left is an expectation for the season.

2. Sources in proximity to one another influence each other. Some sources might have very few comments. In those cases, its assumed to look more like the sources around it in the 50-mile stretch. In turn, each stretch is weighted to look like the trail in general, where the more data given in a stretch shofts the trail toward being more unique. That means that sources with a lot of information will have their own shape.

3. Some uncertainty bounds are generated and shown. The shaded band on each chart is an approximate 80% range, and it widens when there is less data about a given water source.

What the model results show

For a typical year and a northbound schedule, here is how the sections break down. Each row is an average over the mapped sources in that section.

SectionChance a source is dryChance it is flowing or better
New Mexico19%64%
Colorado5%82%
Wyoming8%74%
Southern Montana / Idaho12%66%
Northern Montana14%67%

Colorado has the best likelihood for reliable water, while New Mexico has the most suspect sources.

Wet and dry years do influence results, but far less than you fear

0%5%10%15%20%Dashed line: an average year (11% chance a source is dry)13%≤20168%201712%20188%201911%202013%202110%20228%20238%202414%202516%2026
Chance that a typical source is dry, by year, with the season and the source held constant. Whiskers are 95% ranges. Green is a better-than-average year, orange is worse. The first bar pools the earliest years with limited data.

The big-snow years of 2019 and 2023 were the best, with a typical source being dry only about 7% to 8% of the time. 2026 was the worst year on record, with about 16% of sources across the trail being dry. Even in that bad year, the average source was dry roughly one time in six.

Which sources have high variability

Some sources behave the same every year. Others may be super unpredictable and unreliable, changing between dry and flowing, depending on the year.

  • Close to half of all water sources with sufficient hiker commentary show consistent water trends, year after year. These sources are reliably flowing (or dry).
  • About 10% of sources fluctuate significantly from one year to the next, and their chance of being dry increases by roughly 30 percentage points.
  • These highly variable sources are more often lower in elevation, more often ponds, stock tanks, and seasonal creeks, and New Mexico has most of them.
  • “Usually dry” is not the same as “higly variable”. A source that is always dry is predictable.

Can the weather help predict highly variable sources?

This was one of the key questions that I had going into this. Can we better predict the probability of a source being wet or dry using other data sources. I looked at gridded snowpack data from NOAA’s snow model as well as daily flow from USGS stream gauges near the trail, including lags so that melt timing was considered.

What I addedBetter quantity predictionBetter chance-of-dry prediction
Snowpack, with lags and melt timing0.4%0.6%
Stream gauge flow, with lags1.7%2.8%
Stream gauge flow, forecast six weeks ahead1.0%1.7%

The ability to predict water quantity using the stream data was better than the snowpack data, but neither meaningfully improve our ability to predict the water quantity for variable sources.

How the Garmin water points were chosen

The garmin courses I’ve added to my CDT page have water points that were selected using an optimization algorithm. The optimizer balances three things: a high chance the source really has water, short carries between points, and as few points as possible, so your watch does not turn into an annoying barage of endless waypoint notifications.

  • Reliability was chosen based on the chance of a given source hacing at least a trickle on the worst date within three weeks of when a northbound hiker typicallly arrives in a typical year.
  • Water carries were chosen to be between every 2 and 5 miles, and are de-priortized if they are further than 5 miles, and even further deprirotized if they are 8 and 15 miles away.
  • Each point has a “cost” for being chosen by the optimizer, and a chance of being dry adds an expected extra carry plus a reliability penalty, so the optimizer avoids highly variable sources.

The waypoints I added to the course yields about 16 points per 100 miles with a median gap of about 5 miles. A name starting with ~ means a source with decent odds (60 to 85%). A ? means it is the best available option in a thin stretch and may be dry.

Model limitations and considerations

  • Data represent a typical year, not today. It tells you the odds, not the conditions.
  • The model only has a limited subset of sources. An unmapped spring does not exist in the model, and a mapped one in the wrong place can be misleading. Where the course runs along a road or an alternate with nothing mapped nearby, long gaps could reflect missing data and not necessarily dry conditions.
  • Managed sources are different. Tanks, windmills, pumps, and caches depend on people, not weather. From the data, weather had very little ability to predict what would happen for these sources. (Thanks New Mexico!)
  • The scoring is not percect. I tried to turn hiker text into meaningful, quantified scores. Some differences in phrasing and words can mean different things to different people, and the language-to-number step makes mistakes.

Making Use of this Analysis

  • The explorer: a heatmap of the whole trail and a seasonal curve for every source.
  • The CDT Water app: a map, your position, and the next water ahead of you, working offline.
  • The Garmin courses: every leg with the water points built in.