About PowderBench

A benchmark you can't train on, scored against mountains that don't care about your model architecture.

Why this exists

Most benchmarks die the same death: their answers leak into training data and the leaderboard turns into a memory test. PowderBench is contamination-free by construction — every round asks about weather that hasn't happened yet. There is nothing to memorize, nothing to overfit, and no way to peek. You, your ML pipeline, or your AI agent make a real forecast, the atmosphere runs its own compute, and a few days later the mountains grade your work.

It's also just fun. Snowfall is one of the hardest quantities in weather forecasting — orographic effects, rain-snow lines, wind loading — and it comes with a built-in villain: the global weather models submit automatically every round. The whole site is one long dare: nobody is beating the weather models. Yet.

Three leagues, three kinds of truth

Every league is named for its ground truth, because the truth is the game:

One set of rules everywhere: daily rounds, 24/48/72-hour horizons, Powder Score against climatology, QC that voids bad data for everyone equally.

Fair by design

How resort scraping works (and its manners)

The Resorts league depends on archiving snow reports that vanish daily, so a scraper visits each resort's own website twice a day. It behaves like a good guest, and the policy is enforced in code, not just promised:

Data sources & licensing

The benchmark runs entirely on GitHub Actions: rounds open, baselines submit, reports get archived, truth gets fetched, scores resolve, and this site redeploys — every day, in public.

Contact

Questions, bugs, a resort that wants in (or out), a station we should add? Open an issue. Bug reporters get eternal glory in the README.