The hail data
Every hail number on this site comes from one place: the
NOAA
National Centers for Environmental Information Severe Weather Data Inventory, specifically the
nx3hail dataset of NEXRAD Level-III hail signatures. It is free, it is public domain,
and it needs no API key.
We download the ten complete annual NOAA SWDI v2 bulk CSV archives for 2016-01-01 through 2025-12-31. NOAA labels these files experimental. We retain records with a positive estimated size and 100% algorithmic hail probability. This probability is not certainty that hail reached the ground.
For each of our 50 cities, we calculate great-circle distance from every candidate signature to the city centre and keep only those within 25 miles. Duplicate observations with the same radar, scan time, cell and coordinates are counted once. Records are grouped by UTC calendar day, retaining the original maximum size in inches, maximum severe-hail probability, signature count and distance range. A day can cover multiple storm cells. Overlapping city areas can share detections; adding their totals does not count unique US storms.
The average uses all ten calendar years, including years with no records. Missing source files, invalid records and failed downloads block publication. A successful import confirms source-file completeness, not complete radar coverage. Absence of a signature does not prove absence of hail.
Data are processed before publication and served from our own site. Source files, checksums and processing rules make this archive traceable. It is a historical record, not a live feed. Retrieved 2026-09-29; the latest covered date is 2025-12-31.
Raw files: /data/index.json,
/data/zip-index.json, and one file per metro under
/data/metros/.
How ZIP codes are matched to a metro
We map ZIP codes to metros by their three-digit prefix. A ZIP3 covers a sectional centre facility area, which is a good deal larger than a city, so this is an approximation and we say so on every result. If a prefix straddles two of our metros we ask you which one you are nearer rather than guessing.
The honest limitation: a ZIP3 match tells you a storm was somewhere within 25 miles of the metro centroid, not that it was over your street. When we can load ZIP Code Tabulation Area centroids from the Census Gazetteer we will query per ZIP instead, and this section will change.
41,000 near-identical pages built from data that is really metro-level would be exactly the kind of thin programmatic SEO that has been getting sites removed from search results. Fifty pages with real, different numbers on them are worth more than forty-one thousand pages with the same sentence and a different place name.
What the radar data cannot tell you
MAXSIZE is the radar algorithm's estimate of the largest stone it believes was aloft in a storm cell. Four things follow from that, and all four matter:
- It is aloft, not on the ground. Stones melt, break up and get carried. What lands is usually smaller than what the algorithm saw, sometimes much smaller.
- It is the cell, not your house. Hail swaths are narrow and sharply bounded. One street can be destroyed and the next one untouched.
- It is an estimate from reflectivity, not a measurement. The algorithm has known biases, and radar coverage quality varies with distance from the radar site and with terrain.
- Absence of a signature is not proof of absence of hail. Low-topped storms and areas far from a radar can produce hail that is never flagged.
Which is why every result page on this site says the same thing: use it to decide whether it is worth looking, then look.
The RoofTake Score
The 0–100 number on every area page is a composite index of how much hail an area has taken. It is a model, not data, and it describes an area rather than a house — two roofs on the same street of different ages face very different risk from the identical storm.
A timeline of sixty-seven dates is accurate and unshareable; a single number is shareable and, on its own, misleading. So the score never appears without the four components that produced it, and here is the whole formula:
score = 100 × (0.35 × frequency + 0.30 × severity
+ 0.20 × recency + 0.15 × concentration)
| Component | What it measures | Weight | Scale |
|---|---|---|---|
| Frequency | Hail days per year at 1.00 in (2.5 cm) or larger | 0.35 | 0 at none, 1.0 at 40 days a year |
| Severity | Largest stone radar estimated in the window | 0.30 | 0 at 0.75 in, 1.0 at 4.00 in |
| Recency | Years since the last 2 in (5 cm) day | 0.20 | exponential decay, half-life 3 years |
| Concentration | Hail days per year at golf-ball size or larger | 0.15 | 0 at none, 1.0 at 20 days a year |
Bands: 0–25 low, 26–50 moderate, 51–75 high, 76–100 very high.
The weights are a judgement call and we state them as one. Frequency leads because roofs fail cumulatively, not from one storm. Recency matters because claims have deadlines. There is no statistical optimisation behind these four numbers — there is no labelled dataset of "roofs that needed replacing" to fit them against, and anyone who implies otherwise is overselling.
The scale thresholds are editorial choices, not a validated damage model or a national percentile. In version 1.03 we revised frequency and concentration thresholds from 9 and 3.5 days per year to 40 and 20. The actual 2016–2025 records across these 50 metros range from 9.7 to 32.2 days per year at 1 inch and from 3.1 to 15.8 days at 1.75 inches. Rounded thresholds above those observed maxima leave headroom and prevent almost every city scoring the maximum. Scores from older demonstration builds are not comparable. Recency is measured from the end of the archive (2025-12-31); an area with no 2-inch event gets zero for that component. This score describes the historical window and does not assess current roof condition. Rank compares only our 50 selected metros.
You can recompute the score from our public JSON using the formula above.
The roof condition quiz scoring model
The quiz is a simple additive model. Each answer carries a fixed number
of points; the total is divided by the maximum possible and expressed out of 100. Bands are: under
36 no strong signal, 36 to 61 worth an inspection, 62 and above get it looked at now. The full
scoring table is in assets/js/quiz.js, which is unminified on purpose so you can read
it. It is a triage heuristic built from the questions a roof inspector asks first. It is not a
diagnosis and it does not replace somebody getting on the roof.
The remaining-life model
The age calculator starts from a published typical
service life for the material category, then applies two disclosed multipliers: attic ventilation
(0.85 for poor, 1.05 for good) and hail exposure (0.78 for frequent, 0.90 for occasional). Nominal
lives by material are listed in assets/js/calc-roof.js.
Installation quality varies more than product does, and two identical roofs installed the same week can be five years apart in real life. Treat the output as a planning number.
The replacement cost model
There is no free, machine-readable dataset of roof replacement prices by US region. This is the single most important sentence on this page, because it means every "average roof cost in your ZIP" you have ever seen — including ours — is somebody's model.
Ours works like this. Roof surface area is the ground footprint multiplied by the exact geometric
pitch multiplier, plus a 12 percent waste allowance. That gives squares. Squares are multiplied by a
published installed cost range per square for the chosen material, then by a region factor (0.82 to
1.65 across four bands) and a complexity factor (1.00 to 1.70 across four bands). Tear-off is added
at $100 to $190 per square per layer, decking at $85 to $150 per sheet, and permit at $150 to $700.
All of those inputs are visible in assets/js/calc-roof.js and every result page prints
the full breakdown rather than a single number.
It is for deciding whether a bid is in the right universe and for asking better questions about scope. It is not a quote, and a real bid can land outside the range for good reasons — a second layer nobody knew about, decking, code upgrades triggered by the permit, or simply a market that is short of crews after a storm.
The ACV and RCV model
The payout calculator uses straight-line depreciation: age divided by expected life, capped at 100 percent, applied to the replacement cost you enter. Real carriers use their own depreciation tables and apply condition judgements no formula can see. The calculator exists so that the adjuster's worksheet is not a mystery, not to predict your settlement.
The impact outcome model
The comparator maps hail size bands to qualitative outcomes per covering type. Those bands come from the physics and from standard industry practice, not from a dataset of measured outcomes, and the tool says so on every result. The UL 2218 test parameters quoted (steel ball diameter and drop height per class) are the published parameters of the standard.
What we deliberately do not publish
Three things, on purpose:
- Per-state claim deadlines. The deadline that binds you is in your policy, the states differ, and a wrong number on a website could cost somebody their claim. The deadline tracker builds a timeline from your own policy terms and points you at the NAIC directory of state insurance departments to verify.
- Cost figures per ZIP presented as data. See above. We publish a model and label it.
- Free-text reviews of named contractors or insurers. Community content on this site is structured fields only. That is a deliberate choice about defamation exposure and about signal quality, explained under the terms.
Community-submitted content
Neighbourhood reports, claim outcomes and damage photos are submitted by readers. Everything goes into a moderation queue and a human approves it before it appears. We do not ask for names, and we do not ask which insurer paid or denied a claim — the outcome bands are structured so that nobody can be identified and no company can be defamed. Photographs have their location metadata stripped on upload.
Corrections
If a number on this site is wrong, tell us and we will fix it and say what changed. Email hello@rooftake.com. Corrections to data pages are picked up in the next validated archive update; corrections to editorial pages are made immediately and the review date is updated.
Common questions
Can I reuse your data?
Yes. The underlying radar data is US Government public domain. Our aggregation, the per-metro JSON files and the per-year files are free to reuse with a link back to the page you took them from.
How often is it updated?
This is a verified historical archive covering January 2016 through December 2025. It does not include 2026 storms or provide live alerts. The retrieval date is shown separately from the period covered. Updates require a new validated archive build.
What happens when NOAA changes the endpoint?
The build fails loudly rather than publishing stale numbers quietly. That is deliberate.
Do you use AI to write the articles?
The editorial pages are drafted with AI assistance and reviewed against primary sources before publication. Where we could not verify a number, we describe the mechanism instead of inventing a figure — that is why you will see ranges and 'this is a model' notices rather than confident single numbers.