Technical Documentation

How HurricaneVuln Is Built

The data behind every number, where each file comes from, how it is cleaned and joined, the wind-field and statistical algorithms, how the models are validated, and every fitted parameter. Figures on this page are read from the same build as the product, so they cannot drift apart.

Version Built Best-track file Buildings Households Storms

1Overview

HurricaneVuln answers two questions for a single building and a single storm. Will the building be damaged, or destroyed? That is the fragility model, fitted at building level to Hurricane Michael. When a home is damaged, what does it cost to repair? That is the household loss model, fitted to FEMA inspections across storms. Both take the same input, the peak sustained wind at the building, which a parametric wind field computes from the official best track.

The product is a vulnerability component. It does not simulate where storms go or how often; that belongs to a hazard model or an event set. Given a wind at a building, HurricaneVuln returns probabilities and loss distributions, each with an uncertainty range.

Outputs

2Data Sources

Eight public sources. Each is retrieved by a script in scripts/, and data/raw/manifest.json records the URL, the retrieval time and the SHA-256 digest of the file as downloaded. Re-running a fetch script and comparing digests shows whether a provider has changed the data since this build.

3Data Dictionary

The two analysis tables the models are fitted to. Types are as stored after processing; distributions are computed on the full table in this build.

3.1 Fragility Table

One row per building in Hurricane Michael's assessed footprint. File: data/processed/fragility.csv, built by scripts/build_fragility.py.

3.2 Household Loss Table

One row per owner-occupied primary residence that registered with FEMA and received an inspection. File: data/processed/household_loss.csv, built by scripts/build_loss.py.

3.3 Category Levels

Counts of every level of every categorical field.

4Processing Pipeline

4.1 Building the Fragility Table

  1. Assessed tiles. FEMA's remote-sensing wind points for Hurricane Michael are binned into 0.1° tiles. A tile is assessed if it holds at least 5 points. Every NSI structure in an assessed tile is pulled from the NSI API by bounding polygon.
  2. One record per building. NSI repeats a footprint once per occupancy inside it. Records are de-duplicated on the NSI building identifier bid, keeping the record with the highest structure value.
  3. Damage labels. Buildings and FEMA points are projected to a local equirectangular plane (metres). A k-d tree finds, for each FEMA point, the nearest building; matches within 30 m are kept. Each building takes the most severe state among its matched points: affected or minor map to 1, major or destroyed to 2. Buildings with no matched point are labelled 0, no visible damage (assumption A-01).
  4. Wind. The peak sustained wind of Michael's track is computed at every building (section 5).
  5. County. Each building takes the county of the nearest 2020 block-group population centre.
  6. Parcel attributes. Florida buildings are placed in their parcel polygon with a spatial index (STR-tree, point-in-polygon). The parcel supplies actual year built, land-use code, construction class and quality grade. Where the parcel gives a construction class, it replaces NSI's.
  7. Post-storm rule. The roll is the 2026 roll. A building whose parcel shows an actual year built of 2019 or later and has a damage point is a rebuild: it keeps its outcome and its attributes are set to unknown. One with no damage point did not exist when Michael struck and is removed (assumption A-04).

4.2 Building the Household Loss Table

  1. Registrations. For each disaster declaration, OpenFEMA is queried for owner registrations with a returned inspection (ownRent eq 'O' and inspnReturned eq true). Rows are ordered by id, a random UUID, so a pull capped at 60,000 per declaration is a simple random sample of it.
  2. Admissibility. Primary residences with a 12-character block-group GEOID are kept.
  3. Location. Each household is placed at its block group's population centre, using the 2010 or 2020 centre file as its censusYear says.
  4. Wind. The storm's peak sustained wind is computed once per block group and assigned to every household in it.
  5. Dollars. Verified real property loss and roof damage amount are multiplied by CPI-U(2024) / CPI-U(storm year).
  6. Flood. A household is flooded if FEMA recorded flood damage or a water level above zero. Water level is in inches above the floor.

Registrations by Disaster Declaration

Consumer Price Index Used

5Wind Field

A parametric field turns the best track into a peak wind at any location. It is a Holland (1980) radial profile, scaled so its maximum equals the best-track intensity, with the forward motion added on the right of the track.

5.1 Track Interpolation

Best-track records are linearly interpolated in time to a 15-minute grid. Records off the 6-hour cycle, such as landfalls, are used as interpolation knots. The radius of maximum wind is interpolated only between records that carry one; elsewhere it is missing. Forward speed and heading come from central differences of the interpolated positions:

$$\Delta y_i=\tfrac12(\phi_{i+1}-\phi_{i-1})R_E,\qquad \Delta x_i=\tfrac12(\lambda_{i+1}-\lambda_{i-1})R_E\cos\phi_i$$ $$V_{t,i}=\frac{\sqrt{\Delta x_i^2+\Delta y_i^2}}{\Delta t},\qquad \theta_i=\operatorname{atan2}(\Delta x_i,\Delta y_i)$$

with latitude \(\phi\) and longitude \(\lambda\) in radians, \(R_E=6371\) km and \(\Delta t=900\) s. One-sided differences are used at the ends of the track.

5.2 Radial Profile

$$S(r)=\sqrt{\left(\frac{R_m}{r}\right)^{B}\exp\!\left[1-\left(\frac{R_m}{r}\right)^{B}\right]},\qquad S(R_m)=1$$

The radius of maximum wind \(R_m\) is the best-track value when HURDAT2 has one (from 2021, and in reanalysed records), converted from nautical miles. Otherwise it follows Willoughby, Darling and Rahn (2006):

$$R_m\,[\text{km}]=46.4\,\exp\left(-0.0155\,V_{\max}+0.0169\,|\phi|\right)$$

with \(V_{\max}\) in m/s and \(\phi\) in degrees. The shape parameter follows Vickery and Wadhera (2008), clipped to \([0.8,\,2.2]\):

$$B=1.881-0.00557\,R_m-0.01097\,|\phi|$$

5.3 Asymmetry and Peak Wind

The best-track intensity includes the storm's motion. Half the forward speed is removed to get the symmetric part, and added back as a cosine of the bearing from the centre, peaking 90° to the right of the heading:

$$V_{\text{sym}}=\max\!\left(V_{\max}-\tfrac12V_t,\,0\right)$$ $$V(r,\psi)=\left[V_{\text{sym}}+\tfrac12V_t\cos\!\left(\psi-\theta-\tfrac{\pi}{2}\right)\right]S(r)$$ $$V_{\text{peak}}(x)=\max_{i}V\!\left(r_i(x),\psi_i(x)\right)$$

Here \(\psi\) is the bearing from the centre to the site, clockwise from north. Distances use the local equirectangular approximation with \(\cos\phi\) of the track point; \(r\) is floored at 0.5 km. In the browser, steps farther than 900 km from the site are skipped, where \(S\) is negligible. The result is a peak 1-minute sustained wind at 10 m in open terrain, the same reference as the best track.

5.4 Check Against FEMA's Wind Classes

FEMA labels each Michael damage point with a wind class. The modelled wind at those points rises with FEMA's class, and runs about one category higher. That is consistent with ground roughness, which FEMA's classes evidently include and an open-terrain reference does not.

6Building Fragility Model

6.1 Model Form

Two logistic stages, both with a partially pooled effect for the assessment county \(c\):

$$\Pr(D\ge 1\mid x)=\sigma\!\left(\mu_1+x^\top\beta_1+u_{1c}\right),\qquad u_{1c}\sim\mathcal N(0,\tau_1^2)$$ $$\Pr(D=2\mid D\ge1,x)=\sigma\!\left(\mu_2+x^\top\beta_2+u_{2c}\right),\qquad u_{2c}\sim\mathcal N(0,\tau_2^2)$$ $$\Pr(D=2\mid x)=\Pr(D\ge1\mid x)\,\Pr(D=2\mid D\ge1,x)$$

\(\sigma\) is the logistic function. The product form guarantees the chance of destruction never exceeds the chance of damage. Wind enters as \(\log(V/50)\) with \(V\) in m/s, so each stage is a fragility curve that is logistic in log wind and monotone by construction.

6.2 Design Matrix

Categorical attributes are one-hot encoded against a reference level. A level enters only if its evidence grade is supported (section 6.5); other levels are merged into the reference. An unknown value never has its own parameter. It is replaced by the training frequencies of the known levels, so it contributes the average effect:

$$x_{j\ell}=\begin{cases}\mathbb 1[\text{value}=\ell] & \text{known}\\ \hat p_{j\ell}=\dfrac{n_{j\ell}}{\sum_{k}n_{jk}} & \text{unknown}\end{cases}$$

Numeric terms are standardised with the training mean and standard deviation. The table lists every column, reference levels and standardisation constants.

6.3 Estimation

For fixed \(\tau^2\), the coefficients and county effects maximise the penalised log-likelihood, with a ridge penalty \(\lambda=1\) on \(\beta\) and the Gaussian prior on \(u\):

$$\ell_p(\mu,\beta,u)=\sum_i\left[y_i\eta_i-\log(1+e^{\eta_i})\right]-\frac{\lambda}{2}\lVert\beta\rVert^2-\frac{1}{2\tau^2}\sum_c u_c^2$$

It is maximised with L-BFGS-B using the analytic gradient. \(\tau^2\) is then updated by a Laplace-approximate EM step (Breslow and Clayton, 1993) and the cycle repeats until \(\tau^2\) changes by less than \(10^{-4}\), or 15 iterations:

$$h_c=\sum_{i\in c}\hat p_i(1-\hat p_i)+\frac{1}{\tau^2},\qquad \tau^2\leftarrow\frac{1}{C}\sum_c\left(\hat u_c^2+h_c^{-1}\right)$$

Standard errors come from the inverse of the penalised observed information for \((\mu,\beta)\), conditional on the county effects.

6.4 Prediction

For a building whose county effect is unknown (any new location), the probability is averaged over the county distribution with 24-node Gauss-Hermite quadrature:

$$\Pr(D\ge1\mid x)=\int\sigma(\eta+u)\,\varphi(u;0,\tau^2)\,du\approx\frac{1}{\sqrt\pi}\sum_{k=1}^{24}w_k\,\sigma\!\left(\eta+\sqrt{2\tau^2}\,z_k\right)$$

The 80% range shown in the product is \(\sigma(\eta\pm1.2816\,\tau)\). Most of it reflects differences in FEMA's imagery coverage between counties, not the building.

6.5 Attribute Grading

Each attribute level's effect on the chance of damage gets a 90% interval from a bootstrap that resamples 0.1° map tiles with replacement ( replicates, county effects refitted each time). A level is supported when its interval excludes zero and, if engineering review gives an expected direction, the estimate agrees with it. Otherwise it is contested: shown, but merged into the reference in the product.

6.6 Fitted Parameters

Final fit on all buildings with the shipped specification. Estimates are on the logit scale per unit of the (standardised) column.

Stage 1: Any Visible Damage

Stage 2: Destroyed Given Damage

County Effects

7Household Loss Model

7.1 Population and Outcomes

The population is owner-occupied primary residences that registered with FEMA's Individuals and Households Program and received an inspection. The model is therefore conditional on registration and inspection, the analogue of a claim severity model. Outcomes:

  • Verified loss \(L\): FEMA-verified real property loss in 2024 dollars. It measures repairs needed to make the home safe, sanitary and functional, which is below a full insured loss.
  • Uninhabitable: FEMA's inspector found habitability repairs were required.
  • Destroyed: FEMA's destroyed flag.

7.2 Model Form

A hurdle model for loss, plus two logistic models, each with a partially pooled storm effect \(s\):

$$\Pr(L>0\mid x)=\sigma\!\left(\mu_p+x^\top\beta_p+u_{ps}\right)$$ $$\log L\mid L>0,\,x\sim\mathcal N\!\left(\mu_l+x^\top\beta_l+u_{ls},\ \sigma^2\right),\qquad u_{ls}\sim\mathcal N(0,\tau_l^2)$$ $$\Pr(\text{uninhabitable}\mid x)=\sigma\!\left(\mu_h+x^\top\beta_h+u_{hs}\right),\qquad \Pr(\text{destroyed}\mid x)=\sigma\!\left(\mu_d+x^\top\beta_d+u_{ds}\right)$$

7.3 Features

All models share one design. Wind \(V\) is in m/s and \(k=V/0.514444\) in knots; \(w\) is water above the floor in inches, capped at 120; \(R=\mathbb 1[\text{year}\ge2019]\) is the inspection regime.

Why a regime term. At similar winds, FEMA's verified losses are about four times higher from 2019, and the roof-damage flag stops being recorded. That is a change in inspection practice, not in buildings. The regime indicator and its interactions absorb it, and the product predicts with \(R=1\). Adding the interactions cut household CRPS from $3,095 to the value in section 9.

7.4 Estimation

The three logistic models are estimated as in section 6.3, with storm in place of county. The severity model is a Gaussian mixed model fitted by EM. Given the storm effects, \(\beta\) solves a ridge-penalised least-squares problem; given \(\beta\), each storm effect is shrunk towards zero; the variances are then updated:

$$\hat\beta=\left(X^\top X+\lambda\sigma^2 P\right)^{-1}X^\top(z-u_{s(i)}),\qquad P=\operatorname{diag}(0,1,\dots,1)$$ $$\hat u_s=\frac{\tau^2}{\tau^2+\sigma^2/n_s}\,\bar r_s,\qquad \tau^2\leftarrow\frac1S\sum_s\left(\hat u_s^2+\left(\frac{n_s}{\sigma^2}+\frac1{\tau^2}\right)^{-1}\right),\qquad \sigma^2\leftarrow\frac1n\sum_i e_i^2$$

with \(z=\log L\), \(\bar r_s\) the mean residual in storm \(s\), and 30 EM iterations.

7.5 Predictive Distribution

For a new storm the storm effect is unknown. The probability of a positive loss is integrated over \(u_{ps}\) as in section 6.4, and the severity is lognormal with the storm variance added:

$$L\sim(1-p)\,\delta_0+p\cdot\text{LogNormal}\!\left(m,\ s^2\right),\qquad s^2=\sigma^2+\tau_l^2$$ $$Q_L(\alpha\mid L>0)=\exp\!\left(m+s\,\Phi^{-1}(\alpha)\right),\qquad \mathbb E[L\mid L>0]=\exp\!\left(m+\tfrac12s^2\right)$$

The browser evaluates \(\Phi^{-1}\) with Acklam's rational approximation (relative error below \(1.2\times10^{-9}\)).

7.6 Fitted Parameters

Final fit on all households.

Probability of a Positive Verified Loss

Log Verified Loss, Given Positive

Uninhabitable

Destroyed

Storm Effects

8Combining the Models

For a building with attributes \(x\) at peak wind \(V\), the product reports:

$$p_{\text{dmg}}=\Pr(D\ge1\mid x,V),\qquad p_{\text{des}}=\Pr(D=2\mid x,V)$$ $$\mathbb E[\text{verified loss}]=p_{\text{dmg}}\cdot\mathbb E[L\mid L>0,\,h(x),V]$$

where \(h(x)\) maps the building to a household profile. The loss part is shown for residential uses only. The habitability probability is reported as it is estimated: for a home that FEMA inspects.

In the portfolio desk, each building's wind is the peak of the chosen storm's full track at the building's coordinates. Book totals are sums of expected values. The uninhabitable-and-uninsured count is \(\sum p_{\text{dmg}}\Pr(\text{uninhabitable}\mid\text{inspected})\) over uninsured residential buildings.

9Validation

9.1 Schemes

  • Fragility, spatial tiles. The 0.1° tiles are randomly split into 5 folds (seed 7). Each fold is predicted by models fitted to the other four, with county effects known. This tests new neighbourhoods inside assessed counties.
  • Fragility, leave one county out. Each county with at least 1,000 buildings is withheld in turn and predicted with its county effect integrated out. This tests a new assessment area.
  • Household loss, leave one storm out. Each of the storms is withheld in turn; models are trained on the rest and scored on up to 20,000 sampled households of the withheld storm (seed 0).

9.2 Metrics

$$\text{Brier}=\frac1n\sum_i(p_i-y_i)^2,\qquad \text{Log loss}=-\frac1n\sum_i\left[y_i\log p_i+(1-y_i)\log(1-p_i)\right]$$ $$\text{CRPS}(F,y)=\int_{-\infty}^{\infty}\left(F(z)-\mathbb 1[y\le z]\right)^2dz=2\int_0^1\left(\mathbb 1[y

CRPS is computed from 99 quantiles at \(\alpha=(k-0.5)/99\), the quantile-score form of Gneiting and Ranjan (2011). AUC is the area under the ROC curve. Share error is the absolute difference between the mean predicted and observed rates in a fold. 80% coverage is the share of observations inside the 10th–90th percentile range. For storm-level ranges, the storm effects of the frequency and severity models are set jointly to \(\pm1.2816\,\tau\).

9.3 Baselines

Gradient boosting never sees an unknown value as a separate signal: unknown categorical fields are replaced by random draws from the known levels of the training data before fitting and scoring.

9.4 Results

Fragility

Fragility, Withheld Counties

Household Loss

Household Loss by Storm

10Implementation

10.1 Browser Engine Parity

The website carries its own copy of the wind field and both models in JavaScript. The check below runs that engine in your browser now, against reference values the Python models produced at build time.

10.2 Reproduce

python scripts/fetch_public.py      # HURDAT2, block-group centres, FEMA damage points
python scripts/fetch_nsi.py         # NSI structures for assessed tiles
python scripts/fetch_parcels.py     # Florida parcel roll, nine counties
python scripts/fetch_ihp.py         # OpenFEMA registrations, 19 storms
python scripts/build_fragility.py
python scripts/build_loss.py
python experiments/run_fragility.py
python experiments/run_loss_validation.py
python scripts/export_to_web.py
python scripts/build_site.py
python tests/test_core.py
node tests/test_js_parity.js

Python 3.11 with numpy, scipy, pandas, scikit-learn, shapely and joblib. Node 18 or later for the parity test.

11Assumptions and Limitations

12Glossary

13References