Texas crash narratives record what coded fields miss: 15,074 phone-use injury crashes against 7,340 coded
Lead
Across 5,601,890 Texas crashes from 2017 to 2025, a model reading the coded table and a model reading officer narratives were joined under one sampling design to produce population counts with intervals: narratives document 15,074 phone-use injury crashes against 7,340 in the coded field.
Story
Two readings of the same crash record are joined into one estimation system that returns statewide counts, a discordance map, a re-read list, and a reading budget rather than a single prediction. The coded field and the officer narrative were almost never measured against each other under one sampling design, so a safety office could not tell how much its counts missed. On 5,601,890 Texas crashes from 2017 to 2025, with estimates over the 5,018,080 crashes carrying a usable narrative, a second human tier drawn with recorded probabilities agreed with all fifteen estimates within its sampling margin.
The coded table is read for every crash, the narrative is read on probability samples, and human judgments recalibrate the narrative reader's probabilities. Tabular models had served prediction rather than inference, and prediction-powered inference had used a single proxy without crash records. The table reader is Kumo Tabular small with a 60,000-row context; the narrative reader is Jev 1.13.0 reading a 150,000-narrative simple random sample and a second wave of 79,510 narratives drawn on the discordance score; the human tier holds 2,309 usable judgments on 400 narratives.
What to watch
Next steps include repeating the design in other states to show how much discordance is Texas practice, and across versions of both readers to show how stable the estimates are when the readers change. Reading the narratives of crashes without coordinates after a cleaning pass would bring the extrapolated group into the measured population, and a larger human tier drawn by the allocation rule would narrow the intervals the recalibration tier dominates. A relational table reader that sees the crash, unit, and person tables jointly is the natural extension of the coded view.
The narrative reader is a closed commercial model whose training data are not disclosed, so its readings were pinned to one version and audited and recalibrated through two human tiers rather than inspected. Extrapolations to crashes without a usable narrative are model-based by construction and reported apart from every estimate. Vehicle defects has 41 percent of its estimate below the probability threshold of the human check, and hit-and-run and intentional crashes have the widest check intervals, so those rows carry the weakest human agreement evidence. The recalibration tier carries most of the variance, 95 percent for alcohol and 99 percent for wrong-way driving, so the table reader narrows the total half-width only modestly.
