How We Score Safety
The FRED Score is a forward-looking model that predicts a carrier's expected severity-weighted crash burden over the next 12 months, then grades it against similarly-sized peers. It is fit on one year of history to predict the following year — out-of-time validated, not scored on the same data it learned from.
Scoring Pipeline at a Glance
Data Sources
Seven FMCSA datasets feed the scoring pipeline
Every score starts with public data from the Federal Motor Carrier Safety Administration. We pull seven distinct datasets, join them on DOT number, and filter to the population of active for-hire property carriers.
| Dataset | What It Contains | Updates |
|---|---|---|
| Census | Carrier registration, fleet size, address, authority type, officer names | Daily |
| Inspections | Roadside inspections with driver and vehicle out-of-service (OOS) counts | Daily |
| Crashes | Reportable crash events with fatality, injury, and tow-away details | Monthly |
| Violations | Individual violation records with category, severity weight, and inspection date | Monthly |
| BASIC Scores | SMS safety measure scores — SMS AB (interstate + intrastate hazmat) and SMS C (intrastate non-hazmat) merged by DOT number | Monthly |
| SMS Census | Authority classification fields: for-hire, exempt, private property, government, etc. | Monthly |
| History | Operating authority orders, revocations, and docket status changes | Daily |
Carrier Eligibility
Who gets scored — and who doesn't
Not every FMCSA-registered entity is a for-hire trucking carrier. We apply hard exclusions that remove non-carriers entirely — but beyond those, nearly every carrier with power units and usable exposure is graded. Thin-data carriers are not dropped to N/A; their estimate is shrunk toward their size-band prior and tagged with a confidence tier. This is why the FRED Score covers about 1.1 million carriers, versus roughly 475k under the prior approach that excluded thin carriers outright.
Pre-Scoring Exclusions
Carriers matching any of these criteria are removed before scoring begins:
Confidence Tiers, Not Exclusion
Rather than dropping carriers with thin data, the model shrinks their estimate toward their size-band prior and labels how much real track record backs the score:
Data & Exposure Normalization
Why raw counts are misleading
Public FMCSA records — crashes, inspections, and violations — are aggregated for each carrier over its most recent observation window. But comparing a 3-truck fleet to a 500-truck fleet by raw event counts is inherently unfair — a larger fleet naturally encounters more events. Exposure is how the model accounts for that: crash burden is measured per power-unit-year and each violation signal per inspection, so a carrier is compared on a rate, not a raw count. We use power-unit-years rather than self-reported mileage on purpose — MCS-150 mileage is frequently blank, stale, or implausible, so it is never the exposure base for the grade.
Exposure is the carrier's mileage over the scoring window, in 100k-mile units. It enters the model as log(E) — an offset that puts every metric on a per-100k-mile footing so small and large fleets are predicted on the same scale:
Credibility & Confidence
How much to trust a carrier’s own history
Even on a per-mile basis, a tiny fleet with 1 crash in 50k miles looks far worse than a large fleet with 10 crashes in 5 million miles — even though the small fleet's rate is mostly just noise. One lucky or unlucky year can swing it wildly. Rather than throwing thin carriers away, the model blends each carrier's own signal with its size-band prior using Bühlmann-Straub credibility.
Each carrier gets a credibility weight $Z = n/(n+K_{\text{band}})$, where $n$ is the carrier's accumulated evidence — the count of real safety observations (roadside inspections and observed crashes / out-of-service results, plus a small fleet-size-and-tenure floor). Credibility comes from how much of a carrier we have actually seen, not from the mileage it self-reports: a one-truck fleet that merely claims 100k miles has almost no evidence and is shrunk to its size-band prior, while a fleet with a heavy inspection record (or real adverse events) is governed mostly by its own history. Carriers with lots of evidence ($n \gg K$) ride their own record; thin-record carriers ($n \ll K$) are pulled ("shrunk") toward the band prior — so a strong grade has to be earned with observations, never granted for the mere absence of data. Because adverse events add to the evidence count, a genuinely bad record still has the credibility to grade a carrier down.
That same credibility weight drives each carrier's confidence tier:
Frequency × Severity Model
Predicting next year’s crash burden
The FRED Score targets a carrier's severity-weighted crash burden — not just how many crashes, but how bad. Each crash is weighted by its outcome, so a fatal collision counts far more than a minor tow-away:
Fatality and injury terms count casualties (capped) and stack — a two-fatality crash weighs more than a one-fatality crash. The carrier's own severity-weighted crash burden, and its vehicle out-of-service, severe-violation, hours-of-service / fatigue, unsafe-driving, and speeding rates, are each turned into a cohort-relative Empirical-Bayes relativity (1.0× = typical for its size cohort), shrunk toward the cohort average by how much real exposure backs that signal — crash burden by power-unit-years, violation rates by inspection count. These combine into one transparent, credibility-weighted risk relativity; a carrier's observed crashes set a floor on that relativity, so a real crash record can never be averaged away.
The grade comes from a transparent geometric blend of the per-signal relativities $r_{s}$ (no black box) — each raised to a weight $w_s$ learned out-of-time from how strongly that signal predicts the following year's crash burden. In order of learned weight, the signals are:
Weights are log-relativity exponents learned out-of-time by a non-negative Poisson fit; refit each cycle. Signals that don't predict forward crashes (e.g. generic equipment / maintenance write-ups) are dropped rather than given face-value weight.
Violations as Fitted Predictors
Roadside violations enter as vehicle out-of-service, severe-violation, hours-of-service / fatigue, unsafe-driving, and speeding rates — each weighted by how strongly it predicts the following year's crash burden. The relativities below are illustrative of the ordering the fit recovers, not pre-set multipliers baked into the score:
| Tier | Violation Type | Forward-crash RR |
|---|---|---|
| Critical — Immediate danger behaviors | ||
| Reckless Driving | 1.49 | |
| Dangerous Driving | 1.37 | |
| Jumping OOS / Driving Fatigued | 1.36 | |
| High — Serious behavioral risks | ||
| Speeding (high & excessive) | 1.30 | |
| Drugs / Alcohol | 1.29 | |
| Alcohol Possession | 1.27 | |
| Moderate Speeding | 1.20 | |
| Phone Call / Texting | 1.16–1.18 | |
| Moderate — Concerning behaviors | ||
| False Log | 1.12 | |
| Seat Belt | 1.11 | |
| Equipment — Vehicle condition | ||
| Lighting | 1.17 | |
| Tires | 1.16 | |
| Brakes (all types) | 1.13 | |
Each figure is the empirical relative risk — the ratio of next-year crash burden for carriers with that violation type versus those without. The model learns how much weight to give each behavioral and equipment feature directly from the year-over-year data, so its influence on the score reflects measured forward risk rather than a fixed assumption.
Per-Band Calibration
Unbiased predictions across every fleet size
A model can rank carriers well yet still systematically over- or under-predict for a given fleet size. To prevent that, each fleet-size band is calibrated so its total predicted crashes (and burden) match that band's own recent observed totals — the observed-over-expected ratio lands at ≈ 1.0:
The forward expectation is then surfaced for underwriting as two figures per carrier — expected crashes and expected severity-weighted burden over the next 12 months:
A parallel Poisson model is fit the same way on fatal crashes (crashes involving a fatality) to estimate each carrier’s forward fatal-crash rate, surfaced for underwriting as the probability of at least one fatal crash in the next 12 months. Fatal crashes are rare, so this estimate leans heavily on exposure, fleet-size band, and behavioral / speeding signals, and is calibrated band-by-band like the burden model:
Grades & Risk Relativity
Where a carrier sits among same-size peers
Each carrier's predicted burden is first expressed as a risk relativity — its predicted burden divided by the level that is typical for its size band. 1.00× means typical for its size; below 1 is safer than peers, above 1 is riskier.
Grades are then assigned from this credibility-weighted risk relativity on fixed thresholds calibrated to realized forward crash burden — below 0.52× the carrier's size cohort earns “Excellent,” 0.52–0.60× “Strong,” 0.60–0.90× “Satisfactory,” 0.90–1.20× (at/just above the cohort average) “Marginal,” 1.20–1.60× “Poor,” and above 1.60× “Critical.” The thresholds sit on the actual forward-burden gradient, so each grade realizes materially higher next-year crash burden than the one above it. Because the relativity is measured within the carrier's size cohort — the in-scope book is banded 10–19, 20–49, 50–99, 100–199, 200–499, 500+ power units, with 1, 2–5 and 6–9 held as advisory — a fleet is judged against genuinely comparable peers, never against carriers many times its size. The bands are set finely wherever the population supports a stable per-cohort estimate. A single-power-unit carrier is capped at “Strong” — one power-unit-year of crash exposure is too thin to credibly earn the top grade (see grading owner-operators). A carrier with no inspections and no crashes on record has no observed safety record to rank, so it is marked “Not Rated” rather than assigned a fabricated grade.
The thresholds are absolute — tied to a carrier's measured risk relativity, not a fixed quota — so a size cohort can have many strong carriers or few, exactly as the data warrants, and the grade is stable rather than shifting as peers move. A carrier whose credibility-weighted burden is below half its cohort earns “Excellent” whether it runs 3 trucks or 3,000. The 0–100 FRED Score (100 = safest) expresses the same relativity on a friendlier scale.
Rolling Refit
Out-of-time, refreshed weekly
The model is refit on the most recent complete year→year pair — coefficients are learned from one year of carrier history paired with the crash burden that actually followed. Every carrier is then scored on its latest 12 months of history to predict the forward 12 months. Because the model is never scored on the same data it learned from, the FRED Score is genuinely out-of-time.
The data pipeline refreshes the score weekly, so each carrier's grade reflects their current safety posture. A carrier that improves will see it as older events age out; one that deteriorates feels the impact within months, not years.
The most recent ~45 days are held out for crash-reporting lag — crashes take time to appear in FMCSA's feed, so the very latest weeks aren't yet mature enough to score against. This keeps the forward target honest rather than artificially low for recent activity.
Automatic Rules & Flags
Hard overrides, eligibility gates, and informational flags
Beyond the statistical model, a set of deterministic rules handle edge cases where the math alone isn't sufficient. These fall into three categories: score overrides, data-quality adjustments, and informational flags.
Score Overrides
These rules supersede the calculated score entirely:
An FMCSA Unsatisfactory (U) or Conditional (C) rating only counts against a carrier when it is current. A rating is treated as historical — shown for context but not penalized — when it is stale or superseded. The key cross-check: a for-hire carrier cannot legally operate under a current Unsatisfactory rating (FMCSA revokes its operating authority), so an Unsatisfactory sitting next to an active operating-authority docket is proof the rating was superseded, regardless of its age. A rating older than the recency window is likewise historical. We derive this from bulk FMCSA fields (SAFETY_RATING, SAFETY_RATING_DATE, REVIEW_DATE, docket status, and revocation history) and, on the carrier page, confirm it against a real-time FMCSA lookup. Only a genuinely current adverse rating drives the grade down; a historical one does not. When an Unsatisfactory or Conditional rating is genuinely current, the carrier is flagged for referral and its displayed grade is capped at “Marginal” for binding — even if its credibility-weighted burden would place it higher — while a worse observed record still shows through.
If a carrier has interstate-only scope, no active operating authority, and is not exempt — the FRED Score is set to N/A (no score produced). N/A is reserved for defunct or no-authority carriers, never for being small.
Eligibility Gates
A thin carrier is shrunk toward its size-band prior, not blocked. These gates only apply when there is no usable exposure at all or the reported data is implausible:
No reported mileage and zero inspections — nothing to anchor even a prior-based estimate. A carrier with any usable exposure is still scored as Provisional on its size-band prior.
Reported mileage is implausible: >300k miles per truck, <1k miles per truck for fleets ≥10, or >500M total miles. Scoring is blocked to prevent extreme rates.
Reports >300k miles per truck and has fewer than 2 inspections — high mileage can't be corroborated by inspection activity.
Reported power units far exceed any real-world carrier (the census occasionally carries values in the hundreds of thousands to millions), a very large fleet is claimed with no usable mileage to corroborate it, or the fleet is internally inconsistent — far more power units than drivers (e.g. 4,150 trucks against 4 drivers), which can’t be operated. Such records have no trustworthy exposure — they are left ungraded rather than handed a fabricated baseline, and excluded from the band calibration.
Data Quality Adjustments
Modifications to component scores when data quality is degraded:
When mileage is missing or implausible but the carrier has observed activity (inspections, crashes, or violations), the model falls back to inspection-based exposure rather than trusting the bad odometer figure.
Calculated exposure is floored at 50k miles (0.5 units) to prevent extremely volatile rates from tiny denominators.
Informational Flags
These flags are attached to carrier records for context but do not directly alter the score:
Chameleon detection cross-references every revoked DOT (any historical REVOCATION order) against active carriers using normalized address, officer name, phone, and DUNS. Matches are tiered by signal strength. The flag is informational and does not affect the FRED score — the underwriter decides what to do with the linkage.
Validation Standards
How we know it works
Every refit is validated out-of-time: the model is fit on one year, then judged on whether it ranks and calibrates the following year's crash burden on a carrier-disjoint holdout (every fifth DOT number is held out of fitting). One hard gate blocks publication — grade monotonicity (below) — and the refit aborts before any score is written if it fails. Discrimination and calibration metrics are recorded for every refit and reviewed; the figures below are from the latest validated run (outcome year ending 2026‑03‑08).
In every fleet-size band, realized forward-year burden must rise across all six grades (Excellent → Critical). A statistically significant inversion — a worse grade genuinely safer than a better one, beyond the sampling error of the holdout buckets — rejects the refit and leaves the live scores untouched. A near-tie within sampling noise (adjacent grades in the smallest holdout cells, where a few dozen carriers can't resolve a fraction-of-a-percent burden gap) is not counted as a failure.
Normalized Gini of ≈ 0.20 overall on the forward year among rated carriers, rising with exposure from ≈ 0.11 on single-power-unit carriers to ≈ 0.64 on 100+ unit fleets — the model discriminates most where the data is richest. (Single-unit ranking is intentionally weak; see grading owner-operators.)
Observed-over-Expected burden lands within 0.97 – 1.06 of 1.0 in every fleet-size band on the holdout — predictions stay calibrated across sizes, not biased by fleet size.
Top-decile lift, per-cohort Gini, and a grade-monotonicity gate — realized forward burden must rise from Excellent to Critical, or the run aborts rather than ship a contradictory score — are computed and logged to a per-run metrics artifact for review. All scoring, calibration, and out-of-time validation is performed by the reproducible pipeline
fred_score_v6.py.