General
Real-Time Visibility and Carrier Scorecards: Why Your Ranking Measures Reporting, Not Performance in 2026
Sep 18, 2026
15 mins read

A carrier scorecard built on exception counts from real-time visibility feeds ranks carriers partly by how often they report, because an exception that is never observed is never counted. Carriers reporting at four-hour or eight-hour milestones surface a fraction of the delays that a carrier on continuous telematics surfaces, so on identical underlying performance the sparser reporter scores better. Locus, the world’s first Decision-Intelligent, Agentic TMS, normalizes owned-fleet telemetry and carrier events into one model with observation frequency carried as data, which is what makes a cross-carrier comparison defensible.
Key Takeaways
- Exception-based carrier scorecards measure observed exceptions, not actual ones. Observation depends on reporting cadence, which differs by an order of magnitude across a typical carrier base.
- In our illustrative model, three carriers with an identical 8% true late rate report 7.00%, 4.00% and 0.00% purely because they update hourly, four-hourly and eight-hourly.
- Ranked on reported exceptions, the order is exactly the order of least reporting to most. The best-scoring carrier is the one you can see least.
- The correction is to normalize by observation density before comparing, and to report reporting cadence alongside every performance figure on the scorecard.
- Terminal outcomes such as on-time delivery and failed first attempts are reported by every carrier regardless of cadence, which makes them the only directly comparable cross-carrier measures.
- Locus maps carrier events and owned-fleet telemetry into one event model with freshness carried as data, so performance and visibility are separable.
Why Carrier Scorecards Inherit a Measurement Problem
Carrier performance management has become a data exercise. Shippers pull exception and milestone data from visibility platforms, compute on-time and exception rates per carrier, and use the result to award volume, negotiate rates and trigger contractual remedies. The logic is sound and the inputs are not comparable.
The scale of the problem tracks the fragmentation of the carrier base. AlixPartners’ 2026 Home Delivery Survey found more than 90% of executives run a mix of last-mile carriers and 32% use four or more. A scorecard spanning four or more carriers is almost certainly spanning four or more reporting standards, and the differences between them are larger than the performance differences the scorecard is trying to detect.
Conditions are also making the underlying performance harder to read. INRIX’s 2025 Global Traffic Scorecard found congestion increased in 254 of the 290 US cities it analyzed, which raises the variance a carrier is operating against. More variance means more exceptions available to be observed, so the gap between a well-reported carrier and a sparsely reported one widens with congestion even when neither carrier’s operation has changed.
What hangs on the ranking is real money. McKinsey’s out-of-home delivery work puts the last mile at 60% to 70% of total parcel delivery cost, so volume allocation across carriers is one of the larger cost levers a shipper controls. Getting it wrong is not a reporting inconvenience, it is a procurement decision made on a distorted input.
The Distortion, Quantified
A delay does not become an exception when it happens. It becomes an exception when it is observed, and it is observed at the next status update. That single fact drives the whole distortion.
We modeled it. The inputs are illustrative rather than measured: a disruption arising at a random point in transit with roughly four hours of transit remaining afterwards, becoming visible at the next status update, and three carriers with an identical true late rate of 8%.
| Carrier | Status update interval | Share of lateness visible before delivery | Late rate the scorecard reports |
|---|---|---|---|
| Carrier A | 1 hour | 87.5% | 7.00% |
| Carrier B | 4 hours | 50.0% | 4.00% |
| Carrier C | 8 hours | 0.0% | 0.00% |
All three carriers perform identically. Carrier C reports a perfect record, and the reason is that a delay arising mid-leg on an eight-hour reporting cycle is typically never visible before the delivery either succeeds or fails. It does not arrive late in the data. It arrives resolved.
Ranked best to worst on reported exceptions, the order is C, B, A. That is exactly the order of least reporting to most. The scorecard has produced a perfect inverse ranking of data quality and presented it as performance.
This is the specific failure mode of visibility programs that succeed technically. Every carrier is integrated, every shipment appears, and the resulting league table rewards the carriers that tell you least.
Why the Best-Scoring Carrier Is Often the Least Visible
The inversion has a second-order effect worth naming, because it changes incentives rather than only measurements.
A carrier that invests in telematics and frequent status reporting makes its own problems visible. A carrier that reports at wide intervals does not. If volume is awarded on reported exception rates, the shipper is paying the second carrier for opacity and penalizing the first for transparency. Over a few contract cycles that is a real signal to the carrier base about what the shipper actually rewards.
It also means the carriers you can manage are the ones that look worst. The whole operational value of frequent reporting is that it creates a window in which recovery is possible, so the carrier producing the most actionable exceptions is the one giving you the most opportunity to protect the customer. On an exception-count scorecard that shows up as a liability.
Finally, the distortion is largest exactly where it matters most. The gap between observed and actual exceptions widens as true performance worsens, because there is more to miss. Two carriers at 3% true lateness look more alike than two at 12%, so the ranking is least reliable in the tail you built the scorecard to find.
How to Normalize a Carrier Scorecard
1. Record reporting cadence per carrier as a scorecard field
Compute the median and 90th percentile interval between status updates, per carrier and per lane. This is the denominator every performance figure on the scorecard depends on, and most scorecards do not carry it at all.
2. Estimate observable share before comparing rates
For each carrier, estimate what proportion of a mid-transit disruption would become visible before delivery given its cadence and typical remaining transit. That proportion is the factor the reported exception rate has been multiplied by.
3. Compare within cadence bands, not across them
Carriers on similar reporting frequencies can be compared directly. Carriers on very different frequencies cannot be, and grouping the scorecard into bands is more defensible than applying a correction factor nobody can audit.
4. Use terminal outcomes as the cross-carrier measure
Delivery outcome at the end of the journey is reported by every carrier regardless of cadence. On-time delivery against the commitment, failed first attempts and claims are comparable in a way that mid-transit exception counts are not.
5. Report exception counts as a visibility metric, not a performance one
Exceptions surfaced per hundred shipments is a useful measure of how manageable a carrier is. Keep it on the scorecard, label it as visibility, and stop using it to rank service quality.
6. Negotiate cadence as a contract term
Reporting frequency is the variable that makes everything else measurable, and it is negotiable at renewal in a way that a carrier’s operating performance is not. A carrier that cannot improve cadence can be moved to lanes short enough that sparse reporting still leaves room to act.
| Also Read: Real-Time Tracking and Visibility for 3PLs |
|---|
A Worked Correction on a Four-Carrier Base
The correction is easier to see applied. Take a shipper running four carriers, pulling exception data from one visibility platform, and awarding volume on the resulting exception rate.
| Carrier | Median update interval | Reported exception rate | On-time against commitment | Raw rank | Rank on outcome |
|---|---|---|---|---|---|
| Regional A | 45 minutes | 9.1% | 94.2% | 4th | 2nd |
| National B | 3 hours | 5.4% | 91.8% | 3rd | 3rd |
| Regional C | 6 hours | 2.8% | 89.1% | 2nd | 4th |
| National D | 90 minutes | 1.9% | 96.4% | 1st | 1st |
The figures are illustrative, but the shape is the point. Raw exception rates rank Regional A last and Regional C second. Terminal outcomes reverse them: Regional A is the second-best carrier on the measure the customer experiences, and Regional C is the worst. The raw ranking placed the two exactly backwards.
National D ranks first on both, which is the case that makes the distortion hard to spot. When a genuinely strong carrier also reports frequently, the raw scorecard produces a defensible top of the table and an inverted middle, so the ranking looks broadly sensible and only misleads on the decisions between adjacent carriers, which is where allocation is actually decided.
The correction takes an afternoon: pull median update interval per carrier, pull on-time against commitment, and re-rank. Where the two rankings disagree, the raw one is the one to distrust.
Raw and Normalized Scorecards Compared
| Scorecard element | Raw exception-based scorecard | Normalized scorecard |
|---|---|---|
| Primary performance measure | Exceptions observed per hundred shipments | Terminal outcome against commitment |
| Treatment of reporting cadence | Not recorded | Recorded per carrier and lane, shown alongside every figure |
| Cross-carrier comparison | Direct, across all carriers | Within cadence bands, or on terminal outcomes only |
| What exception counts are used for | Ranking service quality | Measuring manageability and recovery opportunity |
| Incentive created for carriers | Report less | Report more, compete on outcomes |
| Reliability in the worst-performing tail | Lowest, because more goes unobserved | Preserved, because outcomes are reported by everyone |
The incentive row is the one to take to a procurement conversation. A scorecard that rewards opacity will get opacity, and the effect is slow enough that nobody attributes it to the measurement design.
What to Check on Your Own Scorecard
Does it record how often each carrier reports? If cadence is not a field, the scorecard cannot distinguish a quiet carrier from a good one, and neither can anyone reading it.
Are mid-transit exceptions and terminal outcomes reported separately? Blending them produces a single number whose comparability depends on the mix, which varies by carrier.
Does the scorecard reward improvement you can actually see? A carrier that invests in telematics will show more exceptions in the first quarter after doing so. If the scorecard reads that as deterioration, the program is penalizing exactly the change it should reward, and the carrier will notice before you do.
Do your best-scoring carriers also have your sparsest feeds? Plot reported exception rate against median update interval across your carrier base. A strong relationship is the distortion, visible directly in your own data.
Does the scorecard drive volume allocation? If it does, and cadence is unrecorded, then some share of your allocation is being decided by reporting frequency. That is worth quantifying before the next award cycle.
Does the raw ranking agree with the terminal-outcome ranking? Re-rank your carriers on on-time against commitment and compare the two orders. Every disagreement is a carrier you are currently mis-rating, and the direction of the error is predictable: sparse reporters sit too high.
Is exception latency measured separately from exception frequency? A carrier reporting hourly but delivering events six hours after they occur is as blind as one reporting six-hourly. Frequency and latency are different fields and most feeds carry only one of them.
Putting Cadence Into the Contract
Measurement design eventually has to reach the contract, because a scorecard that does not change what carriers are paid for will not change what carriers do.
Three terms are worth negotiating and are rarely present. The first is a stated minimum update frequency, expressed as a maximum interval between status events rather than as a vague commitment to provide tracking. The second is event coverage, meaning which event types must be reported at all: many carrier feeds report arrival and delivery reliably and report exceptions inconsistently, and an exception feed that omits the exceptions is the specific failure this article describes. The third is latency, the gap between an event occurring and the event reaching you, which is distinct from frequency and often worse.
Where a carrier genuinely cannot meet a cadence, the useful response is allocation rather than penalty. Short lanes and wide delivery windows tolerate sparse reporting because there is less transit in which an unobserved problem can develop. Tight windows and long lanes do not. Matching carriers to lanes by reporting capability turns a data limitation into a routing decision, which is cheaper than either penalizing a carrier for infrastructure it does not have or accepting blind spots on the work that can least afford them.
It is also worth stating the scorecard methodology to carriers explicitly. A carrier that understands it is ranked on terminal outcomes and measured separately on visibility has no incentive to suppress reporting, and several will improve cadence without a contractual requirement once they know it is not being scored against them.
Common Mistakes in Carrier Performance Measurement
Comparing exception rates across carriers with different reporting frequencies. This is the core error, and it systematically favors the carriers you can see least.
Treating a quiet feed as a clean operation. On a sparse feed, quiet means nothing has been reported recently, which is also what a serious problem looks like until the next scan.
Rewarding low exception counts contractually. It creates a direct incentive against the reporting investment that makes recovery possible.
Assuming integration completeness fixes comparability. Connecting every carrier standardizes the pipe, not the cadence. Two fully integrated carriers can still be observed at wildly different rates.
Reading a year-on-year improvement in exception rates as better performance. If carrier mix shifted toward sparser reporters, or a carrier reduced its reporting, the exception rate falls without anything improving. Check cadence before celebrating the trend.
How Locus Makes Carrier Performance Comparable
Locus, the world’s first Decision-Intelligent, Agentic TMS, maps carrier events and owned-fleet telemetry into one event model rather than presenting each feed in its own vocabulary. ShipFlex provides pre-integrated carrier access within an ecosystem of more than 1,000 carriers, so events arrive already normalized to a common state model, and the Control Tower carries the observation time and update interval on each state. That is what allows a performance figure and a visibility figure to be separated rather than blended into one misleading number.
The Carrier agent holds allocation decisions against that normalized view, which means carrier selection is computed from comparable inputs rather than from a ranking distorted by reporting frequency. The DiSCO governance mechanisms, particularly Explainability and Traceability, record which observation drove each exception, so a performance figure presented to a carrier can be evidenced with the data that produced it rather than asserted.
A Fortune 50 enterprise running a driver pool of more than 4,500 across captive and third-party capacity raised weekly plan execution from 75% to 92% and surfaced more than $14M in capacity it already owned through centralized dispatch, on a network where mixed captive and contracted capacity is the permanent condition. A Canadian grocery brand delivering through contracted third-party fleets in more than 30 cities cut customer support resolution 10 to 20 times faster through carrier orchestration, alongside 33% faster deliveries and 15% lower fulfillment cost.
Locus has been recognized by Gartner for seven consecutive years across multiple research categories, including the 2026 Gartner Market Guide for Multicarrier Parcel Management Solutions, where ShipFlex is featured as a Representative Vendor, and Representative Vendor status in the 2026 Gartner Hype Cycle for Supply Chain Execution and Logistics Technologies. QKS Group positions Locus as the Leader in its SPARK Matrix for Transportation Management Systems 2025, and G2 ranked Locus number one in Route Planning in its 2026 Best Software Awards. The platform has run more than 1.5 billion deliveries for 360+ enterprise customers across 30+ countries at 99.99% uptime.
In October 2025, Ingka Investments, the investment arm of Ingka Group, the world’s largest IKEA retailer, acquired Locus. Locus continues to operate independently.
The change this argues for costs nothing and can be made before the next review cycle. Put reporting cadence on the scorecard as a field, rank carriers on terminal outcomes rather than on observed exceptions, and keep exception counts as a measure of how manageable each carrier is. The carriers that currently look worst are frequently the ones giving you the most chance to protect the customer. Locus normalizes carrier events and owned-fleet telemetry into one model with freshness carried as data, across 30+ countries and more than 1,000 carriers. Schedule a demo to see it against a live carrier base.
FAQs
Why do carrier scorecards favor carriers that report less? Because exception-based scorecards count observed exceptions, and an exception is only observed at the next status update. A carrier reporting every eight hours surfaces almost no mid-transit delays before delivery, so it reports a near-perfect record on performance identical to a carrier reporting hourly.
How do you compare carriers with different tracking frequencies? Compare terminal outcomes such as on-time delivery against commitment, failed first attempts and claims, which every carrier reports regardless of cadence. Keep mid-transit exception counts grouped within similar cadence bands, and record each carrier’s median update interval on the scorecard.
What is the right way to use exception counts in carrier management? As a visibility measure rather than a performance one. Exceptions surfaced per hundred shipments tells you how manageable a carrier is and how much recovery opportunity it gives you, which is useful information provided it is not used to rank service quality.
Does integrating every carrier fix the comparability problem? No. Integration standardizes how data arrives, not how often. Two fully integrated carriers can report at intervals differing by an order of magnitude, and the scorecard distortion depends on interval rather than on integration.
Where is a carrier scorecard least reliable? In the worst-performing tail. The gap between observed and actual exceptions widens as true performance worsens, because there is more to miss, so the ranking is least trustworthy precisely among the carriers it was built to identify.
Should reporting cadence be a contractual term? It is worth negotiating, because it is the variable that makes every other performance measure comparable and it is more tractable at renewal than a carrier’s operating performance. Where cadence cannot improve, allocating that carrier to shorter lanes preserves the ability to act.
Aseem, leads Marketing at Locus. He has more than two decades of experience in executing global brand, product, and growth marketing strategies across the US, Europe, SEA, MEA, and India.
Related Tags:
General
Real-Time Visibility and ETA Confidence: Why a Point ETA Hides What You Need in 2026
Every visibility platform shows one ETA. Two shipments with the same ETA can have very different chances of missing, and the point estimate cannot tell you which.
Read more
General
Real-Time Visibility Across European Borders: Why Your ETA Error is Zero or Two Hours in 2026
On ferry and shuttle legs an ETA is either right or wrong by a whole departure interval. Average error describes almost no shipment, which is why the metric misleads.
Read moreInsights Worth Your Time
Real-Time Visibility and Carrier Scorecards: Why Your Ranking Measures Reporting, Not Performance in 2026