General
Last-Mile Delivery Efficiency Benchmarks: How to Compare Couriers Beyond Transit Time
Aug 21, 2026
13 mins read

Key Takeaways
- Transit time, rate, and coverage are procurement metrics. They describe what you bought, not how the operation performed.
- Five measures compare couriers on efficiency: first-attempt rate, re-delivery burden, WISMO contact rate, on-time-in-full by promise window, and exception lead time.
- Published benchmark ranges for these metrics do not exist at research grade. Every circulating version traces to software vendors rather than to research firms or government sources.
- That absence matters less than it appears, because an industry average would be the wrong comparator anyway. Courier efficiency is lane-specific, so a carrier at 88% in dense urban may be outperforming one at 93% in suburban.
- The benchmark that works is your own network: carriers compared against each other on matched lanes, with zone type held constant.
The transit time trap
Procurement compares couriers on transit time, rate card, and coverage. Those are the right metrics for a purchasing decision and the wrong ones for an efficiency review, because all three describe the contract rather than the outcome.
Transit time in particular is a service commitment. It tells you what the carrier undertakes to do under normal conditions. It says nothing about what share of your parcels complete on the first attempt, how often your support team fields a status enquiry, or whether the carrier tells you about a problem before your customer does. Two couriers with identical published transit times can produce materially different operational costs on the same lane.
This matters more than it used to, because customer priorities moved. McKinsey found that speed fell from consumers’ number one delivery priority in 2022 to fifth by 2024, displaced by reliability and predictability, with around 90% of consumers willing to wait two to three days when delivery is free and arrives within the stated window. If reliability outranks speed for the customer, a carrier comparison weighted on speed is optimising for the wrong thing on both sides.
The stakes justify getting the comparison right. Capgemini Research Institute puts last-mile delivery at 41 to 53% of total logistics and shipping cost, which makes carrier performance one of the larger controllable variables in the network.
Why published benchmarks will not help you
You will find benchmark tables for every metric below. Strong urban first-attempt rates, acceptable WISMO contact rates, target on-time percentages by sector.
None of them trace to a research firm, a government source, or peer-reviewed work. First-attempt delivery failure rates, on-time benchmarks by sector, WISMO cost and volume figures, and cost per failed delivery all originate with software vendors and aggregator pages, generally without a stated methodology. Planning against a number whose derivation you cannot inspect is not measurement.
The more useful point is that an industry average would be a poor comparator even if a credible one existed. Courier efficiency is a function of the work, not just the carrier. Density, building type, recipient availability, access difficulty, and parking all vary enormously by zone, and they move these metrics by more than the difference between two competent carriers. A carrier running 88% first-attempt in dense urban may be outperforming a carrier running 93% in low-density suburban, and an absolute benchmark would tell you the opposite.
The comparator that works is your own network. Compare carriers against each other on matched lanes, holding zone type constant, and use your best performer per zone as the internal benchmark. That number is defensible, specific to your mix, and impossible for a competitor to replicate.
Five measures that actually compare couriers
1. First-attempt delivery rate, segmented by zone type
The share of deliveries completed without a second attempt. Segment by zone before comparing anything, because a network-wide figure is a weighted average of your zone mix rather than a measure of carrier quality.
Pair it with structured failure reasons. Recipient unavailable, access refused, address problem, and damaged require different responses and point to different owners. Without reason codes the metric identifies that there is a problem and nothing about whose it is.
2. Re-delivery burden
Total attempts per completed delivery, rather than a redelivery count. Attempts are the unit of cost, because each one consumes approach, access, and driver time regardless of outcome.
Most of that cost sits outside the vehicle. Urban Freight Lab research at the University of Washington, based on more than 1,800 real deliveries, found urban commercial vehicles spend around 80 percent of daily operating time parked, with most of the driver’s time spent outside the vehicle. A failed attempt consumes that whole cost and produces nothing, which is why attempts per completion is a better efficiency measure than a failure percentage.
3. WISMO contact rate per thousand deliveries
Support contacts asking about order status, normalised per thousand deliveries and attributed to the carrier handling them.
This measures the carrier’s communication quality as your customers experience it. A carrier with adequate delivery performance and poor status reporting generates support cost that never appears in its rate card, and it is frequently the largest hidden cost difference between two otherwise comparable providers.
4. On-time-in-full by promise window
Not on-time against the carrier’s service standard, which is what carrier scorecards report, but against the window you committed to your customer.
Segment by window width. A carrier that reliably hits a four-hour window and misses a one-hour window is telling you something specific about where to use them, and a blended figure hides it.
5. Exception lead time
How long before the customer would have noticed does the carrier surface a problem. This is the measure that separates carriers who report from carriers who warn, and it determines whether an exception is recoverable.
It is also where most stacks stop. Gartner found that while 95% of supply chains must react quickly to change, only 7% can execute decisions in real time. A carrier that surfaces an exception forty minutes before the promised window closes has given you a decision. One that reports it afterwards has given you a record.
| Also Read: The First-Attempt Delivery Rate: A Key Metric That Decides Last-Mile Profitability in 2026 |
|---|
Zone type is the control variable
Before any carrier comparison means anything, classify your delivery geography. Four categories are usually sufficient.
Dense urban core. High stop density, low drive distance, high access friction. Parking scarcity and building navigation dominate the stop. First-attempt rates run lower here for structural reasons and re-delivery is cheaper because the return trip is short.
Suburban residential. Moderate density, moderate drive distance, low access friction, high variance in recipient availability. First-attempt rate is driven mainly by whether anyone is home.
Commercial and business districts. Receiving windows, docks, and mailrooms. Performance depends on window adherence rather than on access, and a missed window can mean a full day lost.
Rural and low density. Long drive distance, few stops, and expensive failures. The US Postal Regulatory Commission has found average cost per delivery in rural areas runs approximately twice that of urban areas, which makes a failed attempt here materially more costly than the same failure downtown.
The practical consequence: build a carrier scorecard per zone type, not per carrier. The same provider will rank differently in each, and the output you want is not a ranking but an allocation rule.
Running a courier efficiency audit
Six steps, all achievable with data you already hold.
- Classify every delivery postcode into one of the four zone types. This is a one-time exercise and everything downstream depends on it.
- Pull the five metrics per carrier per zone for a complete quarter. A quarter smooths seasonal variance without going stale.
- Identify matched lanes, meaning zone and volume combinations where two or more carriers both operate. These are the only genuinely comparable data points, and they are usually a smaller share of the network than expected.
- Set your internal benchmark as the best performing carrier per metric per zone. That is your realistic target, since it has been achieved on your work by a carrier you already use.
- Quantify the gap for each underperforming carrier per zone, converting attempts, contacts, and window misses into hours and cost rather than percentages. Percentages do not survive a budget conversation.
- Convert the finding into an allocation rule, not a procurement conversation. The output of the audit should change which carrier gets which order in which zone, and only then feed the contract discussion.
Step three is where most audits stall, because carriers are frequently allocated by geography in a way that leaves almost no overlapping lanes. If that describes your network, deliberately overlapping two carriers on a subset of lanes for a quarter is worth the small cost, because it is the only way to generate comparable data.
Why the right courier varies by zone
The output of the audit is rarely a best carrier. It is a map of which carrier is strongest where, and the pattern is usually consistent with the structural characteristics of each provider type.
National networks bring universal coverage and standardised process, which travels well into rural and mixed geographies where density cannot support a dedicated operation. Regional carriers frequently outperform inside their footprint, because density and local knowledge compound. Owned or contracted urban fleets can hold tight windows and resolve exceptions directly, at a cost that only makes sense where volume density supports it.
None of that is a ranking. It is a reason to stop looking for one, and to build an allocation rule instead. The mistake worth avoiding is treating the audit output as a procurement decision when it is an operational one: the same carrier mix, allocated differently, frequently produces more improvement than a renegotiated rate card.
Where the orchestration layer fits
Two things have to be true for this framework to change anything. The metrics have to be measurable across carriers in one place, and the allocation rule has to be executable order by order rather than as a quarterly policy.
Both are orchestration-layer functions. Carrier statuses have to be normalised into one taxonomy before any cross-carrier comparison is valid, since each provider reports different events at different granularity. And once the zone-level pattern is known, allocation has to apply it at dispatch time, including reallocating when a carrier is underperforming on a lane this week rather than last quarter.
Locus, the world’s first Decision-Intelligent, Agentic TMS, operates at that layer. Within DiSCO, the Carrier agent scores providers on cost and service and allocates each order to the best fit, the Dispatch agent re-sequences on live events, and the Customer agent manages the promise when a plan changes. Carrier statuses are harmonised at ingestion, which is the prerequisite for the comparison above rather than a reporting convenience. Six governance mechanisms bound autonomous action, including explainability and traceability, so an allocation decision can be explained when a carrier disputes a scorecard.
Locus has been recognized by Gartner for seven consecutive years, featured in the 2026 Hype Cycle for Supply Chain Execution and Logistics Technologies, named a Leader in TMS by QKS Group (SPARK Matrix), and ranked #1 in Route Planning on G2’s 2026 Best Software Awards. In October 2025, Ingka Investments, the investment arm of Ingka Group, the world’s largest IKEA retailer, acquired Locus. Locus continues to operate independently.
Two deployments show the same carrier network producing different results under different allocation. A leading Canadian grocery brand delivering perishable food across more than 30 cities through contracted 3PL carriers had been selecting carriers by manual judgement against serviceability sheets, comparing rates and ETAs order by order, with the logic living in planners’ heads. Moving selection into autonomous allocation on the same carrier network produced 33% faster deliveries and 15% lower fulfilment costs. The carriers did not change; the allocation did.
A leading ASEAN apparel retailer shows the measurement prerequisite. Every carrier had been reporting delivery events in its own status codes, so operations tracked shipments carrier by carrier and internal systems never saw a common status, which makes cross-carrier comparison impossible. With statuses harmonised into one standard set and allocation running on live serviceability, cost, and performance rules, delivery SLA held above 99% while WISMO and returns queries fell more than 40%.
The comparison worth running first
Take one zone type where two carriers both operate, pull first-attempt rate and attempts per completion for both over a quarter, and convert the difference into driver hours.
If the gap is material, you have found a reallocation worth making before any contract conversation, using carriers you already have under contract. That is the fastest available improvement in most multi-carrier networks, and it does not require anyone to sign anything.
Learn more, visit locus.sh
Frequently Asked Questions (FAQs)
How do you compare courier efficiency beyond transit time?
On five measures rather than on service commitments: first-attempt delivery rate segmented by zone type, attempts per completed delivery, WISMO contact rate per thousand deliveries, on-time-in-full against your customer promise window rather than the carrier’s service standard, and exception lead time measuring whether the carrier warns you before the customer notices. Transit time, rate, and coverage describe what you bought; these describe what happened.
What is a good first-attempt delivery rate?
There is no credible published benchmark. Circulating figures for first-attempt rates trace to software vendors and aggregators rather than to research firms or government sources. More usefully, an industry average would be the wrong comparator anyway, because the metric is dominated by zone characteristics. Set your benchmark as the best performing carrier per zone type in your own network, since that target has been achieved on your actual work.
Why does delivery zone type matter when comparing carriers?
Because zone characteristics move these metrics more than carrier quality does. Dense urban stops carry access friction and parking scarcity, suburban performance is driven by recipient availability, commercial deliveries turn on receiving windows, and rural failures are expensive, with the Postal Regulatory Commission finding rural cost per delivery runs approximately twice urban. Comparing carriers without holding zone constant measures your zone mix rather than their performance.
How do you run a courier efficiency audit?
Classify delivery postcodes into zone types, pull the five metrics per carrier per zone for a full quarter, identify matched lanes where two or more carriers both operate, set the best performer per metric per zone as your internal benchmark, quantify the gap in hours and cost rather than percentages, and convert the finding into an allocation rule before it becomes a procurement conversation.
What if our carriers do not overlap on any lanes?
That is common, since carriers are frequently allocated by geography in ways that leave no comparable data. Deliberately overlapping two carriers on a subset of lanes for a quarter generates the comparison, and the cost of doing so is usually small against the value of knowing which provider performs better on your work rather than in general.
Should the audit result change our carrier contracts?
Change allocation first. The same carrier mix allocated by zone-level performance frequently produces more improvement than a renegotiated rate card, and it can be implemented immediately without a contract cycle. The audit output then strengthens the contract conversation, because you are negotiating with lane-level evidence rather than with a general impression of service quality.
Ishan, a knowledge navigator at heart, has more than a decade crafting content strategies for B2B tech, with a strong focus on logistics SaaS. He blends AI with human creativity to turn complex ideas into compelling narratives.
Related Tags:
General
Your Legacy TMS Renewal is Not a Selection Decision: A 2026 Framework for North American Shippers
A TMS renewal is a different decision from a TMS selection, and the contract date sets the timeline. Three options, the five factors that decide, and how to plan backwards from your notice deadline.
Read more
General
Real-Time Visibility or Just Another Dashboard? How to Tell the Difference
Logistics blind spots live in the handoffs between legs, not inside them. Why carrier feed aggregation cannot close them, the four handoff types that matter, and three questions to ask any vendor.
Read moreInsights Worth Your Time
Last-Mile Delivery Efficiency Benchmarks: How to Compare Couriers Beyond Transit Time