Logistics Automation & Orchestration
Why Automating 70% of Exceptions Only Removes 32% of the Work
Sep 11, 2026
15 mins read

Logistics automation removes human effort from a workflow by handling cases without intervention, and logistics orchestration coordinates those handled cases across systems so one decision accounts for all of them. Both deliver, and the business cases built on them are usually wrong in the same direction, because they forecast effort from case volume. Those are not the same quantity. Automation takes the tractable cases first, by construction, since tractability is what makes a case automatable at all. What remains is the residual that resisted, and it costs more per unit than the average case did before anything was automated. The arithmetic is unforgiving. On a realistic exception mix, automating 70% of cases removes only 32% of the effort, and the mean handling time of what is left rises 2.27 times. Locus, the world’s first Decision-Intelligent, Agentic TMS, runs more than 12 million automated decisions a day across dispatch, routing and exception management, and the operations that get the most from that volume are the ones that planned for the residual rather than assuming it away.
Key Takeaways
- Effort and case volume are different quantities. Automating 70% of exceptions removes 32% of the work, not 70%.
- The residual is the hardest residual by construction, since tractability is what made the other cases automatable.
- Mean handling time rises as automation rises: 2.27x at 70% automation and 6.06x at 95% on the same mix.
- A plan budgeting headcount from case volume budgets 4.13 FTE where 9.38 are needed, a 2.27x shortfall.
- The gap runs 33 to 40 percentage points across every exception mix tested, so no realistic distribution makes the linear assumption safe.
- Automation removes the apprenticeship. The simple cases were the training volume for the people who must handle the hard ones.
- Locus automates routine allocation while Autonomy Levels and Human Review decide which decisions reach a person, so the residual is a designed set rather than a leftover.
Why the residual gets harder as automation gets better
This failure mode is documented and it is older than logistics software. In 1983 Lisanne Bainbridge published Ironies of Automation in Automatica, one of the most cited papers in human factors, and its argument is precisely this: automating most of a system leaves the operator responsible for the tasks that could not be automated, which are the harder ones, while removing the routine practice that built the skill to handle them. Her conclusion was that operators need more training after automation, not less, which is the opposite of what most automation business cases assume about the people left behind.
The companion finding comes from Parasuraman and Riley’s classification of automation use, misuse, disuse and abuse, where misuse is over-reliance producing failures of monitoring. Read together, the two describe the position an automated logistics operation puts its remaining staff in: fewer cases, each harder, with less practice and more monitoring.
Work it through on a concrete mix. An operation handles 1,000 exceptions a day. Seventy percent are simple and take three minutes, 25% are moderate at ten minutes, and 5% are complex at forty. That is 6,600 minutes a day, 110 hours, or 13.75 full-time equivalents at an eight-hour day, with a mean handling time of 6.6 minutes.
| Automation stage | Automation rate | Cases remaining | Effort remaining | Mean handling time |
|---|---|---|---|---|
| Nothing automated | 0% | 1,000 (100%) | 13.75 FTE (100%) | 6.6 min |
| Simple cases automated | 70% | 300 (30%) | 9.38 FTE (68%) | 15.0 min, 2.27x |
| Simple and moderate automated | 95% | 50 (5%) | 4.17 FTE (30%) | 40.0 min, 6.06x |
The second row is the one that breaks business cases. Case volume fell 70% and effort fell 32%. Both numbers describe the same successful project, and which one you put in the plan determines whether the plan survives contact with the operation.
The obvious question is whether this depends on the particular mix chosen, and it does not. Running the same calculation across four distributions shows the gap persists everywhere.
| Exception mix | Cases removed | Effort removed | Handling time |
|---|---|---|---|
| Even spread, 50/35/15 | 50% | 14% | 1.73x |
| Moderately skewed, 60/30/10 | 60% | 20% | 1.99x |
| Typical, 70/25/5 | 70% | 32% | 2.27x |
| Heavily skewed, 85/12/3 | 85% | 52% | 3.23x |
The gap between cases removed and effort removed runs between 33 and 40 percentage points on every distribution tested, so no realistic mix makes the linear assumption safe. It is worth noting the effect is not monotonic in skew: an operation whose workload looks highly automatable, at 85% simple cases, still gives back only 52% of its effort, while an evenly spread operation gives back 14%. Whichever end you sit at, forecasting from case volume overstates the saving.
The consequence for headcount is direct. A forecast that scales staffing with case volume budgets 4.13 FTE after automating the simple band. The work actually remaining needs 9.38. That is a 2.27 times shortfall, or 5.25 people, and it shows up as an overloaded exception desk roughly one quarter after go-live.
Cost pressure makes the miss expensive rather than merely awkward. ATRI’s 2026 report puts the industry-average cost of operating a truck at $2.336 per mile in 2025, a record for the series and 3.4% above the prior year, so an exception that sits in a queue because the desk is understaffed converts into miles that were not planned. And the carrier base an enterprise coordinates across is fragmented, with the American Trucking Associations reporting almost 580,000 active US motor carriers as of June 2025, of which 91.5% operate 10 or fewer trucks, which is a large share of the complexity that lands in the complex band.
How to forecast automation effort correctly
1. Measure your exception mix by handling time, not by count
Most operations report exception counts and no durations. Instrument handling time per exception type for a representative period, then build the distribution. Without it every forecast below is a guess, and the distribution is usually more skewed than teams expect, because the rare complex cases absorb a disproportionate share of the day.
Two practical notes on doing this. Measure elapsed handling time rather than touch time, because a complex exception typically involves waiting on a carrier, a client or a warehouse, and that waiting occupies the handler’s attention even when it does not occupy their hands. And measure across a full cycle including at least one disrupted week, since a sample drawn from clean weeks will understate the complex band, which is exactly the band that determines the answer.
2. Establish which band automation actually takes
Ask the vendor, and then verify against a sample, which specific exception types the system resolves without a person. Automation does not remove a uniform slice across the distribution. It removes the tractable end, and the boundary of that end is a property of your data quality and your process, not of the vendor’s capability claim.
3. Forecast effort, not volume
Multiply the remaining cases by their own handling times rather than scaling total effort by the automation rate. This is a five-minute calculation that changes the number materially: on the mix above it is the difference between 4.13 FTE and 9.38. Do it per exception type rather than in aggregate, because the complex band drives the result and an average hides it. Put both figures in the business case and label which is which, because the case-volume number is the one everyone will quote.
| Also Read: ROI of Logistics Technology Investments |
|---|
4. Budget for skill, not only for headcount
The residual needs more capable people, which is a different budget line from fewer people. An exception desk handling only complex cases requires deeper product knowledge, more authority to make commercial calls and better judgment under time pressure. Operations that cut headcount and hold the skill profile flat discover the gap during their first disrupted week.
5. Protect the apprenticeship you are about to automate away
This is Bainbridge’s point and it is the one most often missed. Before automation, a dispatcher saw around 700 simple cases a day, and that volume was how the skill was built. Automate it and the training ground disappears while the remaining work demands more skill than before. Replace it deliberately, through structured exposure, shadowing on complex cases, or rotation, because the pipeline that used to produce senior exception handlers no longer exists.
6. Instrument the residual so the mix stays visible
Automation coverage changes as the system learns and as your data improves, which moves the boundary between automated and residual. Report the mix monthly, not the automation rate alone, so that a rising mean handling time is read as the expected consequence of progress rather than as a performance problem on the desk.
What falls and what rises when automation lands
| Measure | Direction | Why |
|---|---|---|
| Exception case count reaching people | Falls sharply | The tractable band is removed |
| Total human effort | Falls modestly | The remaining cases are the expensive ones |
| Mean handling time per case | Rises | The mix shifts toward complexity |
| Skill required per handler | Rises | Only hard cases reach a person |
| Routine practice available | Falls to near zero | The simple cases were the practice |
| Monitoring load | Rises | Supervising a system that is usually right |
| Planning and rekeying labor | Falls genuinely | This is the part automation truly removes |
The last row matters for fairness. Planning labor does fall, and that saving is real. The error is applying the same logic to exception labor, where volume and effort move at very different rates.
Five questions to ask before signing an automation business case
Does the forecast use effort or case volume? If the headcount line was derived by multiplying current staff by one minus the automation rate, it is wrong by roughly the ratio between those two quantities.
Which exception types are actually automated? A named list, verified against a sample of your own data, not a category. The boundary determines everything downstream.
What is the projected mean handling time after go-live? If nobody has computed it, the desk has not been sized. Expect it to roughly double at a 70% automation rate.
| Also Read: Automated Dispatch Software: Complete Guide |
|---|
What is the plan for skill development once routine volume disappears? Ask specifically, because it is nobody’s default responsibility and it fails quietly over 18 months.
How is the residual mix reported? Monthly composition, not just an automation percentage, so that rising average difficulty is interpreted correctly.
What this looks like in enterprise deployments
A leading North American retailer running multimodal logistics automation across several hundred stores reduced manual dispatch by more than 80% while resolving exceptions in under two hours, alongside 99%-plus on-time store delivery and 95%-plus route compliance, after replacing six legacy systems. Both halves belong in the same sentence. An 80% reduction in manual dispatch is the automation result, and a sub-two-hour exception resolution is the residual capability that had to exist for the first number to be safe. Operations that deliver the first and not the second have automated the easy work and left the hard work unresourced, which shows up as exceptions aging in a queue while the automation dashboard reports a healthy number. The sequencing lesson is that exception capability has to be built before or alongside the automation, not after it, because the moment the simple cases stop arriving is the moment the desk composition changes.
A Fortune 50 parcel operation running centralized dispatch across a 120-country network and 51 sites lifted weekly execution adherence from 75% to 92% and surfaced more than $14 million of unused capacity, including $565,000 at a single site. Adherence is the useful metric here because it measures whether the plan survived the day, which is where residual exceptions actually land. Moving it 17 points across a network that size is a statement about exception capability as much as about planning quality.
Four mistakes operations make on automation forecasting
Deriving headcount from the automation rate. It is the most available number and the wrong one. Effort is the product of remaining cases and their own handling times, and on a skewed mix those diverge by more than a factor of two.
Treating the residual as a temporary backlog. It is not a transition state that resolves as the system matures. Further automation raises the difficulty of what remains, so the effect strengthens as the program succeeds.
Cutting headcount and seniority together. The two move in opposite directions. Fewer people is consistent with automation; less capable people is not, because only the hard cases now reach a human.
Reporting automation rate as the health metric. It describes the input. The mix of what remains, and the mean handling time inside it, describe whether the operation can actually absorb what the automation hands back. A program reporting 94% automation with no view of the residual is reporting the half of the system that is working.
How Locus handles the work automation hands back
Locus, the world’s first Decision-Intelligent, Agentic TMS, runs more than 12 million automated decisions a day, and the governance layer exists because the residual has to be a designed set rather than whatever fell through. Autonomy Levels run per agent and per domain, with L1 requiring human approval, L2 acting inside guardrails and L3 operating autonomously, which means the operation decides which decision classes reach a person instead of discovering that after go-live.
Explainability and Traceability record the trigger, context, reasoning, action and outcome for each decision, which is what lets an exception arrive with its context attached rather than as a bare alert, and that difference is most of the handling time on a complex case. Human Review supplies the escalation pathway itself. Because the route planning system re-optimizes in roughly two minutes against more than 250 real-world operating constraints, many exceptions can be answered by re-planning rather than by manual reconstruction, which is the main mechanism for holding down the handling time of the residual rather than merely shrinking its count.
Two boundaries belong here. Locus does not size your exception desk, and no platform can, because the right number depends on the mix measured in your own operation rather than on a benchmark. And the apprenticeship problem is organizational rather than technical. A platform can deliver context-rich exceptions that are easier to learn from, but deciding how your next generation of dispatchers builds judgment when the routine volume is gone is a workforce decision that belongs to you.
Locus supports more than 360 enterprise customers across 30-plus countries, with over 1.5 billion deliveries optimized, more than $320 million in documented client logistics savings and 99.99% uptime. It has been recognized by Gartner for seven consecutive years, featured in the 2026 Hype Cycle for Supply Chain Execution and Logistics Technologies, named a Leader in TMS by QKS Group (SPARK Matrix), and ranked #1 in Route Planning on G2’s 2026 Best Software Awards.
In October 2025, Ingka Investments, the investment arm of Ingka Group, the world’s largest IKEA retailer, acquired Locus. Locus continues to operate independently.
So how much work does logistics automation actually remove? Less than the automation rate suggests, and the gap is predictable rather than disappointing. Automation takes the tractable cases first because tractability is what makes them automatable, so on a mix of 1,000 daily exceptions at 70% simple, 25% moderate and 5% complex, automating the simple band removes 70% of cases and 32% of effort while mean handling time rises from 6.6 to 15.0 minutes. A plan that scales headcount with case volume budgets 4.13 FTE where 9.38 are needed. Push automation to 95% and handling time rises 6.06 times, because only the complex band is left. The fix is to forecast effort rather than volume, budget for skill rather than only headcount, and deliberately replace the routine practice that automation removes, which is the point Bainbridge made in 1983. Locus supports the residual through Autonomy Levels that make it a designed set, Explainability and Traceability that deliver exceptions with their context attached, and a route planning system that re-optimizes in roughly two minutes so many exceptions are answered by re-planning rather than by hand. Request a Locus assessment to model your own exception mix before you size the desk.
Frequently Asked Questions
Why does automating 70% of exceptions not remove 70% of the work? Because automation takes the tractable cases first, and tractable means fast. On a mix of 70% simple at three minutes, 25% moderate at ten and 5% complex at forty, removing the simple band removes 70% of cases but only 32% of total minutes, since the expensive cases are all still there.
What happens to handling time as automation increases? It rises, because the mix shifts toward complexity. Mean handling time goes from 6.6 minutes to 15.0 at a 70% automation rate, a 2.27 times increase, and to 40.0 minutes at 95%, which is 6.06 times the original average.
How wrong is a headcount forecast based on case volume? On the mix above, it budgets 4.13 FTE against an actual requirement of 9.38, a shortfall of 2.27 times or 5.25 people. The error is systematic and always in the same direction, so it is predictable and correctable before go-live.
Does this mean automation is not worth doing? No. Planning and rekeying labor genuinely falls, fleet utilization improves and cost per delivery drops. The error is applying the same linear logic to exception labor, where case volume and effort move at very different rates. The return is real and smaller than a naive model shows.
What is the apprenticeship problem? Before automation a dispatcher handled around 700 simple cases a day, and that routine volume was how judgment was built. Automation removes it while raising the skill the remaining work demands, so the training pipeline disappears exactly when the requirement goes up. Bainbridge identified this in 1983 and concluded that operators need more training after automation, not less.
Should we reduce exception desk headcount after automating? Usually yes, but by less than the automation rate and with a higher skill profile. Cutting headcount and seniority together is the common failure, because only hard cases now reach a person and those need more authority and product knowledge, not less.
What should we report instead of automation rate? The residual mix and its mean handling time, monthly. Automation rate describes the input. The composition of what remains tells you whether the desk can absorb what the system is handing back, and a rising average handling time should be read as evidence the automation is working.
Ishan, a knowledge navigator at heart, has more than a decade crafting content strategies for B2B tech, with a strong focus on logistics SaaS. He blends AI with human creativity to turn complex ideas into compelling narratives.
Related Tags:
General
An Ultimate Guide to Smart 3PL Delivery Orchestration for Enterprise Logistics Leaders & 3PL Operations Teams
Learn how 3PL delivery orchestration improves control, cost efficiency, and performance. Explore how Locus helps enterprises orchestrate logistics at scale.
Read more
Logistics Automation & Orchestration
Nobody is 70% Automated: Node Coverage and Flow Coverage Are Different Numbers
Flow automation is node automation raised to the path length. A network 70% automated by node runs 24% of its four-node flows end to end. The sequencing fix.
Read moreInsights Worth Your Time
Why Automating 70% of Exceptions Only Removes 32% of the Work