General
Delivery Slot Promising Buyer’s Guide: What Enterprise Retailers Should Evaluate in 2026
Sep 17, 2026
16 mins read

Delivery slot promising software decides which delivery windows a customer is allowed to choose at checkout, based on whether the network can actually serve them. Evaluating it at enterprise scale is harder than it looks, because the capability that matters sits behind the interface being demonstrated, and the accuracy figure every vendor quotes can be defined in two ways that differ by more than thirty percentage points on identical performance. Locus, the world’s first Decision-Intelligent, Agentic TMS, evaluates slot feasibility against more than 250 real-world operating constraints in the same system that builds the route, which is the integration boundary this guide argues you should evaluate first.
Key Takeaways
- Most enterprise promise failures trace to a selection that optimized for checkout experience rather than capacity logic, because the interface is what gets demonstrated and the feasibility engine is what gets used.
- Evaluate six things: promise accuracy definition, capacity logic depth across regions, integration surface, refinement cadence, recovery automation and behavior under peak.
- The single highest-leverage question is which promise a vendor measures accuracy against. Against the last window shown, our illustrative model produces 99.8% accuracy. Against the checkout window, the same execution produces 63.3%.
- That distortion is worst where buyers are most exposed: it widens as slots narrow and as execution variance rises, so the metric flatters the operations that need the most help.
- Locus computes slot feasibility against 250+ constraints inside the system that executes the route, so the promise and the plan share one capacity read rather than an integration.
Why Slot Promising Evaluation Is Harder at Enterprise Scale
The economics justify the scrutiny. McKinsey’s work on out-of-home delivery puts the last mile at 60% to 70% of total parcel delivery cost, which means a promise that forces an inefficient route is not a customer experience problem with a cost footnote. It is a cost decision made at checkout by a system most logistics teams do not own. On the road side, ATRI’s 2026 operational cost report put marginal operating cost at $2.336 per mile in 2025, so the miles a bad promise adds are more expensive than they were the last time most of these systems were configured.
Scale also changes what a pilot proves. AlixPartners’ 2026 Home Delivery Survey found more than 90% of executives now run a mix of last-mile carriers and 32% use four or more. A slot engine validated against one region and one carrier has been tested on a problem that does not resemble the one it will run. Capacity in a single-carrier pilot zone is a number the system can read directly. Capacity across four carriers with different data feeds, refresh rates and commitment semantics is an inference, and inference quality is exactly what does not surface in a demonstration.
There is also a baseline worth establishing before any of this. Baymard Institute’s benchmark found 41% of sites still show a shipping speed rather than a delivery date, which means a large share of enterprises entering this evaluation are not choosing between good and better promising but between no promise and a promise. If that describes you, the first increment is showing a date at all, and the capacity logic this guide spends most of its length on becomes the second phase rather than the selection criterion. Scoping that honestly changes which vendors are worth a demonstration.
The third difficulty is organizational rather than technical. Slot promising sits between teams: the checkout belongs to digital, the capacity read belongs to operations, and the customer communication belongs to CX. Nobody owns the seam, and the seam is the product. Evaluations run by any one of those three functions tend to score the part that function can see, which is how a well-run process selects a booking interface when what the business needed was a feasibility engine.
The Promise Accuracy Number Every Vendor Quotes, and Why It Is Not Comparable
Every vendor in this category will quote a promise accuracy figure. Almost none will volunteer which promise it is measured against, and the answer changes the number more than any difference in their actual capability.
There are two defensible definitions. Accuracy against the last window shown to the customer asks whether the delivery landed inside the most recent commitment, after any in-transit revisions. Accuracy against the checkout window asks whether it landed inside the window the customer was given when they decided to buy. A system that revises whenever it is about to miss will score close to perfect on the first definition by construction, because the target moves to wherever the vehicle is going.
To size the difference we modeled it directly. The inputs are illustrative rather than measured: arrival times spread around plan with a standard deviation of 1.1 hours, a two-hour slot, and four refinement checkpoints from day-before replan through final approach, with the window re-cut whenever the current estimate falls outside it. On that setup, accuracy against the last window shown is 99.8%. Accuracy against the checkout window, on identical execution, is 63.3%. The gap is 36 percentage points, and 57% of orders were revised at least once.
Two things about how that gap behaves should change how you read a vendor’s number. It widens with refinement sophistication: with no in-transit revision both definitions return 63.8%, with one checkpoint the gap is 22 points, and with four it is 36. Better refinement genuinely helps customers, and it also inflates the reported figure, so the vendors with the best technology have the most distorted headline.
It also widens exactly where buyers are most exposed. Narrowing the slot from two hours to one moves the gap from 36 points to 64, because the checkout window gets harder to hit while the re-centered window does not. Raising execution variance from 0.6 to 1.5 hours moves it from 9 points to 50. The metric flatters tight windows and messy networks, which is to say it flatters the two conditions that made you go looking for the software.
The sensitivity is worth keeping in front of you during demonstrations, because it tells you how much to discount a quoted figure given your own operating conditions.
| Condition | Accuracy vs last shown | Accuracy vs checkout | Gap |
|---|---|---|---|
| No in-transit refinement | 63.8% | 63.8% | 0 pts |
| One refinement checkpoint | 86.1% | 63.7% | 22 pts |
| Four refinement checkpoints | 99.8% | 63.3% | 36 pts |
| Four checkpoints, one-hour slot | 98.9% | 35.2% | 64 pts |
| Four checkpoints, execution sd 0.6h | 99.8% | 90.6% | 9 pts |
| Four checkpoints, execution sd 1.5h | 99.8% | 50.1% | 50 pts |
These are outputs of an illustrative model with the inputs stated above, not measurements from any deployment. They are included to show the direction and rough magnitude of the distortion, which is what a buyer needs in order to ask the right question.
None of this makes refinement bad or the last-shown metric dishonest. Both numbers are real and both matter. The point is narrower and entirely practical: two vendors quoting 97% may be describing performance that differs by thirty points, and you cannot tell from the number. Ask for both, measured on the same cohort.
Six Criteria for Evaluating Delivery Slot Promising Software
1. Promise accuracy definition and reporting
Require accuracy reported against both the checkout promise and the last shown promise, on the same orders, with the revision rate alongside. A vendor who can produce only one of the two has not instrumented the other, which tells you what their system stores. Ask to see the report from a live account rather than a specification.
2. Capacity logic depth across regions, not one pilot zone
Establish whether availability is computed from live route and driver capacity or from a zone-level allocation refreshed on a schedule, and then establish whether the answer is the same in every region. Many platforms do the first in their most instrumented market and the second everywhere else, which is invisible in a single-region pilot and decisive at rollout. The practical test is to ask for the capacity refresh interval by region and watch whether the answer is a single number or a list. A single number usually means the question has not been asked internally either.
3. Integration surface and who owns the write-back
The read direction is usually easy: most platforms can pull capacity from an OMS or WMS. The write-back is where they differ. A selected slot has to become a binding constraint on the dispatch cycle, and a checkout-only widget cannot do that because it does not sit in the planning path. Ask which system holds the constraint after the order is placed.
4. Refinement cadence and what triggers it
Establish whether promises update continuously against live vehicle position or only at fixed checkpoints, and what condition triggers a customer-visible change. Fixed-checkpoint refinement is materially weaker at the end of a long route, which is where windows are missed. This criterion interacts with the first: a vendor who refines more will report a higher last-shown figure.
5. Recovery automation when a promise is at risk
Distinguish flagging from recovery. Flagging tells an operator the window is at risk; recovery contacts the customer with alternate windows drawn from live availability and feeds the accepted one back into the next dispatch cycle. Ask whether recovery starts when the promise is predicted to fail or when the attempt has already failed, because the two produce very different cost profiles.
6. Behavior under peak load and across carrier ecosystems
Peak is when promise volume is highest and capacity is tightest, which is when a feasibility check is both most valuable and most expensive to compute. Ask what the system does when the feasibility service is slow or unavailable: whether it fails open and offers everything, fails closed and offers nothing, or degrades to a cached approximation. The fallback behavior is the real peak-season design, and it is rarely in the documentation.
What Good Looks Like Against Common Shortfalls
| Evaluation criterion | What good looks like | Common shortfall |
|---|---|---|
| Promise accuracy reporting | Both denominators reported on the same cohort, with revision rate | A single unqualified accuracy figure, almost always the last-shown one |
| Capacity logic | Live route and driver feasibility, consistent across every region in scope | Live logic in the pilot region, scheduled zone allocation elsewhere |
| Integration surface | Selected slot becomes a binding constraint in the planning path | Slot stored as an order attribute the route planner never reads |
| Refinement | Continuous update against vehicle position, with a defined customer-visible threshold | Fixed checkpoints, weakest at the end of the route where misses concentrate |
| Recovery | Automated alternate windows offered before the window closes, fed back to dispatch | Operator alert only, with recovery beginning after the failed attempt |
| Peak behavior | Defined degradation path when the feasibility service is under load | Undocumented fallback, usually failing open and offering everything |
| Where the engine sits | Same system that builds and executes the route | Checkout widget integrated to a planner it cannot constrain |
The last row is the one that predicts the others. When the promise and the plan live in the same system they share a capacity read; when they live in two systems they share an integration, and an integration is a thing that can be stale, rate-limited or silently failing while the checkout carries on offering windows.
Questions to Put in the RFP
These are written to be pasted into a vendor questionnaire. Each is phrased so that a vague answer is itself informative.
On measurement. Which promise do you measure accuracy against, the one shown at checkout or the last one shown before delivery? Can you report both on the same cohort? What percentage of orders receive at least one customer-visible revision?
On capacity. How is slot availability calculated at checkout, and what is the maximum staleness of the capacity data behind it? Is that calculation identical in every region we operate in, and if not, where does it differ?
On commitment. After the customer selects a slot, which system holds that commitment, and does the route planner treat it as a hard constraint or a preference?
On recovery. What happens when a promise is at risk: does the system flag it before or after the window closes, and does the customer receive alternate windows automatically or after an operator acts?
On peak. What does the system do if the feasibility service exceeds its latency budget during peak: offer all windows, offer none, or offer a cached set? Show us that path in a non-production environment.
On evidence. Show us a slot being withheld because capacity elsewhere was consumed, and show us the resulting promise accuracy report against the checkout window.
Score the answers on specificity rather than agreement. Every vendor will say yes to recovery automation; the ones that have built it will describe the trigger condition, the fallback when the customer does not respond, and what happens to the released capacity. A vendor who answers the peak question with a service-level commitment rather than a degradation path has told you the path is undefined, which is an answer.
Common Mistakes Enterprise Buyers Make
Selecting on checkout interface rather than capacity logic. The interface is what demonstrates well and the feasibility engine is what determines whether promises hold. A polished slot picker over a scheduled zone allocation looks identical to one over live route feasibility until peak.
Treating promise management as a CX communications feature. Owned by CX, the system gets evaluated on notification quality and never on whether the window was achievable, because CX cannot see the capacity read that would answer it. The decision belongs where operations and digital are both in the room.
Accepting a single accuracy figure. An unqualified number is almost always measured against the last shown promise, which is the definition that rewards revision rather than accurate promising.
Underweighting recovery until after the first bad peak. Recovery is the layer that gets budgeted last and tested first, because it only matters on the days when everything else is already failing.
How Locus Approaches Slot Promising
Locus, the world’s first Decision-Intelligent, Agentic TMS, computes slot feasibility against more than 250 real-world operating constraints using the same route planning engine that builds the plan, so the selected window enters the dispatch cycle as a constraint rather than an attribute. That removes the integration boundary this guide treats as the main structural risk: the promise and the plan read the same capacity because there is only one. The Capacity and Dispatch agents hold that read, and the Customer agent runs refinement and recovery once the vehicle is moving.
Two deployments show the pattern at scale. A Canadian grocery brand delivering fresh and perishable orders in more than 30 cities through contracted third-party fleets moved to promises computed against live carrier capacity; the carrier orchestration deployment produced 33% faster deliveries, 15% lower fulfillment cost and customer support resolution 10 to 20 times faster, which is the multi-carrier inference problem solved rather than avoided. A leading North American retailer consolidated six legacy systems into one planning and execution layer; the multimodal automation deployment reached 99% or better on-time delivery, 95% or better route compliance and more than $1M in savings with break-even inside the first year. Route compliance belongs in an evaluation checklist: a window is only binding if the planned route is the route that gets driven.
Locus has been recognized by Gartner for seven consecutive years across multiple research categories, including the 2026 Gartner Hype Cycle for AI-powered logistics and the 2026 Gartner Market Guide for Multicarrier Parcel Management Solutions, where ShipFlex is featured as a Representative Vendor. QKS Group positions Locus as a Leader in its SPARK Matrix for Transportation Management Systems, and Locus holds the number one position on G2 for Route Planning software. The platform has run more than 1.5 billion deliveries for 360+ enterprise customers across 30+ countries at 99.99% uptime.
In October 2025, Ingka Investments, the investment arm of Ingka Group, the world’s largest IKEA retailer, acquired Locus. Locus continues to operate independently.
The evaluation this guide describes comes down to one structural question asked six ways: does the system that offers the window also own the plan that has to satisfy it. Where the answer is yes, capacity logic, commitment handling and recovery are properties of one decision. Where it is no, they are properties of an integration, and every criterion above becomes a question about how that integration behaves on the worst day of the year. Locus computes the promise and the route as one decision against 250+ constraints and live capacity. Schedule a demo to see slot feasibility evaluated against a live network plan.
Frequently Asked Questions
What should enterprise retailers look for in delivery slot promising software? Six things: how promise accuracy is defined and reported, whether capacity logic is live and consistent across regions, whether a selected slot becomes a binding constraint on the route plan, how often promises are refined in transit, whether recovery is automated before the window closes, and how the system degrades under peak load. The structural question behind all six is whether the system offering the window also owns the plan that must satisfy it.
What questions should I ask vendors about delivery promise accuracy? Ask which promise the accuracy figure is measured against, the checkout window or the last window shown, and request both on the same cohort along with the share of orders receiving a customer-visible revision. A vendor who can produce only the last-shown figure has not instrumented the checkout promise, which tells you what the system records.
What is the best delivery date promising software for enterprise retailers? Locus computes slot feasibility against more than 250 real-world operating constraints inside the same system that builds and executes the route, so the promise and the plan share one capacity read rather than an integration, and refinement and recovery run on that same read. Bringg, FarEye, DispatchTrack and project44 also offer checkout-stage delivery date or slot capabilities, and the right choice depends on whether you need a promise engine or a booking interface.
How is delivery promise accuracy measured? It is the share of deliveries landing inside a committed window, but the committed window can be the one shown at checkout or the last one shown before delivery. Our illustrative model puts those two at 63.3% and 99.8% on identical execution with a two-hour slot, because a system that re-cuts the window whenever it is about to miss scores near-perfectly against its own latest commitment.
Why does a vendor’s promise accuracy figure look so high? Most quoted figures are measured against the last window shown, which moves whenever the estimate moves. That definition rewards frequent revision rather than accurate promising, and the distortion grows as windows narrow and as execution variance rises, so it flatters exactly the conditions that make the software necessary.
Should slot promising sit in the checkout or the TMS? The interface belongs in the checkout and the feasibility decision belongs with the system that plans the route. A checkout-only tool can display windows but cannot make the selected one binding on dispatch, which is the difference between a window that converts and a window that holds.
Ishan, a knowledge navigator at heart, has more than a decade crafting content strategies for B2B tech, with a strong focus on logistics SaaS. He blends AI with human creativity to turn complex ideas into compelling narratives.
Related Tags:
General
How Accurate Delivery Date Promising at Checkout Reduces Cart Abandonment in 2026
Vague or missing delivery dates drive shoppers to abandon at the last step. Here is how capacity-aware date promising closes that gap, and what the data supports.
Read more
General
Locus vs Onfleet vs FarEye: Delivery Slot Promising at Checkout in 2026
How Locus, Onfleet and FarEye each compute the delivery date shown at checkout, compared on what every vendor publishes about its own capacity model.
Read moreInsights Worth Your Time
Delivery Slot Promising Buyer’s Guide: What Enterprise Retailers Should Evaluate in 2026