Five Questions to Ask Before You Believe Another AI Pitch for TMS
QKS Group analyst sessionSeptember 20264 min read
We recently concluded our webinar “Five Questions to Ask Before You Believe Another AI Pitch for TMS.” We asked these questions to Sanjeevi Cuddalore, Associate Vice President and Principal Advisor at QKS Group, and Nithin B., Senior Analyst at QKS Group. Below is an excerpt from the insightful conversations we had.
Sanjeevi CuddaloreAssociate Vice President & Principal Advisor, QKS GroupNithin B.Senior Analyst, QKS Group
Question 1What the word actually covers
What does autonomous actually mean in this market?
A Tuesday at 6am. A driver calls in sick, no replacement can be made in time, and 40 stops are in jeopardy, 11 of them early-morning deliveries. Before asking whether the software can recommend or act, ask whether it even knows something has gone wrong. At 6:05 the driver texts the dispatcher. That is not a structured event in the system.
Detection is the first of four levels: record the issue, alert the right person, recognize and prioritize it, then recommend an action and execute it. The market is strong on the first two and increasingly capable at the third. True autonomous execution, the fourth, is still rare.
The word “recommendation” needs care, because it covers a lot of ground. There is a large difference between re-running an optimizer and saying: here are the options, here is what each protects, here is the business trade-off. Routing is not the difficult part. The hard part is knowing which customer commitment matters more, and what to sacrifice when something goes wrong.
The real measure of autonomy is the point where the system stops deciding and a person has to step in. That boundary matters more than the label on the technology.
Accountability matters just as much as that boundary. The moment a system moves from recommending to deciding, somebody owns that outcome, and usually that ownership is not clear. Then something goes wrong, everyone wants to know who approved the autonomy, and the organization reverts to manual. The statement that surfaces is “AI did not work.” Usually the technology worked and the accountability model did not.
Most autonomy today sits in last-mile dispatch, high-density routes and short time horizons where decisions can usually be reversed. Re-tendering a load is a different risk: once a carrier accepts, you have affected a commercial relationship and it is much harder to undo. In a 3PL the same decision can hit a customer SLA and trigger penalties. Autonomy is also knowing how safely a system fails.
Question 2Explainability, and the EU AI Act
If an automated decision changes a customer commitment, what do I need to see before I can stand behind it?
Explainability is where trust starts, and it is also where the market has a gap. Most products have a change log. What buyers need is a decision log.
A change log tells you what happened. A decision log tells you why it happened: what triggered the change, what information the system used, what constraint it considered, what action it chose, and whether a human approved it.
Ask the vendor
Replay Tuesday's decision with the same input and policy. Do you get the same answer? If the vendor cannot reliably reproduce it, you cannot defend it later.
The explanation has to be in business terms. A store manager does not want to hear that AI optimized the route. They want to know why the delivery was rescheduled: the stop was moved because keeping it in the original slot would have affected the other three delivery commitments.
On the EU AI Act. There is an assumption the high-risk obligations came into force in August 2026. They did not. They were deferred to December 2027. What changed is the deadline, not the requirement, and a contract signed today will very likely still be running in 2027. So that conversation belongs in the contract you are negotiating now, not in 2027.
Buyers naturally look at the routing engine, but that is not where to start. The Act covers AI that allocates work, and AI that monitors or evaluates worker performance. In transportation that means driver dispatch, driver allocation and driver scoring. If you use gig workers or owner-operators in Europe, legal should look at that layer specifically. Data comes after that, and it is where you should have the most questions and where you should demand the answers in writing.
Ask the vendor, and get it in writing
Where exactly is my data processed: which region, which cloud, and what changes when a third-party model provider is introduced?
Are you using my data? Ask separately about training, fine-tuning, evaluation, and aggregated or de-identified use. A “no” to training does not answer the other three.
Who can access it? Can vendor support reach production during an incident, and if so, is that logged and am I notified?
How long do you retain the decision record, can I export it, and what happens when the contract ends? Disputes and audits outlive contracts. If your audit trail disappears when you leave, you were renting it, not owning it.
What is the evidence? A responsible-AI page on a website is good, but it is still marketing material. Ask for an actual exportable decision record from a live customer environment, redacted.
Question 3Guardrails, and who can move them
What should the system never be allowed to do on its own, and who is allowed to change that?
Q4 is peak season, and a good system should be able to re-route freely. But not all decisions are equal. Say twelve accounts have tight contractual windows with real penalties attached. Those are orders a business will not let anyone touch without approval.
Avoid autonomy at a system level. Autonomy should sit at a decision level. Re-sequencing a route, changing a delivery commitment, re-assigning a driver, re-tendering a load and spending money are completely different risk levels. Do not give all five the same permission.
Start where decisions are frequent, reversible and carry low consequences: exception triage, ETA re-calculation, proactive notifications, re-sequencing routes before they start. If AI wants to spend money, change a contractual commitment, make a safety-related decision, or allocate and score driver work, a human needs to be involved. Do not give AI the company credit card on day one.
Guardrails cannot be configured at implementation and then forgotten, because six months later when something goes wrong, everyone will want to know what the guardrail was. Four things matter:
The four guardrail questions
Can I test the policy before production, against historical scenarios, so I know what rule applied when a decision was taken?
Can I roll it back?
Who is allowed to change it? The person under pressure at 6am should not be able to quietly loosen a guardrail because it is slowing things down. Then you do not have a guardrail. You have a suggestion.
Is every policy change captured, attributed and reviewable, exactly like the automated decision itself?
Governing the guardrail is as important as having one.
Question 4Demo against production
Will what I saw in the demo survive in my production environment?
Three checks. Is it officially released and versioned? If there is no release note or version number, treat it as a prototype. Is any other customer running it at scale, and if so, since when? And what is native versus what is partner-delivered?
That third one matters most, because mapping, traffic data, ETA, address validation and even the AI models may come from partners. When something breaks during peak, you need to know who provides what, and who is solving it for you.
Ask the vendor
Stop asking for a demonstration. Ask for evidence that moves the conversation from “we can demonstrate it” to “we can prove it in production.”
One test you can run yourself is to feed the vendor's system a bad address, a missing geocode, a duplicate order and a conflicting service rule, then watch what happens. Does the system flag the problem and escalate it to a human, or does it confidently make a decision using the bad data?
Then turn the same scrutiny on yourself. Assess your own data before you select a vendor, because what your data supports should influence what you buy. Does your geocode point to the real entrance or somewhere on the street? Do you know real service times per stop? Are your delivery windows realistic? Are driver qualifications available as structured data, or are they sitting in the dispatcher's head? Those determine whether AI works in production.
Question 5Ownership and the 90-day case
Who owns this after go-live, and what will the CFO actually see in 90 days?
Operations should own the decision policy and the KPIs. IT should own the environment. Neither should hand this to the implementation team, because the project will end and accountability will not.
Before expanding autonomy, set the criteria upfront. Has it run volume through a full seasonal cycle? Is the override rate stable? Have we avoided unplanned rollback? Expand on evidence, not enthusiasm. Write it down while everyone is calm, because after an incident everyone remembers the policy differently.
The business cases that survive contact with a CFO get four things right, and the ones that struggle are always missing one:
The four that decide whether the case survives
A baseline actually measured, not re-created six months later from memory. If you do not know cost-per-drop, planner hours or whichever KPIs you are trying to improve before go-live, then nine months later you will not be discussing whether the investment worked. You will still be debating which number is correct.
Every benefit has an owner in the business, not the project team, because that team eventually moves on.
Realistic timing. No network-level benefit in the first quarter simply because the AI went live.
The full cost, not the license cost.
Then comes attribution, which is what the CFO actually has in mind. They will ask how much of the improvement came from the technology, because along the way you also cleaned the data, changed the processes and trained people, and all of those improve performance on their own. Decide beforehand who the numbers are attributed to.
The first signals at 90 days will not be a reduction in freight cost. Expect improvements in planner and dispatcher productivity, from faster route planning, less manual intervention and more exceptions handled without a human. With proactive notifications enabled, expect fewer inbound customer calls. Vehicle utilization improvement, more stops per route and delivery success rate improvement usually surface after three quarters. Avoid putting long-term benefits into a 90-day case.
Interested in knowing how Locus balances autonomy, audit, governance and human-in-the-loop execution?
Across all five answers the analysts highlighted a critical point. The question is rarely whether the AI can make the decision. It is whether you can see why it did, roll it back, and say who owns the outcome when it is wrong.
Answers recorded during a live session with Sanjeevi Cuddalore, Associate Vice President and Principal Advisor, QKS Group, and Nithin B., Senior Analyst, QKS Group, hosted by Locus. Compressed for length; substance unchanged. The EU AI Act deferral the analysts describe is confirmed by the European Commission: the Annex III high-risk obligations, which cover AI used in employment and worker management, were postponed from 2 August 2026 to 2 December 2027↗.
Ishan Bhattacharya
Lead - Content
Ishan, a knowledge navigator at heart, has more than a decade crafting content strategies for B2B tech, with a strong focus on logistics SaaS. He blends AI with human creativity to turn complex ideas into compelling narratives.