Ingka Group acquires Locus! Built for the real world, backed for the long run. Read here>Read the full story>
Ingka Group acquires Locus! Built for the real world, backed for the long run. Read the full story
locus-logo-dark
Schedule a demo
Locus Logo Locus Logo
  • Platform
    • Transportation Management System
    • Last Mile Delivery Solution
  • Products
    • Fulfillment Automation
      • Order Management
      • Delivery Linked Checkout
    • Dispatch Planning
      • Hub Operations
      • Capacity Management
      • Route Planning
    • Delivery Orchestration
      • Transporter Management
      • ShipFlex
    • Track and Trace
      • Driver Companion App
      • Control Tower
      • Tracking Page
    • Analytics and Insights
      • Business Insights
      • Location Analytics
  • Industries
    • Retail
    • FMCG/CPG
    • 3PL & CEP
    • Big & Bulky
    • Other Industries
      • E-commerce
      • E-grocery
      • Industrial Services
      • Manufacturing
      • Home Services
  • Resources
    • Guides
      • Reducing Cart Abandonment
      • Reducing WISMO Calls
      • Logistics Trends 2024
      • Unit Economics in All-mile
      • Last Mile Delivery Logistics
      • Last Mile Delivery Trends
      • Time Under the Roof
      • Peak Shipping Season
      • Electronic Products
      • Fleet Management
      • Healthcare Logistics
      • Transport Management System
      • E-commerce Logistics
      • Direct Store Delivery
      • Logistics Route Planner Guide
    • ROI Calculator
    • Product Demos
    • Whitepaper
    • Case Studies
    • Infographics
    • E-books
    • Blogs
    • Events & Webinars
    • Videos
    • API Reference Docs
    • Glossary
  • Company
    • About Us
    • Global Presence
      • Locus in Americas
      • Locus in Asia Pacific
      • Locus in the Middle East
    • Analyst Recognition
    • Careers
    • News & Press
    • Trust & Security
    • Contact Us
  • Customers
en  
en - English
id - Bahasa
Schedule a demo
  1. Home
  2. Blog
  3. The Agentic TMS RFP Scorecard: 30 Questions to Separate Real AI from Rebranded Legacy in 2026

General

The Agentic TMS RFP Scorecard: 30 Questions to Separate Real AI from Rebranded Legacy in 2026

Avatar photo

Ishan Bhattacharya

Aug 4, 2026

10 mins read

Key Takeaways

  • Standard TMS RFP templates evaluate features designed for a pre-AI era. Every vendor supports multi-stop consolidation, EDI 214, and a stated uptime SLA, so every vendor scores the same and the template tells you nothing.
  • Evaluating a system that decides requires architectural questions: where decisions come from, how the system learns, what governs its autonomy, what it can execute, how it behaves under disruption, and what evidence backs its claims.
  • This scorecard contains 30 questions across those six sections, each scored 0 to 3 on evidence strength, for a maximum of 90 points. Vendors below 45 are running rebranded rules regardless of what the marketing says.
  • Send the questions verbatim, score the answers on evidence rather than fluency, and require the strongest claims to be demonstrated on your data, not the vendor’s demo dataset.

Why Your Current TMS RFP Can’t Detect AI

The standard enterprise TMS RFP is a feature checklist, and feature checklists were the right tool for the systems they were built to evaluate: rules-driven platforms whose value lived in coverage. Does it support multi-stop consolidation? EDI 214? A 99.9% uptime SLA? Reasonable questions in 2018. In 2026, every vendor answers yes to all of them, including vendors whose “AI” is a rules engine with a rebranded interface, and the RFP produces a tie that gets broken on price and demo polish.

The problem is categorical. A feature checklist evaluates what a system covers. An agentic TMS has to be evaluated on how it decides, because the product is the decisioning: whether it senses live conditions, chooses among options it was never explicitly configured for, executes within governed boundaries, and improves from outcomes. None of that is visible in a feature matrix, and all of it is testable if you ask the right questions and score the answers on evidence.

The 30 questions below are designed to be sent verbatim. Each section explains what it tests and what strong evidence looks like. The scoring rubric follows. Two honest notes before you use it: first, this scorecard is vendor-agnostic by design, and any credible agentic vendor, Locus included, should welcome being scored against it. Second, for the process wrapped around these questions (stakeholders, sequencing, proof-of-concept design), see our companion RFP framework for US enterprises; this piece is the instrument, that one is the process.

Section 1: Decision Architecture (Questions 1-5)

What it tests: whether decisions are computed or configured.

  1. For a given day’s dispatch plan, describe precisely what produces the plan: a rules engine executing configured logic, an optimization engine computing decisions, or autonomous agents. Which decisions does each component own?
  2. How many real-world constraints does your optimization model natively (vehicle, driver, regulatory, customer, commercial)? Provide the list, not the count.
  3. Show us a decision your system made that was not anticipated by any configured rule. How did the system arrive at it?
  4. When two objectives conflict (cost vs. service level, speed vs. margin), how does the system resolve the tradeoff, and who sets the weights?
  5. What happens to plan quality as our order volume doubles: does solve time degrade, does the constraint set thin out, or does the architecture hold?

Strong evidence: a named optimization and agent architecture with owned decision domains, a constraint catalog in the hundreds (Locus models 250+ in production), and a live demonstration on your data. Rebranded legacy tells: “our AI enhances your business rules,” constraint counts that collapse under enumeration, and tradeoff questions answered with configuration screens.

Section 2: Learning Loop (Questions 6-10)

What it tests: whether the system improves without reimplementation.

  1. Show us one decision your system makes differently today than six months ago, and the outcome data that caused the change.
  2. What specific signals feed back from execution into future decisioning (actual service times, travel times, carrier performance, failure patterns)?
  3. How often do models recalibrate, and is recalibration per customer or global across your install base?
  4. How does the system distinguish a genuine operational shift from noise before changing behavior?
  5. What guardrails prevent the learning loop from drifting into decisions that violate our commercial or compliance constraints?

Strong evidence: a concrete before/after decision with data lineage, named feedback signals, and guardrail architecture. Tells: “the system gets smarter over time” with no example, or learning that turns out to mean quarterly manual retuning by the vendor’s services team.

Also Read: Best TMS for Shippers in the Logistics Industry: TMS Software Comparison 2026

Section 3: Governance and Explainability (Questions 11-15)

What it tests: whether autonomy is deployable, auditable, and controllable.

  1. For any decision the system made last week, can an operator see why: the inputs, constraints, and tradeoffs behind it? Show us.
  2. What autonomy levels exist, and can we configure which decision types execute autonomously versus route to a human, per region and per decision class?
  3. Is there an execution sandbox where agent behavior can be tested against live data before production exposure?
  4. How are agent decisions evaluated on an ongoing basis, and what triggers automatic de-escalation to human control?
  5. Provide the audit trail for one historical decision end to end: signal, decision, execution, outcome.

Strong evidence: named governance mechanisms demonstrated live. Locus formalizes six (Explainability, Traceability, Evaluation, Autonomy Levels, Execution Sandbox, Human-in-the-Loop), and any vendor claiming agentic capability should produce an equivalent structure. Tells: governance answered as “full audit logs” (logging is not explainability) or autonomy with no de-escalation path.

Also Check: Locus, world’s first Agentic TMS in Action

Section 4: Integration and Execution Depth (Questions 16-20)

What it tests: whether decisions can actually execute across your stack.

  1. Which of the systems we run (OMS, WMS, ERP, telematics, driver apps) do you connect to today, in production, at a named customer? List them.
  2. How many carriers are pre-integrated and production-live on your network, and what does onboarding a carrier we use that you have never integrated actually take?
  3. When your system decides to switch an at-risk shipment to another carrier, does it execute the switch or generate a recommendation a human keys into another system?
  4. What is your data latency by source type, in minutes, in writing, and how does the system detect and behave under silent feeds?
  5. Walk us through one exception end to end in a live environment: signal to executed intervention, counting the human steps between.

Strong evidence: named production integrations, a pre-connected carrier network at scale (Locus’s ShipFlex connects 1,000+ carriers with 160+ pre-integrated), and an exception traced with zero unexplained human bridges. Tells: “open API” as the answer to every integration question, and recommendations dressed as execution.

Section 5: Disruption Response (Questions 21-25)

What it tests: whether the plan is a living object or a morning artifact.

  1. A driver with 40 assigned orders goes dark mid-shift. Demonstrate, on live or sandbox data, what your system does in the next ten minutes without human initiation.
  2. How long does re-optimizing an affected territory take at our scale, and what stays stable (unaffected routes) while it happens?
  3. During a demand spike at three times baseline volume, what degrades first: solve time, constraint fidelity, or alert quality? Show peak-season evidence.
  4. How does the system decide which customers to proactively notify during a disruption, and which notifications does it send autonomously?
  5. Provide a reference customer we can ask specifically about last Q4 peak, including uptime and re-planning performance under surge.

Strong evidence: a live mid-shift recovery demonstration, re-optimization in minutes with surgical scope, and peak references with contractual uptime (Locus operates at 99.99%). Tells: disruption handled by “alerts to your team,” or peak evidence that is a load-test report rather than a customer.

Also Read: Reducing CPG Supply Chain Costs Through Agentic Transportation Management System Architecture in 2026

Section 6: Proof and ROI Accountability (Questions 26-30)

What it tests: whether outcomes are evidenced, attributable, and contractually backed.

  1. Provide three production customers at our scale and operational profile, with the specific measured outcomes each achieved and the baseline methodology.
  2. For your headline ROI claims, what was measured, over what period, against what baseline, and what portion is attributable to the platform versus concurrent changes?
  3. What outcome metrics are you willing to put in the contract (execution rate, on-time performance, planning-effort reduction), and what happens if they are missed?
  4. What did your last three implementations at our scale actually take: elapsed time, customer effort, and time to first measured value?
  5. Which analyst research covers you in this category, and in what capacity: named evaluation, representative vendor, or paid placement?

Strong evidence: referenceable outcomes with methodology (the benchmark class here: plan execution lifted from 75% to 92% at a Fortune 50 enterprise with 4,500+ drivers; 80%+ dispatch effort reduction and year-one break-even at a retail enterprise), contractual outcome language, and independent analyst coverage. Tells: ROI claims with no baseline, references that all predate the “agentic” rebrand, and analyst mentions that dissolve under question 30’s framing.

The Scoring Rubric

Score every question 0 to 3 on evidence, not fluency:

ScoreMeaning
0No answer, deflection, or claim contradicted by the demonstration
1Roadmap or claim without evidence
2Demonstrated in vendor environment or single reference
3Demonstrated on your data, or production-proven with a checkable reference

Maximum: 90. Read the bands as follows. Below 45: rebranded legacy; the AI is a label on a rules engine, whatever the demo looked like. 45 to 67: transitional; genuine capability in some sections, usually decisioning, with gaps in governance, execution depth, or proof, so scope the contract to what scored 3. 68 and above: agentic-ready; hold the vendor to contractual outcomes per section 6. Weight sections equally on the first pass; if your operation is disruption-heavy, a second pass with sections 4 and 5 upweighted is defensible.

Two usage rules from the field. Score independently across your evaluation team before comparing, because vendor fluency anchors group scoring. And require every 3 to be evidence you could hand to procurement: a demo on your data, a reference call, or a document.

Also Read: Why TMS Migrations Fail: 7 Architecture Mistakes That Kill Digital Transformation in 2026

Build the Customized Version

The 30 questions above are the general-purpose instrument. If you want the version tailored to your operation, the interactive TMS RFI Builder takes a few context questions about your network (industry, fleet mix, carrier structure, disruption profile) and returns a customized question set ready to send to vendors: find it here and build your own RFI. For the evaluation process around the instrument, see our practical RFP framework for US enterprises.

Learn more about agentic TMS, explore Locus here.

Frequently Asked Questions (FAQs)

What should an RFP for an agentic TMS include?

Six sections of architectural questions rather than feature checklists: decision architecture, learning loop, governance and explainability, integration and execution depth, disruption response, and proof with ROI accountability. Score answers 0 to 3 on evidence strength, with the strongest claims demonstrated on your own data.

How do I tell real AI from rebranded legacy in a TMS evaluation?

Ask where decisions come from, demand a decision the system makes differently than six months ago with the outcome data behind it, trace one exception from signal to executed intervention, and require a mid-shift disruption demonstration. Rules engines with AI branding fail all four, fluently.

What is a good score on the agentic TMS scorecard?

Out of 90: below 45 indicates rebranded legacy, 45 to 67 indicates transitional capability worth scoping carefully, 68 and above indicates agentic readiness worth holding to contractual outcomes. The distribution across sections matters as much as the total; strong decisioning with weak governance is not deployable.

Should ROI guarantees be part of a TMS contract?

Outcome language should be, where the vendor’s evidence supports it: plan execution rate, on-time performance, planning-effort reduction, with defined baselines and remedies. A vendor unwilling to contract any outcome after claiming them in the RFP has answered the question.

How is this scorecard different from a standard TMS RFP template?

Standard templates evaluate feature coverage, which every 2026 vendor passes identically. This scorecard evaluates decisioning architecture and evidence, which is where agentic platforms and rebranded rules engines diverge. It is designed to be sent verbatim and scored on demonstration, not narrative.

MEET THE AUTHOR
Avatar photo
Ishan Bhattacharya
Lead - Content

Ishan, a knowledge navigator at heart, has more than a decade crafting content strategies for B2B tech, with a strong focus on logistics SaaS. He blends AI with human creativity to turn complex ideas into compelling narratives.

Related Tags:

Previous Post Next Post

General

5 Ways Logistics Automation Breaks Down at Scale in 2026 (and the Orchestration Layer That Fixes Each One)

Avatar photo

Aseem Sinha

Aug 4, 2026

Five operational failure modes of logistics automation at enterprise scale: the mid-shift plan collapse, carrier handoff disputes, SLA blind spots, orphaned exceptions, and margin-blind optimization. And the orchestration fix for each.

Read more

General

How to Run Driver Performance Scorecards in Last-Mile Logistics in 2026 (With Metrics That Actually Matter)

Avatar photo

Ishan Bhattacharya

Aug 5, 2026

How to run driver performance scorecards in last-mile logistics: the metrics that actually matter, weighting, coaching triggers, review cadence, and the mistakes that turn scorecards into surveillance.

Read more

The Agentic TMS RFP Scorecard: 30 Questions to Separate Real AI from Rebranded Legacy in 2026

  • Share iconShare
    • facebook iconFacebook
    • Twitter iconTwitter
    • Linkedin iconLinkedIn
    • Email iconEmail
  • Print iconPrint
  • Download iconDownload
  • Schedule a Demo
glossary sidebar image

Is your team spending more time on fixing logistics plan than running the operation?

  • Agentic transportation management from order intake to freight settlement
  • Route optimization built on 250+ real-world constraints
  • AI-driven dispatch with automatic execution handling
20% Cost Reduction
66% Faster Planning Cycles
Schedule a demo

Insights Worth Your Time

General

Locus 2026 US Consumer Survey: Generative AI isn’t Just Changing How Consumers Shop, it’s Breaking the Demand Patterns US Retail Was Built On

Avatar photo

Ishan Bhattacharya

May 29, 2026

General

Embedded vs Bolted-On AI: The Architecture Question European Logistics Buyers Are Asking

Avatar photo

Aseem Sinha

May 21, 2026

General

Hybrid Fleet Management: How Owned, 3PL, Gig, ICE, and EV Capacity Actually Operate at Most Enterprises

Avatar photo

Aseem Sinha

May 7, 2026

General

US Returns Hit $850 Billion in 2025: Why US Retailers Are Restructuring Reverse Logistics in 2026

Avatar photo

Ishan Bhattacharya

May 7, 2026

SUBSCRIBE TO OUR NEWSLETTER

Stay up to date with the latest marketing, sales, and service tips and news

Locus Logo
Subscribe to our newsletter
Platform
  • Transportation Management System
  • Last Mile Delivery Solution
  • Fulfillment Automation
  • Dispatch Planning
  • Delivery Orchestration
  • Track and Trace
  • Analytics and Insights
Industries
  • Retail
  • FMCG/CPG
  • 3PL & CEP
  • Big & Bulky
  • E-commerce
  • E-grocery
  • Industrial Services
  • Manufacturing
  • Home Services
Resources
  • Use Cases
  • Whitepapers
  • Case Studies
  • E-books
  • Blogs
  • Reports
  • Events & Webinars
  • Videos
  • API Reference Docs
  • Glossary
Company
  • About Us
  • Customers
  • Analyst Recognition
  • Careers
  • News & Press
  • Trust & Security
  • Contact Us
  • Hey AI, Learn About Us
  • LLM Text
ISO certificates image
youtube linkedin twitter-x instagram

© 2026 Mara Labs Inc. All rights reserved. Privacy and Terms

locus-logo

Cut last mile delivery costs by 20% with AI-Powered route optimization

1.5B+Deliveries optimized

99.5%SLA Adherences

30+countries

Trusted by 360+ enterprises worldwide

Get a Complimentary Tailored Route Simulation

locus-logo

Reduce dispatch planning time by 75% with Locus DispatchIQ

1.5B+Deliveries optimized

320M+Savings in logistics cost

30+countries served

Trusted by 360+ enterprises worldwide

Get a Complimentary Tailored Route Simulation

locus-logo

Locus offers Enterprise TMS for high-volume, complex operations

1.5B+Deliveries optimized

320M+Savings in logistics cost

30+countries served

Trusted by 360+ enterprises worldwide

Get a Complimentary Network Impact Assessment

locus-logo

Trusted by 360+ enterprises to slash costs and scale operations

1.5B+Deliveries optimized

320M+Savings in logistics cost

30+countries served

Trusted by 360+ enterprises worldwide

Get a Complimentary Enterprise Logistics Assessment