Introducing the Long-Horizon AuditBench for Supply Chain AI

Loop Team
Industry Experts

Sep 1, 2026

.

5 minutes to read

Today we're launching Loop's Supply Chain AI Benchmark, a research program to measure and track how general-purpose AI models perform on real supply chain and logistics work compared to Loop's vertical AI harness, DUX™ 2.0. Our first benchmark is the Long-Horizon Freight AuditBench, a complex task that involves auditing a real-world freight invoice from unlabeled documents.

The initial benchmark results show that the vertical AI harness greatly outperforms general-purpose frontier models today. Loop correctly audits 94.8% of our basic freight invoices (the Production Dataset), and it correctly audits 65% of our invoices that include corner cases meant to challenge both AI systems and humans (the Challenge Dataset). The best general-purpose model, Claude Opus 5, correctly audits 16.1% and 10%. That's a gap of 78.7 percentage points on the Production Dataset and 55 points on the Challenge Dataset—and Loop gets there at up to 96% lower cost per audit.

This is the first benchmark in a longer series. Later rounds will extend to specific AI tasks required to understand and provide value from critical supply chain data.

Introducing Long-Horizon Freight AuditBench

AuditBench is built around LTL (less-than-truckload) freight audits. A single audit requires multiple dependent stages and hundreds of decisions, from ingesting data from raw carrier paperwork to calculating a final expected charge. Any missed step along the way, and everything after it will be wrong.

It also has something most agentic tasks don't: a verifiable answer. Either the invoice matches the contract and the supporting documents, or it doesn't. That lets us score a system objectively instead of against a rubric.

DUX™ 2.0 is Loop's vertical AI harness. It uses frontier models and dynamically selects the right one for each step of the work. Around the model sits numerous agents with deep supply chain context.

To find out what that context is worth, we ran DUX™ 2.0 and the frontier models against the same 200 real production freight invoices. The input is what a carrier typically sends a shipper. One packet, containing a core PDF that holds the invoice, the bill of lading, the delivery receipt, EDI files, and other documents. Alongside it sits the shipper's signed pricing agreement and full rate package. The files can span different file formats, including Excel files, PDFs, and images.

Each system audited the set twice: once on the Production Dataset of standard invoices, once on the Challenge Dataset of corner cases and hard audits.

This is the first benchmark we know of that compares vertical AI to general-purpose AI on long-horizon supply chain work using real supply chain documents. It gives supply chain and AI leaders a way to price what domain context is actually worth, and it gives us a baseline to track as models improve. Later rounds will extend to other supply chain tasks where the answer can be verified.

Reporting on two key numbers: % correct rate and cost per audit

Quality, measured as % Correct. The audit's job is to identify invoices ready to be paid and flag the ones where the bill has an incorrect charge or other discrepancy. The % correct rate is how many audit tasks the system prices correctly and then correctly classifies as right or wrong. For example, identifying an incorrect invoice as wrong is considered a succesful audit. Because ground truth is verifiable by a human, we can measure that objectively rather than against a rubric.

Cost-efficiency, measured as USD cost per audit. AI costs vary widely and AI spend is under more scrutiny, so cost per task is the second metric that matters: more intelligence per dollar. We measure total inference cost across every step taken and token consumed. A model that struggles with the task burns tokens, so a low per-token price can still produce a high cost per audit. That's why we report cost per task, not cost per token.

Testing methodology: Two unique datasets

Every model tested received the same materials: the customer's signed pricing agreement and full rate package; the carrier's documents exactly as they arrived, including PDFs, EDI feeds, emailed invoices; and web search to look up rates or published rules. 

The models were tested against two separate datasets:

  • Production Dataset: This set pulls a sample of 200 invoices from a month of real client audits, with 100 invoices having known errors and 100 without errors. Based on LTL audits finding errors in 24% of audits, afterwards we reweighted to match the true mix of correct and incorrect carrier bills. This dataset contains invoices that can be calculated with available documents and contain standard terms and fees,, representing simpler freight audits. 
  • Challenge Dataset: This set is a hand-curated sample of 100 real-world, challenging invoices, built to cover a wide range of charge types and failure conditions that are common in audits. It also contains an equal split of invoices with errors and those without. This set is weighted toward the audits where the information is less clear, data is missing, and other edge cases that challenge even the best vertical AI systems. 

Both datasets are composed of real-world LTL (less-than-truckload) freight invoices produced between July 15, 2026, and August 14, 2026. Each invoice was audited and reviewed so the ground truth (correct answer) of the audit was established.  

The results: Quality

The following chart illustrates the correct rate percentage of both Loop and the SOTA models against both datasets. The results show that general-purpose models, while extremely capable, still lack the context (and the ability to easily find the correct context through web searches) to perform a complex supply chain audit. The advantages of a vertical harness, like what Loop DUX™ 2.0 adds around a core LLM, resulted in a 78.7 percentage point advantage over the best performing model (Opus 5) for the production dataset. 

Production Dataset details

For the Product Dataset tests, we anticipated high levels of performance from Loop’s agentic system given the vertical AI harness includes context-aware agents built to understand and manipulate complex supply chain data. On the flipside, general-purpose models that lack context or supporting agents are expected to struggle even with this more straightforward dataset. 

Based on the above results, we can see that some of the most advanced AI systems today, like Opus 5, can have some level of success on the basic audit required for the Production Dataset. However, as we examine the next chart and table, we reveal how Opus 5 and the other models performed in auditing the 100 samples of invoices with known errors and the 100 without errors. We see that the more complex tests against the “Invoices with errors”—those with billing discrepancies—caused significant challenges. On audits where the carrier actually billed wrong, the frontier model landed on the correct amount only 1 time in 100.

Opus 5 scores 21% on clean invoices largely by agreeing with the carrier, which is available without pricing anything at all. In other words, it requires significantly less reasoning and context to calculate the charges from the shipment job facts included in the documents and the rates in the contracts, and then match that to the invoice total. 

Note on comparing AI to human quality

It’s also important to note that human-based audit systems, such as those you encounter with legacy freight audit and payment (FAP) service providers, never have a 100% correct audit rate either. According to Gartner, the typical error rate in manual, repetitive work—such as data entry and review—ranges from 1% to 4% per transaction, resulting in a correct audit rate of 96% to 99%. That percentage of correct audits grows much smaller when people are under time pressure or managing complex workflows (like a detailed audit with multiple documents). 

As AI systems approach the 95% accuracy mark, they reach a point where they are matching the quality of human work, with the ability to deliver higher performance at a more consistent and predictable rate. 

Challenge Dataset

In contrast to the straightforward audits above, the Challenge Dataset was curated to contain real world edge cases that challenge even human experts with unpressured time to investigate. The goal here is to measure agentic system performance against tasks that are difficult for even the most experienced human experts. All of the answers are ground truth, verified by experts, so we can view and track the performance of AI for these tasks over time.

The results: Cost-efficiency

Cost-efficiency is the other key metric that is becoming increasingly important as AI costs rise and teams are under great pressure to control (or at least monitor) their token spend. 

Along with the base AI cost per audit as shown in the first table below, we analyze and report on the data in two other key ways in this section. The second table in this section (“AI cost per correct audit task”) divides total AI spend by the number of audits answered correctly, so a system that is cheap per audit but usually wrong looks expensive per right answer. The third table (“AI cost per billing error caught”) divides total AI spend by the number of billing errors caught, which is the outcome a shipper is paying for.

The results below have been normalized to maintain relative results but hide confidential information on actual costs. We’ve reset the Loop cost for each dataset. The cost shows the total token cost to complete all tasks and activities for each invoice audit.

Costs are shown as mean cost per audit relative to Loop (Loop = 1.00x). Production figures are reweighted to the true billing-error mix.

Cost per task across models varies significantly. Opus 5 cost 47x more than GLM 5.3 Flash with less than 5 percentage points improvement in quality. One reason for these high costs can be attributed to the limited domain knowledge, forcing the models to perform numerous web searches that amplify costs but do little to help the actual performance.

Cost per correct audit relative to Loop

GLM 5.3 Flash makes one call per audit on about 50,000 input tokens, returns an empty response on 13% of audits in the Production Dataset (counted as wrong), and only catches 1 of 100 billing errors. In other words, some of the models are struggling to perform the necessary work and thus stopping and not giving results, leading to a much lower cost paired with a very low correct audit rate.

Cost per billing error caught relative to Loop (Production Dataset)

This is the cost of the outcome a shipper pays for one incorrect invoice found. It differs from cost per correct audit in what counts as a hit. A model that agrees with the carrier on a clean invoice scores a correct audit but has caught nothing, so a cheap one-shot model can look efficient per correct audit while finding almost no billing errors. Per billing error caught, every bare model is more expensive than Loop.

The growing importance of context for the next wave of AI

Frontier models are extraordinarily capable, and they keep getting better. Loop’s vertical AI harness benefits from more capable and efficient frontier models. By using intelligent routing, Loops is designed to dynamically deploy its vertical harness around the model that provides the best performance and value for a given job. As part of our ongoing research, which we didn’t explore here today, we measure how the different models perform within the Loop harness so we can choose the optimal model at any given time. As new models, like GLM 5.3 are released, we can quickly take advantage of the latest models, both open-weight and closed-weight models.  

The results clearly show us that general intelligence doesn't carry domain context with it, and on a long, complex task like freight audit, context is critical. Data ingestion has become relatively easy for SOTA models. Collecting ground truth for training data and developing true supply chain context is harder. 

Around the frontier model within Loop’s Logistics Data Platform sits more than a dozen agents that each hold one piece of the domain: DUX™ 2.0 starts by turning a stack of mixed carrier paperwork into clean, normalized data and then specialized agents perform key tasks to improve the data, align to unique business requirements, and stress test if the right fees are applied.

An AI model wrapped in that much domain context is what true supply chain AI actually looks like. Our testing showed that the power lies in the AI harness, proven by the fact that non-frontier models with a specialized harness outperform the best frontier models.  

What we are building next

Our release of the Long-Horizon Freight AuditBench is the start of Loop’s Supply Chain AI Benchmark and shows the importance of context to performing freight audits. We are exploring future tests that expand to other frontier models, test audit across other transportation modes (ocean, air, and drayage), and explore decomposed audit tasks that are representative of foundational tasks needed for true supply chain intelligence and action.

Additional task families, like dispute handling, GL coding, and accrual reporting, are all long-horizon tasks with verifiable outcomes that can be modeled after this Long-Horizon Freight AuditBench.

The results of this benchmark demonstrate the value of vertical AI. For complex, data-rich industries, like supply chain and logistics, the right vertical AI dramatically outperforms general-purpose AI both in its ability to complete key tasks but to also optimize AI spend. 

Table of contentS
Share article: