Synlogica Book a call

Contents

  1. Executive summary
  2. What the evidence actually says — and what it does not
  3. Why the threshold exists
  4. Seven error categories, all pointing the same way
  5. The category generic freight audit does not read
  6. Pre-payment checking is not faster post-audit
  7. The barrier that actually stops these projects
  8. Anatomy of a checked line
  9. Why this belongs under a regulated audit trail
  10. Takeaways

1. Executive summary

Freight invoices in regulated pharma are approved against a value threshold, not against the contract. Below the threshold, nobody opens the rate card. Above it, somebody opens whichever version of the rate card they can find.

This is not carelessness, and treating it as carelessness is the fastest way to lose the room. It is a rational response to a workload that cannot be met: several hundred invoices a month, each with six to twenty chargeable lines, checked against terms that live in a PDF from procurement, an amendment in somebody's mailbox and a fuel index that changes weekly. The threshold is the only mechanism a finance team has for deciding what to look at.

The consequence is structural rather than occasional. Errors below the threshold are not caught late — they are never seen. And because the errors are not random, what accumulates is not noise.

This paper sets out three things. What the published evidence actually supports, checked at primary source rather than repeated from vendor blogs. Why checking before the payment run is a different process from recovering afterwards, not a faster version of it. And what a line-level check requires in a regulated environment, where "we found money" is not on its own an acceptable answer to an auditor.

Disclosure, up frontSynlogica sells software that does this. Where a figure comes from a firm that sells freight audits, we say so, because the incentive matters. Where we have no measurement of our own, we say that too. REFRAKT has no completed deployments at the time of writing, so every claim here is either sourced to a third party or explicitly marked as a hypothesis we have not yet tested.

2. What the evidence actually says — and what it does not

The freight-audit literature has a citation problem. A small number of figures circulate widely, get quoted between vendor blogs, and acquire authority through repetition rather than through anyone checking them. Before building an argument, it is worth separating what has a traceable source from what does not.

2.1 The one freight-specific estimate worth quoting

Tompkins Ventures, an advisory firm, estimates that 5 to 10 percent of freight and parcel invoices carry some kind of error, with cases running as high as 40 percent.

Two caveats belong to that number and should travel with it. It is an experience-based estimate from a firm that audits freight for a living — the incentive runs one way. And it is an estimate, not a measured study with a published method and sample.

We quote it because it is the most defensible freight-specific figure available, not because it is strong evidence. If your own invoices produce a different rate, your rate is the better number.

2.2 A real accounts-payable benchmark, checked at source

The adjacent figure is more solid, and it is worth going to the original rather than to the summaries.

Ardent Partners' State of ePayables 2025 reports an invoice exception rate of 18.4 percent across all accounts-payable invoices, alongside a cost of $9.84 to process a single invoice, 8.2 days of processing time, and 35.4 percent of invoices processed straight through.

Ardent Partners, State of ePayables 2025 — Table 1Figure
Invoice exception rate18.4%
Cost to process a single invoice$9.84
Processing time8.2 days
Straight-through processing35.4%

Two observations. First, this is all AP, not freight specifically — freight is a harder category than most, because the pricing has more moving parts than a typical supplier invoice. Second, the figure is frequently quoted in secondary sources as "about 14 percent," which appears to be an older edition carried forward. We read the 2025 report; the number is 18.4 percent. If you are going to cite it, cite it correctly.

The straight-through figure is the one worth sitting with. If roughly a third of invoices clear without human touch, the remaining two-thirds consume someone's day — and that is before anyone asks whether the terms were right.

2.3 The figures we will not repeat

The following circulate constantly and have no verifiable primary source. We list them so you can recognise them in a vendor deck:

Each traces back to another blog post citing another blog post. None resolves to a study you can read. A tool that is meant to make your invoice checking defensible should not be sold to you with numbers that are not.

Recovery rates are a softer case. The commonly stated range — 2 to 8 percent of audited freight spend — appears without attribution even in otherwise careful write-ups. Treat it as a market observation, not a benchmark.

3. Why the threshold exists

Understanding why threshold approval is rational is the difference between a tool that finance adopts and a tool that finance resents.

Consider what checking one freight invoice against its contract actually requires:

  1. Locate the agreement in force on the shipment date, not today's version.
  2. Find the applicable rate for that lane, service level and weight break.
  3. Establish which fuel index applied that week, and the formula that converts it into a surcharge percentage.
  4. Determine which accessorials are included in the base rate, which carry free-time allowances, and how much free time had been consumed.
  5. Confirm the shipment is not already invoiced under a different document number.
  6. Compare line by line and quantify each variance.

At perhaps twenty minutes per invoice when everything is findable — and it rarely is — several hundred invoices a month is not a task. It is a headcount request that will not be approved, competing against a payment deadline that will not move.

So the team sets a threshold. Everything above it gets scrutiny; everything below clears. This is a defensible allocation of scarce attention, and it is what any competent controller would do.

The structural flawErrors are not distributed by invoice value. A phantom re-icing charge on a small shipment looks identical to a phantom re-icing charge on a large one. The threshold filters by size, and error frequency does not correlate with size — so the filter selects almost at random with respect to the thing it is meant to catch. What it reliably does catch is the large, obvious, one-off mistake. What it reliably misses is the small, repeated, systematic one. Which is the expensive one, because it recurs.

4. Seven error categories, all pointing the same way

Freight billing errors fall into recognisable types:

CategoryWhat it looks like
Duplicate invoicingThe same movement billed twice under different document numbers
Wrong rate or classA lane billed at a rate from a different service level or weight break
Phantom accessorialsCharges for services included in the base rate, or not performed
Fuel surcharge miscalculationWrong index, wrong week, or the right index applied with the wrong formula
Reweigh and dimensional upchargesWeight or dimension adjustments applied without evidence
Demurrage and detentionTime billed past a free-time allowance that was not tracked
Currency and conversionRate agreed in one currency, invoiced at a rate applied on the wrong date

The categories matter less than a property they share. The direction of error is not symmetrical. In a genuinely random error process, roughly half the mistakes would favour the shipper — an accessorial forgotten, a surcharge under-applied, a weight break rounded the wrong way. In practice, the errors that survive to the invoice overwhelmingly favour the carrier.

There is no need to reach for bad faith to explain this. Errors that favour the shipper get corrected by the party that loses money — the carrier's own billing controls catch them. Errors that favour the carrier are only caught by the shipper, and the shipper is operating a threshold. Two asymmetric control systems produce an asymmetric result. Nobody has to be dishonest.

This is also why the systematic errors are the ones worth catching. A carrier billing structure that misreads a free-time allowance does not misread it once.

5. The category generic freight audit does not read

Freight audit as a service is a mature market. What it is not built for is cold chain.

Pharmaceutical shipments carry charges that exist nowhere else in general freight: re-icing and dry-ice replenishment, active container conditioning and rental, temperature-monitored service premiums, deviation handling when a lane is re-routed to protect a load, and qualification surcharges for GDP-compliant equipment.

To a generic freight-audit engine, these arrive as accessorial codes it has no rule for. The safe behaviour — passing anything unrecognised — means the pharma-specific charge categories are precisely the ones that go unchecked. Which is inconvenient, because they are also the ones with the least price transparency and the fewest comparable benchmarks.

There is a second-order effect specific to regulated shippers, and it is the one we find most interesting.

A temperature-controlled shipment generates two independent records: a quality record showing whether the temperature was held, and a financial record showing what was charged for holding it. In almost every organisation these live in different systems, owned by different functions, reconciled never.

The question that is currently unaskableDid we pay a temperature-control premium on a shipment that suffered an excursion? It is not a rhetorical question. It has a number attached, it recurs, and answering it requires only that the quality assessment and the invoice check sit on the same platform against the same shipment.

Synlogica Terminus runs both — the quality module (M4) and REFRAKT — on one engine and one audit trail, which is the only reason the question resolves rather than being escalated to two departments who each hold half the answer.

We have not yet measured how often the two records disagree. We expect it to be worth measuring.

6. Pre-payment checking is not faster post-audit

The freight-audit market is overwhelmingly a post-audit market: invoices are paid, then reviewed, then claims are raised against the carrier for what was overpaid.

Checking before the payment run is not the same process moved earlier. The economics differ at every step.

Post-audit recoveryPre-payment check
PositionYou are a claimantYou are the payer
LeverageThe carrier holds the moneyYou hold the money
Process costClaim, dispute, chase, reconcile the creditLine held, resolved, released
Typical outcomeA credit note against future invoicesA payment that was correct in the first place
Relationship effectRecurring adversarial contactA pricing conversation before money moves
Small variancesWritten off — below the cost of pursuitChecked at the same cost as large ones

That last row carries most of the value. In post-audit, a €90 variance is not worth pursuing — the process costs more than the recovery, so it is dropped. Since the variances that recur are typically small, post-audit systematically abandons exactly the errors that repeat. Pre-payment checking has no per-variance pursuit cost, so the threshold problem does not reappear in a new form.

There is a real objection here, and it should be met directly rather than argued away: a check placed before the payment run can delay payments. Late payment damages carrier relationships and, in some jurisdictions, incurs interest. Payables teams are right to raise this, and any vendor who waves it off has not run a payment cycle.

The answer is scope. An exception should hold the disputed line, not the invoice. The remaining lines keep their original payment date. A three-line disagreement on a forty-line invoice should not stop thirty-seven correct lines from paying on time — and a check that stops the whole document will be switched off inside a quarter, correctly.

7. The barrier that actually stops these projects

Ask a controller why they have not implemented invoice-level checking and the answer is rarely "we do not believe there are errors." It is:

"Nobody is going to transcribe our contracts."

This is the real constraint, and it is usually decisive. The terms exist as prose: a master agreement, a rate annexe, three amendments, an email confirming a lane extension. Turning them into structured, machine-comparable terms is unglamorous work requiring both contract literacy and freight literacy. No finance team has that capacity spare, and outsourcing it internally means asking procurement for people they do not have either.

Any honest assessment of this software category has to acknowledge that the digitisation cost is the project. The checking engine is comparatively easy. Deciding what "included in the base rate" means for a specific carrier's accessorial schedule is not.

Two consequences follow for anyone evaluating a tool:

Ask who does the extraction. If the answer is "you upload structured terms," the vendor has moved the hard part onto you and the project will stall in month two. Extraction belongs on the vendor side, with the customer confirming rather than transcribing.

Ask how terms are versioned. A shipment must be checked against the agreement in force on the day it moved. A tool holding only the current version will produce confident, wrong exceptions on every shipment predating the last amendment — and nothing destroys trust in an exception queue faster than exceptions that turn out to be the tool's fault.

8. Anatomy of a checked line

What a finance team should expect to receive is not a report. It is a queue of exceptions, each of which is a complete argument.

For a single flagged line, that means: the line as invoiced; the term it was checked against, with the contract version and its effective dates; the computed expected value with its inputs; the variance in currency; the shipment reference tying it to a physical movement; and an owner, a status and a deadline.

That last group is what separates a tool from a spreadsheet. A dispute that lives in one person's mailbox cannot be counted, aged or handed over. As a record with an owner and a deadline, the open-disputes position becomes a number the controller can state — which, in most organisations we have spoken to, is a number that does not currently exist anywhere.

A worked example, illustrative rather than measured:

LineContractInvoicedVariance
Line haul FRA → MXP2 450.002 450.00
Fuel surcharge14.2%18.5%+105.35
Temperature-controlled service380.00380.00
Waiting time1 h free3 h billed+90.00
Re-icingincluded120.00+120.00

Three exceptions on a five-line invoice, totalling €315.35 against an invoice of €3 275 — 9.6 percent above contract. Two of the three are below any realistic review threshold. All three are individually small enough that, discovered a quarter later, none would be worth a claim.

Note what the fuel surcharge line requires: not a rate comparison but a recomputation from the index and formula in the agreement, for the shipment date. That is the sort of check that is trivial for software and effectively impossible at volume for a person.

9. Why this belongs under a regulated audit trail

In most industries, invoice checking is a cost-control function and the output is savings. In regulated pharma the same activity carries an evidentiary obligation, and this changes what "done" means.

An internal auditor asking about payment controls is not asking whether you recovered money. They are asking on what basis the payment was released, and whether you can demonstrate it. "Someone in payables looked at it" is not a control — it is the absence of one described in the past tense.

Three properties follow, and they are the same three that apply to any record in a regulated environment:

Attributable and contemporaneous. Every comparison records what was compared, against which contract version, by whom and when — written at the time, not reconstructed later.

Reproducible. Re-running the same shipment against the same contract version must produce the same result. This requires versioning both the terms and the checking rules, and pinning both into the record.

Advisory, with a human decision. The engine raises the exception; a person decides. That decision is recorded with a reason. Nothing is disputed, held or released automatically — partly because it is better governance, and partly because a finance team will not accept a system that transacts on its behalf, and is right not to.

None of this is exotic. It is the same posture the quality side of a pharma organisation has applied to computerised systems for years, under EU GMP Annex 11 and 21 CFR Part 11. The observation we would offer is that finance functions in regulated companies are increasingly asked evidentiary questions that quality functions have been answering for two decades — and that the tooling has not caught up.

Running the financial check under the same governance as the quality assessment is not gold-plating. It means the answer to "show me the basis for this decision" has the same shape whichever decision is asked about.

10. Takeaways

  1. The threshold is rational and structurally wrong. It filters by invoice value; errors do not scale with invoice value. The filter and the target are uncorrelated.
  2. Cite the evidence honestly or not at all. One defensible freight-specific estimate exists (5–10 percent, Tompkins Ventures, experience-based, from a firm that sells audits). One solid adjacent benchmark exists (18.4 percent AP exception rate, Ardent Partners 2025, checked at source). The rest of what circulates does not survive a look at the footnote.
  3. Before and after are different processes. Pre-payment checking changes your position from claimant to payer and removes the per-variance pursuit cost that causes post-audit to abandon exactly the errors that recur.
  4. Hold the line, not the invoice. A check that stops correct lines from paying on time will be switched off, and should be.
  5. Contract digitisation is the project. If the vendor does not do the extraction, the project does not start. If terms are not versioned by effective date, the exceptions will be wrong.
  6. In regulated pharma the output is evidence, not just savings. The audit trail is the deliverable; the recovered money is a by-product.

What we do not yet know

REFRAKT has no completed deployments as of August 2026. We cannot tell you what error rate your invoices carry, how long extraction takes for a contract estate of your shape, or what proportion of raised exceptions survive carrier challenge. Anyone quoting you those numbers for your business, from any vendor, is quoting someone else's.

What we would rather do is compute the first one with you, on your invoices, before you buy anything.

Sources

Tompkins Ventures, Why Freight and Parcel Shipping Require Post-Auditing (July 2026), via GingerControl, Freight Invoice Audit: Where the 5-10% Billing Errors Hide — experience-based estimate, not a measured study.

Ardent Partners, The State of ePayables 2025 — 2025 AP Benchmarks, Table 1: invoice exception rate 18.4%, cost per invoice $9.84, processing time 8.2 days, straight-through processing 35.4%. Figures read from the report, not from secondary summaries.

Disclosure. Synlogica develops Synlogica Terminus REFRAKT, software that performs the pre-payment checking described here. Illustrative figures are labelled as such. No claim in this paper rests on Synlogica measurement, because at the date of publication we have none.

Want the first number computed on your own invoices?

Send one month of freight invoices and the rate card they should have been checked against (anonymised lanes are fine). We return a line-level check with the variance quantified and the contract version each line was compared to — before you buy anything.

Book a meeting →