All posts

Cost, , 5 min read

Why per-token AI pricing breaks at shipment-event volume

A pilot that costs a few hundred euros a month can become a line item nobody can forecast once every shipment event, email and document runs through it.

Per-token pricing is a good deal when usage is small and unpredictable. You pay only for what you use, and there is no infrastructure to run. For many early AI experiments in logistics, that is exactly the right choice.

The trouble starts when a pilot succeeds.

Logistics volume is relentless

A mid-sized forwarder can generate hundreds of thousands of events a month: milestone updates, EDI messages, carrier emails, customer queries, invoices, delivery notes and customs documents. Add telematics pings and dock camera events, and the numbers climb again. Every one of them is a candidate for AI: classify it, summarise it, check it, draft a reply.

Under per-token pricing, each of those becomes a small charge. Individually they are trivial. Together, and growing with every new customer and every peak season, they become a cost that finance cannot forecast and operations cannot control.

The more volume you win, the less you can afford to automate. That is the wrong incentive.

Where the break-even sits

Every AI workload has a break-even point. Below it, paying per token is cheaper than running your own capacity. Above it, owned or reserved capacity wins, because the cost of a GPU running in your cloud does not rise with each extra message it processes.

In logistics, a large share of work crosses that line early. It is high in volume, repetitive in shape and steady across the day, which is exactly the profile where private capacity pays for itself.

Four levers that keep private AI lean

  • Small models first. Most requests never need the largest model. Status triage and document checks run well on small, fast models.
  • Capacity that follows your day. Computing scales with daily cut-offs, overnight batch runs and peak season, so you do not pay for idle hardware.
  • Shared hardware. Several models and tasks share the same processors instead of each needing their own.
  • Managed AI only where it pays. Public-only work that genuinely benefits from a frontier model still goes to one, under a monthly cap.

Cost is not the only reason

At very low volumes, pay-per-use can remain cheaper, and it is worth checking honestly where your organisation sits. Our break-even calculator is a good place to start. But for logistics companies, cost is often the second argument. Customer contracts, pricing data and customs information frequently justify keeping AI private before the numbers do.

The practical answer is hybrid: private capacity for the steady, confidential bulk of the work, managed models for the public research that benefits from them, and a policy check deciding each request.

See It on Your Own Data

A supervised pilot shows how hybrid AI performs on your shipments, documents and volumes, measured against how your team works today.

Discuss a Pilot