Delhivery is one of the largest tech-driven logistics companies, backed by FedEx and SoftBank. Know More →

    Back to blogs

    Why No Language Model Counts Your Demurrage Days

    October 9, 2026•6 min read

    Image: Intermodal terminal loading by David Wilson from Oak Park, Illinois, USA, licensed under BY 2.0.

    A carrier bills six days of demurrage on a container. Your team puts the correct figure at four. That four then travels: into the short pay, into the mitigation request, and into the charge complaint if the carrier holds its position. Your company's name sits under it every time.

    So how the four was produced is not an implementation detail. If a language model wrote it by predicting plausible text rather than subtracting one timestamp from another, you have asserted a number to a carrier, and possibly to a federal regulator, that nobody can reproduce on demand.

    The count is a regulated field, not a working figure

    Under 46 CFR 541.6, a demurrage or detention invoice has to state the allowed free time in days, the start date and the end date of that free time, the container availability date on imports or the earliest return date on exports, and the specific dates for which the charge was raised. It has to name the governing tariff rule or service contract section and the specific rate under it. It has to carry a statement that the billing party's own performance did not cause or contribute to what it is billing you for.

    Section 541.5 gives those fields teeth. Omitting required minimum information, in the rule's own words, "eliminates any obligation of the billed party to pay the applicable charge."

    Dispute one of these invoices and you are filing a competing version of the same regulated fields. The Commission's Charge Complaint Procedures rule, published in the Federal Register on September 1, 2026, requires a charge complaint to carry the bill of lading numbers and the invoices, and puts the burden on the common carrier to prove its charges reasonable once the matter proceeds as a traditional complaint. Your filing is the thing that opens that burden, and a number you cannot derive twice is a poor way to open it.

    The stakes are large enough to justify the formality. In the final rule at 89 Fed. Reg. 14330, the FMC recorded that nine of the largest carriers in the US liner trades charged roughly $8.9 billion in demurrage and detention over a two-year span between 2020 and 2022, and collected about $6.9 billion of that.

    Transformers get arithmetic wrong quietly

    The case against generating the count is not that language models reason badly. They reason well, which is the whole reason to build agents on them. It is that arithmetic is the wrong job for the architecture, and the failures arrive without a warning label.

    Dziri and colleagues measured GPT-4 at 59 percent accuracy on three-digit by three-digit multiplication in "Faith and Fate: Limits of Transformers on Compositionality," published at NeurIPS 2023. The broader finding in that paper matters more than the headline figure: accuracy on compositional tasks falls toward zero as the number of required steps grows, and asking the model to show its work recovers less of the loss the harder the problem becomes.

    Counting days looks easier than multiplying, and on a clean pair of dates it is. Real counts are not clean. They skip the weekend days a terminal tariff excludes, honour a holiday the terminal observes while the carrier's clock keeps running, begin at an availability date that differs from the discharge date, and pause while a container sits under a customs exam. That is a chain of conditional steps, which is precisely the shape the research shows degrading.

    None of that decay is visible in the output. A wrong day count comes back in a well-formed sentence, in the same confident register as a right one, and nothing in the text marks which you received.

    You also cannot run it again

    Set correctness aside for a moment, because reproducibility fails first. An audit finding has to survive being asked for twice, since six weeks later a carrier's billing desk will want it rebuilt.

    Thinking Machines Lab published an experiment in September 2025 that settles the architecture question on its own. Sampling one prompt a thousand times at temperature zero, the setting meant to remove randomness entirely, produced 80 distinct completions from Qwen3-235B-A22B-Instruct-2507. All thousand runs agreed through the first 102 tokens. At token 103, 992 of them said one thing and 8 said another. The cause is not sampling: inference kernels are not invariant to batch size, so the unrelated requests sharing your batch move your result.

    Their remedy, batch-invariant kernels, drove all 1,000 completions to identical output. That is promising research rather than the API anyone is buying this quarter. A subtraction inside a deterministic engine, by contrast, returns the same integer on any hardware, under any load, in any year the claim is still open.

    Handing the model the document is not enough

    The standard reassurance is retrieval. Give the model the tariff rule and the gate records, keep it grounded in supplied text, and the risk goes away. Grounding helps considerably. It does not reach zero.

    Vectara's hallucination leaderboard scores how often a model introduces content its source document does not support. In the snapshot updated on September 22, 2026, measured over a set of more than 7,700 articles, the best entry on the board sat at 1.8 percent. That is first place, on summarisation, with the source text in front of the model.

    Call it one statement in fifty-five under the most favourable conditions anyone tests. Then decide which fifty-five invoice lines you would be comfortable being wrong about in writing, to a counterparty who keeps the correspondence.

    Where TransportOne draws the line

    Penny is an agent, and the agent layer does what agents are genuinely good at. Reading a tariff clause written in prose. Deciding which of four documents governs a given box on a given sailing date. Judging whether a terminal's congestion note accounts for a two-day gap. Drafting a mitigation request in language a billing desk will actually act on.

    No arithmetic happens in that layer. Free time windows, day counts, tier boundaries, rate lookups, currency conversion and line totals are computed by deterministic engines the agent calls. Penny does not write a figure she was not handed.

    Question on the invoiceAnswered by
    Which tariff rule or contract section governs this container on this dateAgent layer
    How many days fall between availability and gate out, net of excluded daysDeterministic engine
    Whether a customs exam or carrier hold explains the days billedAgent layer, read against the event record
    Which rate tier each billed day falls into, and the resulting totalDeterministic engine
    Whether every field required by 541.6 is present and internally consistentDeterministic engine
    How to phrase the dispute so it resolves rather than escalatesAgent layer

    The practical consequence is that every figure leaving in a dispute letter traces back to its inputs and to the function that consumed them. Run it again next year and it returns what it returned today. Chase supplies the discharge, gate and hold events the computation stands on, which is why attribution and audit stay joined on container work. Our post on demurrage as a causality problem takes up what to argue once the count itself is settled.

    Three questions to put to any vendor

    • Show me this container's day count as a computation rather than a sentence. What were the inputs, and which function consumed them?
    • Run the same audit twice on the same files. Do both passes return the same integer?
    • When the model and the engine disagree on a total, which one reaches the carrier?

    A vendor who cannot answer the first is generating your evidence instead of deriving it. One who fails the second cannot defend that evidence when it is challenged. The third question is the one that shows whether the separation is architectural or decorative, and it is worth asking in the room rather than in email.

    Book a walkthrough to see exactly where Penny's reasoning stops and her arithmetic begins, against your own demurrage invoices.