Delhivery is one of the largest tech-driven logistics companies, backed by FedEx and SoftBank. Know More →

    Back to blogs

    Another Vendor's Model Grades Penny, and You Never See the Score

    October 11, 2026•6 min read

    A vendor sits across from you and puts an accuracy figure on the slide, somewhere in the high nineties. Ask how it was produced and you will usually learn that the vendor ran its own system against its own test set and graded the output with its own model.

    That arrangement has a name in the evaluation literature, and the research on it is unflattering. It is also why two things are true about how Penny is assessed here: the grader comes from a different vendor, and the result never appears on your screen.

    The grader cannot also be the student

    Penny audits freight invoices before payment, so every output she produces is a claim with money attached. This line-haul charge exceeds the contracted rate by 412 dollars. This detention belongs to the carrier, not to you. Claims of that kind need checking at a volume no review team reaches by hand.

    Analysts check a sample. A model checks everything. The only real question is which model, and the answer stops being obvious once you read what happens when a system marks its own homework.

    Arjun Panickssery, Samuel Bowman and Shi Feng presented "LLM Evaluators Recognize and Favor Their Own Generations" at NeurIPS 2024. They define self-preference as an evaluator scoring its own outputs higher than others' while human annotators rate them equal in quality. The sharper result is causal: fine-tuning revealed a linear correlation between how well a model recognises its own writing and how strongly it prefers it. Recognition and flattery travel together.

    Translate that into a freight back office. If the model drafting Penny's disputes misreads a free-time clause, a judge from the same family, trained on overlapping data, tends to misread it identically and marks the error correct. You learn about it from a carrier's rebuttal weeks later, on an invoice you already paid.

    Changing vendors removes one bias, not all of them

    Cross-vendor grading is necessary and nowhere near sufficient. The MT-Bench and Chatbot Arena paper by Lianmin Zheng and colleagues, published at NeurIPS 2023, names three failure modes in LLM judges: position, verbosity and self-enhancement. The same paper reports that strong judges reach over 80 percent agreement with human preferences, matching the level humans reach with each other. Useful, and well short of a licence to run the grader unattended.

    Position bias is the most mechanical of the three and the best measured. "Judging the Judges", posted to arXiv in June 2024, ran 15 judge models across 22 tasks drawn from MTBench and DevBench, producing over 150,000 evaluation instances. Its findings: the bias is not random chance, it varies significantly across judges and tasks, prompt length influences it only weakly, and it is strongly affected by the quality gap between the solutions being compared. Close calls are where ordering decides the verdict, and an invoice audit is full of close calls.

    So the harness does four things. Presentation order is randomised per item. Authorship is withheld, so the judge never learns which system wrote the finding. Every dimension demands a citation to a document rather than an impression. And a stratified sample goes to human adjudicators each cycle, with each model-analyst disagreement read individually instead of averaged away.

    Five dimensions, because one number hides the failure

    A single composite tells you something went wrong without telling you what. Penny's work is scored on five axes, each answerable from the file in front of the grader.

    • Rate accuracy. Does the charge reconcile to the contracted rate for that lane, service level and effective date, against the fuel basis in force that week?
    • Evidence sufficiency. Is every factual assertion tied to a specific document: the rate table row, the signed delivery receipt, the gate-out timestamp?
    • Causality. For an accessorial, does the evidence establish who caused the delay rather than only that a delay occurred? This axis separates a recoverable charge from an argument, and it is the subject of what a detention charge has to prove.
    • Citation fidelity. Does the quoted contract clause exist, at the section cited, saying what the finding claims it says?
    • Dispute readiness. Would the drafted dispute survive the carrier's first reply without a second round of document gathering?

    They fail independently, which is the point of keeping them apart. A finding can be arithmetically correct and evidentially naked, or it can cite a real clause and draw the wrong causal conclusion from it. Averaging the five into one figure destroys the only information that tells an engineer what to repair.

    Why the result stays inside the building

    The formulation usually credited to Marilyn Strathern's 1997 paper on audit in the British university system runs: when a measure becomes a target, it ceases to be a good measure. Publish a judge score next to a product and it becomes a target inside a week.

    The gaming that follows needs no bad intent from anyone. Verbosity bias sits in the Zheng paper as a named defect: judges favour longer answers. A team optimising a published figure would gradually produce longer findings, more structure, more hedging, and the figure would climb. Carrier recovery would not move. An accounts payable clerk does not reward prose.

    The second reason concerns your reviewers rather than our engineers. Jingshu Li and co-authors, in "Understanding the Effects of Miscalibrated AI Confidence on User Trust, Reliance, and Decision Efficacy" (arXiv, revised September 2025), report that miscalibrated confidence impairs appropriate reliance and reduces the efficacy of AI-assisted decisions, and that users struggle to detect the miscalibration themselves. Their second experiment tried the obvious remedy of disclosing the calibration level. Users do then spot the problem, but trust in an uncalibrated system drops enough to produce heavy under-reliance, so decision quality still fails to improve.

    Taken together, those findings are uncomfortable for anyone selling a confidence badge. Hiding a bad number harms the reviewer, and disclosing it harms them differently. The escape is to keep the number out of the workflow.

    In the reviewer's handsA confidence badgeAn evidence trail
    What appears on the item0.93Charge code, contract clause, timestamped document
    How the reviewer verifies itTrusts the vendor's calibrationOpens the document
    How it failsQuietly, across every finding at onceVisibly, on the one item affected
    What happens when it becomes a targetRises without anything improvingCannot rise without a real document behind it

    A reviewer who sees 0.93 beside a finding reads less of the finding. That is the same erosion that hollows out an approval queue, described in full autonomy is a setting, not a product.

    What you get instead of a number

    The judge's output drives engineering, not marketing. It works as a regression gate: a prompt revision, a new charge-code handler, a new carrier invoice template, each runs the five dimensions before reaching a live account, and a drop on any single axis blocks the change. Miles and Chase sit behind the same gate with their own rubrics.

    What reaches you is different in kind. Each finding arrives with the charge code, the contracted rate it was measured against, the clause, the document and timestamp that settle the question, and a draft dispute you can read end to end in under a minute. The check available to you is the one worth having: open the evidence and decide.

    The number you should actually hold a vendor to is produced on your invoices rather than theirs. Recovery identified per hundred invoices. Dispute win rate with your carriers. False positives counted as findings your own team rejects on review. Those sit on your lanes, your contracts and your freight terms, and no internal grade predicts them.

    What to do next

    The next time a vendor shows you an evaluation figure, ask three questions. Whose model produced the grade. Whether the rubric has separate dimensions, and whether they are reported apart or collapsed into one. What the agreement rate is between the model grader and human adjudicators on a sample. A team that has done the work answers in a minute; a team that has not will call the methodology proprietary.

    Then, whatever the answers, run a shadow period on your own invoices and write the pass criteria down before it starts. If you want to see the five dimensions applied to a week of your freight, book a walkthrough and bring the invoices you argue about most.