Image: Untitled, licensed under CC0 1.0.
Someone has asked you to approve an autonomous agent for a freight workflow, and the approval form has one box: yes or no. The demo ran end to end without a human touching anything, which is impressive and also the wrong thing to evaluate.
The useful question is not whether an agent can run unattended. It is which of your workflows it should run unattended, inside what bounds, and what happens the first time one of those bounds is hit. Autonomy is a dial you set per workflow, not a property of the software you bought.
The driving taxonomy, read properly
Everyone in freight knows the automation levels; the trucking press has spent a decade explaining them. SAE J3016, currently in its April 2021 revision, defines six, running from no automation up to full. Most summaries stop at that list, and in stopping there they drop the part that matters.
The standard's own abstract says the levels apply to the driving automation feature or features engaged in a given instance of on-road operation. The level belongs to whatever is switched on at that moment, not to the truck. One vehicle can show two levels in the same hour depending on which feature is active. SAE also asks that its levels carry the prefix, as "SAE Level 2", so treat the L0 to L4 shorthand below as borrowed structure rather than the standard itself.
The highest-autonomy freight deployment in the United States makes the same point physically. Aurora began commercial driverless runs between Dallas and Houston on May 1, 2025, with Hirschbach Motor Lines and Uber Freight as launch customers. TechCrunch reported that one truck had covered 1,200 miles without a driver at that stage, while more than 30 supervised autonomous trucks hauled over 100 commercial loads a week. The company did not become autonomous. One lane did.
Five tiers in a freight back office
Out of the cab, the tiers describe who acts and who checks. The examples are where a procurement conversation actually gets decided.
| Tier | Who acts | Example |
|---|---|---|
| L0 Manual | Your analyst, start to finish | Invoice and contract shown side by side. A person reads both. |
| L1 Assisted | Agent surfaces, person decides | A week of invoices ranked by variance against contract. The analyst picks what to work. |
| L2 Drafted | Agent prepares, person approves | A dispute written with charge code, contract clause and evidence attached. A human sends it. |
| L3 Supervised | Agent acts, person can reverse | A missed delivery appointment rebooked, the planner holding an hour to undo it. |
| L4 Bounded | Agent acts, person reads reporting | Line-haul charges matching contract rate exactly, under a stated ceiling, on carriers with no open dispute. |
There is no L5 on that list, and the omission is deliberate. A condition-free tier would mean a workflow with no edge: no dollar ceiling, no carrier exclusions, no lane where the rules differ. Freight operations do not contain one, and naming the edge is most of the work.
The tier belongs to the workflow
Ask what tier your AI runs at and the question has no answer, because the subject is wrong. Penny matching a line-haul charge to a contracted rate is arithmetic against documents you already hold, with the rate table and the signed delivery receipt in evidence. That can sit at L4 under a ceiling in the first month. Penny deciding whether four days of detention at a consignee's dock are the carrier's fault or yours is a causality judgment worth a few thousand dollars, and it belongs at L2 for a long time.
Same agent, same week, two tiers. The workflows differ in how reversible a mistake is and how cleanly evidence settles the question.
The pattern repeats. Chase pushing a revised ETA into a customer portal is high-frequency and low-consequence, which is L4 territory from the start. Chase choosing to reroute a load around a closed terminal moves money and reorganises somebody's afternoon, which is not. Miles assembling a carrier shortlist produces a draft. Miles saying a number out loud to a dispatcher produces a commitment, a separate problem covered in what stops a voice agent committing you.
Hard rules sit above the tier
A tier is a default. A hard rule is a veto, evaluated before the tier is consulted at all. That ordering decides whether your guardrails survive a genuine exception.
- Hard rules. Never approve payment without a signed delivery receipt. Never release a dispute against a carrier under legal review. Never commit above the authority the lane owner holds. These apply at every tier, L4 included.
- Tier default. What the agent may do on this workflow when nothing above has fired.
- Agent judgment. How it decides inside that permission.
An L4 workflow that trips a hard rule does not get clever about it. It stops and routes to a named person, and the restriction does not soften because the tier is high. That is what makes a tier safe to raise.
Regulators have begun writing the same shape down. Article 14 of the EU AI Act requires that high-risk systems let the people overseeing them disregard, override or reverse an output, and interrupt the system through a stop button or equivalent so it halts in a safe state. Article 14(3) splits delivery between measures the provider builds in and measures the provider specifies for the deployer to operate. Whether or not your use falls in scope, it is a sound checklist for any tier you switch on.
Watch L3, where approval stops being a decision
Article 14(4)(b) names a failure worth taking personally: the tendency to rely or over-rely on an output, especially where a system recommends decisions to a person. In a freight back office it shows up as a reviewer who clears an entire L3 queue without opening a single item in it.
If your team approves 99 percent of what the agent produces, you are running L4 already. You are paying latency for a signature that carries no scrutiny, and your audit trail records a review that did not happen. Two honest responses exist: promote the workflow and spend the recovered attention on exceptions, or work out why the check went hollow and repair it. Leaving it as it stands buys nothing.
Moving a workflow up, and back down
Tiers should ratchet on evidence, with the evidence agreed in advance.
- Run the workflow in shadow for a fixed period. The agent decides nothing; you compare what it would have done against what your team did.
- Write the promotion threshold before that run starts, as a number: agreement rate across a stated sample, every disagreement reviewed individually.
- Step to L2, then to L3 with a reversal window long enough to catch a real error and short enough not to be theatre.
- Step to L4 inside a bound narrower than feels necessary. Widening later is cheap. An unwind is not.
Demotion has to be as easy as promotion, and it needs an owner who can act without convening anyone. A dial that only turns one way is not a control, it is a ceremony.
None of this is abstract governance talk. Gartner forecast in June 2025 that more than 40 percent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. Programmes die of missing controls far more often than missing models.
What to do next
Take your three highest-volume freight workflows and write each one out in this order. Hard rules first, as plain sentences built around the word never. Then the starting tier, which for most teams is L2. Then the evidence that would justify a move, and who can move it back. That page is worth more than any demo, and it takes an afternoon.
If you want to see the tiers running against real invoices and real exceptions, book a walkthrough and bring one of those three workflows with you.
