In one paragraph
A disruption — a port strike, a river running low, a corridor closing, a tariff filed — reaches a freight desk through a dozen different institutions, in a dozen different formats, and usually later than it reached somebody local. SQRlane runs on the bookings in a forwarder's TMS: it reads sixty free sources across six families — news, river gauges, weather and sea state, natural hazards, government filings and reference rates — decides which of those bookings are affected, and produces three things for each one: a decision, the reasoning behind it, and the carrier and customer emails it implies. Each of those is expressed as a change to the booking record it came from and queued back into the TMS — the exception on the booking, the new discharge port and ETA, the drafted mail on its communication log. Nothing is sent and nothing is written until a person approves it. Beside that loop sits the everyday desk: agents that work every inbound mail — quotes, bookings, documents, milestones, invoices, customs — hand each other the work as recorded messages, check every output against the customer's standing instructions, and learn from each correction a person makes. The risk loop's decisions are handed to that desk to carry out. The whole system runs on open-weight models that can be hosted inside the European Union, so the shipment data never has to leave.
The problem this addresses
A mid-size forwarder's operations desk does two jobs badly, through no fault of its own. It watches for disruption, which means someone reading trade press between other work — one narrow slice of a signal that actually arrives from broadcasters, waterway authorities, weather services, seismic networks, government registers and central banks, in that many different formats. And it reacts to disruption, which means opening the booking system, checking which boxes are exposed, weighing the options, writing the same six emails it wrote last time, and then keying every one of those decisions back into the TMS booking by booking — because until that last step happens, nothing that was decided has actually happened.
Existing risk platforms do the first job and stop. They raise an alert and leave the desk exactly where it was — knowing something is wrong, with all the work still to do. They are also priced for enterprise shippers, not for a forwarder with a few hundred bookings in the water.
So there are two gaps, and the second is the expensive one. The first is that detection is narrower than the problem — a monitor reading headlines is reading a summary of some of the signal, late. The second is that detection is where every tool stops: the decision, the two emails and the write-back are left to a person, and the write-back is the part that eats the day.
How the loop works
Six stages, run in order, on demand. The numbering is real: each stage consumes the previous stage's output — and the first and last are the same system, because the agents work on the forwarder's records rather than on a book of their own.
Every stage records why it did what it did — the loop is auditable, not just automatic
The part that makes it a loop
Risk platforms stop at stage 01. The TMS and the execution stack own stages 00 and 05 and assume someone already decided the middle. Welding the two together — reading the book out of the system of record, deciding, writing down the reasoning, and putting the result back on the same record — is the whole idea. The agents do not run beside the TMS; they run on it. A shipment rerouted on Tuesday is on the new route on Wednesday, and is no longer exposed to the thing that moved it. A shipment that is held accrues a day of delay for every day it waits, and carries that delay when it resumes. Holding is not free, and the system says so.
What is real, and what is staged
This is a prototype. Being precise about which half is which is what makes the demonstrable half worth anything.
| Component | Status | What that means |
|---|---|---|
| News monitoring | Real | Sixty live sources across six families, read on every run, all keyless. Prose is classified by a real model call against real current headlines; gauges, gusts, wave heights and magnitudes are classified by threshold instead. |
| Routing decisions | Real | The model genuinely weighs the trade-off and writes the justification. Nothing is pre-written. |
| Email drafts | Real | Generated per shipment, per audience. Never sent — there is no email library anywhere in the codebase. |
| The shipments | Synthetic | Seven authored bookings. Real ones need a customer's TMS, which sits on the far side of the trust wall. |
| The TMS link | Demo connector | Both directions are modelled and the read is the only door to the book, so every component follows it. But no TMS is contacted: there is no client, no credential and no endpoint, and a write-back is a described change that stays queued. |
| The triggering event | Scripted | So a disruption can be shown on demand instead of waiting for one. It uses the same data shape as a live event and flows through the same pipeline. |
| The everyday desk | Real | Ten agents work every inbound mail on every run: the inbox agent classifies it (one batched model call, rules as fallback), the owner does the work, a playbook agent checks each output against the customer's standing instructions and sends it back until it complies. Every handoff is a recorded message, and each one carries the one-line reason behind it — the cues that classified a mail and whether rules, the model or a lesson decided, the customer rule behind a send-back, the slack arithmetic behind an escalation. Prices, weights and dates are computed, never generated. |
| Learning from corrections | Real | A correction becomes a lesson — a rule and a line of context — and the whole inbox is replayed with it. It is kept only if it fixes the mail it came from and every earlier lesson still holds. No model is retrained. |
| The inbound mail | Synthetic | A morning of thirteen authored mails and each customer's playbook, with two mistakes left in on purpose so the learning loop has something honest to correct. No mailbox is read. |
| Per-booking desk panels | Scripted | On the disruption board, the desk agents' view of the selected booking — rate card, milestones, document fields — is authored content, and is labelled as such where it appears. |
The line we hold
“The risk detection is real — it runs against live news right now. The shipments are synthetic, so I can show you a disruption on demand instead of waiting for one.”
Where the model decides, and where it must not
The most common way an AI logistics tool loses trust is by letting a language model do arithmetic. SQRlane splits the work along a hard line.
Code computes the facts. Which chokepoints a route passes through. Which of them carry active risk. How many days an alternative adds. Whether that fits inside the schedule slack. What the revised arrival date is. What each option costs — the premium for taking it, plus what the days it lands late are worth to that customer. All ordinary, testable code. Dates and sums never go near a model.
The model makes the call and explains it. Given those facts, is the right answer to move the box, hold it, or leave it alone — and why, in language an operations lead would accept.
Pricing the options is the sharpest case of the line. A rate is what an option costs to take; it is not what the decision costs, and ranking routings by days answers only which one lands soonest. So the advisor sums a premium, a per-day cost of being late and a one-off cost of missing the date at all, then refuses any alternate that does not beat staying put in money. Every term is arithmetic and no probability is applied anywhere — the moment a model is asked how likely a cold chain is to break, the number it returns is invented, and an invented number is worse than none. The commercial terms are authored on each booking, like the bookings themselves: the arithmetic is exact, its inputs are synthetic, and nothing here is measured.
This split has a useful side effect: a booking whose route carries no active risk never reaches the model at all. That is a genuine answer, not a shortcut, and it keeps a meaningful share of the board deterministic.
Guard rails around the model's answer
- If the model proposes a route it was never offered, the reroute is refused and downgraded to a hold.
- If it returns a decision value that isn't one of the three allowed, the result falls back to no-action.
- If the provider fails entirely, transparent rules decide instead and every affected card is badged as rule-decided.
- Each of those interventions is written into the shipment's reasoning record, so a refused answer is visible rather than silently swallowed.
The model layer
Six jobs with genuinely different requirements — and most of the sixteen agents need no model at all. Sizing each job separately, and not calling a model where code gives the same answer every time, is where nearly all the cost and latency savings live.
| Task | Shape of the work | Model | In / Out | Why this one |
|---|---|---|---|---|
| Screening headlines | High volume, short inputs, strict structured output. Batched eight at a time. | Qwen3.5-9Bor Gemma-3-27B | $0.15 / $0.20 | Nearly every item is irrelevant. This is a filter, not a thinker — and it is the only stage whose call volume scales with the number of sources. |
| The routing decision | Real judgement over supplied facts, structured JSON plus prose. Three calls a cycle. | Qwen3-235B-A22Bstep up to DeepSeek-V4-Pro if the reasoning needs it | $0.20 / $0.60 | The one call worth spending on. It produces the output a customer will read and challenge, and it runs rarely enough that price barely matters. |
| Drafting the emails | Prose quality, two distinct voices, six calls a cycle — the output-heaviest stage. | Llama-3.3-70Bor Qwen3-235B-A22B | $0.13 / $0.40 | Tone is the deliverable. This stage generates the most output tokens, so the output price is the one that actually moves the bill. |
| Reading non-English sources | Short multilingual text, one pass, must not distort meaning. | Qwen3.5-9BTeuken-7B / EuroLLM if model provenance must be European too | $0.15 / $0.20 | Handled inside the screening call rather than as a separate translation step — one pass is cheaper and loses less than translate-then-classify. |
| Classifying the inbox | Every inbound mail in one batched call: which of eight intents, and one line of why. Strict JSON. One call per inbox run. | Qwen3.5-9Bor Gemma-3-27B | $0.15 / $0.20 | A labelling job with a fixed answer set, like screening. Rules answer it first anyway, and a person’s corrections are applied after the model, so a small model’s mistake is caught rather than trusted. |
| Routing and wording a question | Two short calls per question a person asks: which of the sixteen agents owns it, and the finished answer reworded so it reads like a colleague wrote it. | Qwen3.5-9Bor Gemma-3-27B; the prototype runs a Llama 3.1 8B class model | $0.15 / $0.20 | The facts are assembled by code from the run. The rewording is checked against them — every number, date, reference and status word must survive, and nothing new may appear — and the desk’s own wording is shown when it does not. Cue-phrase rules route the question when no model is configured. |
Sixteen agents, and which of them need a model
The agents are named for the job a forwarding desk already has, in two layers: the risk layer, which runs when a lane moves, and the everyday desk, which runs on every inbound mail. Five of the sixteen call a model in this build, and two more would with work this build does not do yet: reading scanned documents, and turning a customer’s written instructions into rules. The rest are code — prices, weights, dates, rule checks and field values are computed, never generated — and code is both the cheapest and the most repeatable model there is.
| Agent | Its job | Model |
|---|---|---|
| Risk layer · when a lane moves | ||
| Risk | Reads sixty sources; flags the exception on the booking a disruption threatens. | Qwen3-235B-A22Bor Qwen3.5-9B; prose only — gauges, gusts, waves and quakes go to a threshold |
| Routing | Reroute, hold or leave on plan, with the reasoning; writes the booking change. | Qwen3-235B-A22BDeepSeek-V4-Pro if the reasoning needs it |
| Comms | Drafts the carrier and the customer mail each decision implies; sends nothing. | Llama-3.3-70Bor Qwen3-235B-A22B |
| Planner | Sweeps the forward book before departure: act now, tripwire armed, or stand down. | None todayscripted |
| TMS link | The only door to the book, and every agent’s output queued back onto it. | Nonecode; demo connector |
| Everyday desk · on every mail | ||
| Inbox | What the mail is, whose it is, which agent owns it. | Gemma-3-27Bor Qwen3.5-9B; one batched call; rules as fallback |
| Playbook | Checks every output against the customer’s standing rules; sends it back until it complies. | Qwen3-235B-A22Bturns a customer’s written instructions into rules, once per customer; every check stays code. Not in this build: rules are authored |
| Rate | Prices a lane from the rate sheet for any agent that asks. | Nonearithmetic |
| RFQ | Reads a rate request into fields, drafts the quote, queues the quotation. | Noneparser and template |
| Booking | Opens or amends the booking from the mail and its documents; holds what it cannot verify. | None |
| Docs | Reads the fields off the documents and checks them against the booking. | Qwen2.5-VL-72Ba vision model for scanned PDFs and photos; field checks stay code. Not in this build, which reads text documents only |
| Milestones | Writes carrier notices onto the booking; answers where-is-my-box from the record. | None |
| Exception | Opens the exception on a rolled box or a mismatched document; escalates what the desk cannot absorb. | None |
| Invoice | Reconciles the carrier invoice against the agreed rate; drafts the dispute. | None |
| Customs | Prepares the entry for the discharge country; escalates a transit, never files. | None |
| Assistant | Takes a person’s question, hands it to the agent who owns it, and says so when none does. | Gemma-3-27Bor Qwen3.5-9B; routes and rewords; the facts are built from the run and checked; rules as fallback |
That split is the recommendation in one line: a capable mid-size model where the answer is a label that sends work somewhere, a large mixture-of-experts model where the answer is a judgement someone will challenge or a source that must not be missed, a strong writer where tone is the deliverable, a vision model where the input is a picture, and no model at all where the answer must be exact — a price, a weight, a rule check, a write to the record. Qwen3-235B-A22B earns the judgement slot because only about 22 billion of its parameters work on each token, which puts a large model’s reasoning at close to a small one’s price. Every agent also records the reason for each message it posts — the cues that classified a mail, the customer rule behind a send-back, the slack arithmetic behind an escalation — and for the agents that use no model, that reason is the computation itself.
What that costs, measured rather than assumed
A cycle under the Hamburg scenario makes fourteen calls — five to screen headlines, three decisions, six drafts. Measuring the actual prompts and outputs gives roughly 12,800 input tokens and 3,700 output tokens per cycle, at the usual four-characters-per-token approximation.
| Approach | Per cycle | 10 cycles/day | 50 cycles/day |
|---|---|---|---|
| Per-task mix as recommended above | $0.0034 | $12 / yr | $62 / yr |
| Frontier model for every call DeepSeek-V4-Pro throughout | $0.0355 | $130 / yr | $648 / yr |
The everyday desk adds one call per inbox run: about 950 input tokens for the thirteen-mail morning, measured from the prompt, and an estimated 350 out — roughly $0.0002 at the screening price, and a little more as a person’s corrections are added to the prompt as context.
The stronger picks in the agent table change one stage of the cycle: screening runs on Qwen3-235B-A22B instead of Qwen3.5-9B. Pricing every token of the cycle at that model’s rate gives an upper bound of $0.0048 a cycle — about $17 a year at ten cycles a day, $87 at fifty — still about seven times under the frontier figure. Gemma-3-27B and Qwen2.5-VL-72B are not priced in the catalogue extract used here; the vision model would be paid per document, not per cycle.
Sizing each stage separately is about ten times cheaper than reaching for a frontier model everywhere, for output a desk would not be able to tell apart on the two stages that do not need it. That ratio, not the absolute figure, is the point: inference is not the cost centre here, but carelessness with it is the easiest avoidable one.
Why there are no version numbers
This project has already been bitten once. A model name was hard-coded, the provider
retired it, and the API answered with a 404 that is indistinguishable
from a broken key — so every decision silently fell through to the rule-based
fallback while the demo appeared to work. The system now asks the provider what it
can run and picks from that list at start-up.
Treat the table above the same way: pick the class, verify the current best member of it at deployment time, and never pin a name you are not prepared to monitor. Any hard-coded model name is a scheduled outage.
Running it in Europe
Shipment data is commercially sensitive: lanes, volumes, customers, rates. The architecture is designed so none of it has to leave the EU.
Every model in the table above is open-weight, which means it can be served from infrastructure you choose rather than reached through a vendor's own cloud. That is the sovereignty argument in full — not a policy promise about where data is processed, but an arrangement in which the question does not arise.
Lyceum Inference
Studio, run from Berlin, serves a catalogue of open-weight models from
eu-north1 on a pay-per-token basis, with no commitment. Three properties
of it matter here more than the price does.
| Property | Why it matters to SQRlane |
|---|---|
| EU data residencydefault models run in eu-north1 | Booking data, customer names and rates are processed in Europe. This is the claim the whole page rests on, and it is a deployment fact rather than an intention. |
| Zero data retentionprompts and outputs not stored or trained on | A forwarder's lane economics are in those prompts. "Not retained, not trained on" is the difference between a demo and something a commercial director will sign off. |
| Structured output and function calling | Not a nicety. Every decision this system makes comes back as JSON that gets validated and guard-railed; a provider without reliable structured output cannot run this pipeline at all. |
| Drop-in OpenAI-compatible APIone base URL | The provider layer here is a single file. Migration is a base-URL change and a model name, not a rewrite — which is the only reason the switch is realistic at prototype stage. |
| Pay per token, no commitment | SQRlane is bursty by design: idle until a disruption is detected or someone asks, then a short burst of calls. Reserved GPU capacity would be mostly paying for silence. |
Why per-token beats renting a GPU here
It is worth being explicit, because the alternative looks cheaper on a spreadsheet. A dedicated GPU is priced by the hour whether or not anything is running. This system makes roughly fourteen calls when a disruption lands and then nothing for hours. At ten cycles a day the recommended model mix costs about twelve dollars a year; a single rented GPU left running would cost more than that before lunch on the first day. Reserved capacity starts to make sense when the desk is large enough to keep the endpoint genuinely busy, or when a fixed latency SLA is worth paying for — and dedicated endpoints exist for that case.
Where these numbers are soft
Token counts are measured from the real prompts this system builds, but converted at the usual four-characters-per-token approximation rather than by the actual tokeniser — expect a margin either way, and more on non-English text, which tokenises less efficiently. Output length varies with how much the model chooses to write. Catalogue prices are dated July 2026 and stated as subject to change. None of this moves the conclusion, which is about the ratio between approaches rather than the absolute figure.
Built in Europe, for Europe
That phrase is doing a lot of marketing work across the sector at the moment, so it is worth being exact about what it means here and what it does not.
Two different claims, often blurred together
There are two separate axes here and the marketing usually collapses them. Residency is where your data is processed. Provenance is where the model weights came from. They are not the same thing, and only one of them is settled by choosing an EU host.
| Axis | What is true | Standing |
|---|---|---|
| Data residency | Processing in eu-north1, no retention, no training on your prompts. Booking data, customer names and rates stay in Europe. |
Achievable now |
| Model provenance | The strongest open-weight models in the catalogue come from Chinese and American labs — Qwen, DeepSeek, Moonshot, Llama, Gemma. They are openly licensed, but they are not European in origin. | Partly |
| European provenance, if required | Teuken-7B came out of a German research consortium and covers all twenty-four official EU languages; EuroLLM is an EU-funded project with the same coverage; Mistral is French. All openly licensed — but smaller, and they would need confirming as available on whichever host you pick. | Trade-off |
For most forwarders the residency axis is the one that matters, because the risk they are managing is commercial confidentiality rather than model lineage. If a procurement process demands European weights as well, that is available — at some cost in capability on the harder reasoning stage.
What is genuinely European here
- The problem is. The lanes, chokepoints and regulations modelled are the North Range ports, the Rhine, the Mediterranean corridors and EU customs entry — not a global abstraction with European labels applied.
- The processing can be. EU-hosted inference with zero retention, from a European provider.
- The data need never leave. Because the models are open-weight, they are served where you choose. That is a property of the architecture, not a clause in a contract.
- It suits the regulatory direction. Recorded reasoning for every decision, and a human approval gate before anything leaves the building, are exactly the properties an EU AI Act conversation asks about.
What it does not mean yet
The prototype as it stands calls a US inference provider. It was built against a free tier for speed of iteration. Everything above describes an architecture that supports EU-only operation, and a provider layer designed to make the switch cheap — but the switch has not been made, and no one should read the current deployment as sovereign. Saying otherwise would be exactly the kind of unbacked claim this project refuses to make everywhere else.
Assumptions, and how this breaks
Every item below is a real failure that occurred during development, not a hypothetical risk register.
| Assumption | What actually happened | Standing |
|---|---|---|
| The model we ship keeps working | The provider retired it. The API returned a 404 identical to a bad-key error, so every decision fell back to rules while the demo appeared healthy. |
FixedModel resolved at start-up from what the key can actually run |
| The API answers when called | Free-tier rate limits. The last call in a cycle — the customer email — reliably hit a 429 and quietly fell back to a template. |
MitigatedBounded retry honouring the provider's own back-off; the fallback now states its reason on screen |
| A dead news source fails fast | It does. A slow one does not: a source that answers but sends a byte at a time never trips a between-bytes timeout, and one such feed hung an entire run indefinitely. | FixedTotal deadline and size cap per source, with a watchdog that closes the socket |
| Reading sources in turn is fine | Not inside a serverless time limit. Sequentially, only the first source would have been read — collapsing the breadth the whole approach depends on. | FixedSources read concurrently under a hard overall budget |
| Scripted and live events share a shape | Broken three separate times, each by adding a field to one and not the other. | WatchedTested both ways; still the most common way a change breaks this system |
| A defensible day is a defensible week | It is not. Across a simulated week a booking bounced between two ports on consecutive days. Every single day's arithmetic was correct; the sequence was nonsense. | FixedA booking is never re-offered the route it just left |
| Colours chosen by eye are readable | Two chart categories sat close enough to be indistinguishable under the most common form of colour blindness. It looked fine on screen. | FixedPalette validated against contrast and colour-vision thresholds rather than judged visually |
The pattern worth taking away
Every one of those was found by looking at real output, not by reasoning about the code. The rate limit in particular survived three rounds of confident wrong diagnosis until the fallback was made to state its own reason on screen. A system that degrades silently is worse than one that fails loudly, because you will demonstrate it in that state without knowing.
What this is not
Scope discipline is what kept this buildable. None of the following are in it, and none are accidental omissions.
- Not route optimisation. Candidate routes are pre-defined. The system chooses among them and justifies the choice; it does not compute new ones.
- Not a sending system. Emails, booking amendments, customs entries and TMS write-backs are all drafted and held. There is no transport layer in the codebase, and a test fails the build if one is ever imported.
- Not connected to a real TMS. Working through the TMS is the design and both directions are modelled — the read is the only door to the book, and every agent action is expressed as a change to a booking record. What is absent is the far end: no vendor, no credential, no endpoint. It contacts nothing.
- Not self-training. The everyday desk learns from corrections, but a lesson is a rule and a line of context a person can read and delete. No weights change, and a lesson that would undo an earlier correction is refused.
- Not always-on. It runs on demand. For a demonstration, a button beats a background job.
- Not measured. There are no accuracy figures, hit rates or time-saved claims anywhere in this document, because none have been measured. Inventing them is the fastest way to lose a room that knows the domain.
What production would require
The distance between this prototype and a product is not mostly model work. It is integration, trust and liability.
- A real TMS connection. The shape is already here — one read path and a write-back per agent action — so the work is the far end of it: a vendor's API, credentials, field mapping against their schema, and the reconciliation that follows when the two systems disagree. Replacing the synthetic pool with a live book is the single highest-value next step, and everything downstream of it already assumes it.
- Inference moved in-region. Deploy the model classes above on EU capacity and benchmark them, replacing the illustrative costs here with measured ones.
- Evaluation the domain would accept. A held-out set of historical disruptions, scored on whether the decision matches what the desk actually did — the only credible answer to “how good is it”.
- Sending, deliberately and last. The approval gate exists because trust is earned in that order. Drafting well for months is the argument for eventually being allowed to send.