Skip to content

Answers to the assignment questions

Same order and headings as assignment.md. The full system is in design.md; decisions are in docs/adr/. “Built” means it is in the code and tested; “designed” means described only.

Film, 0:59 Load 481207, step by step

The 5 steps of the assignment timeline as the agent processed them, with the ETA ask beside each one.

Scene list (text version)
  1. 0:00 Load 481207: Client A, Riverbend Foods, dry van, Riverbend plant to Valley DC.
  2. 0:12 7:01 AM: Ping turns delay risk high: SMS the driver for the ETA and set the 7:31 AM follow-up. Actions: send sms "Load 481207: what is your ETA to Riverbend plant? Please reply with the time."; timer for 7:31 AM.
  3. 0:21 7:13 AM: Driver asks for the PO: answer it from metadata, the ETA ask stays open. Actions: send sms "Load 481207: the PO number is 55821.".
  4. 0:32 7:31 AM: Timer: no ETA yet, ask again and set the 8:01 AM follow-up. Actions: send sms "Load 481207: what is your ETA to Riverbend plant? Please reply with the time."; timer for 8:01 AM.
  5. 0:40 8:01 AM: Timer: two requests unanswered, email the dispatcher in the main thread and create an urgent Hero task. Actions: send email "The driver on load 481207 has not answered our texts. Please send the ETA to Riverbend plant."; create hero task "Load 481207: Call driver for eta: Call the driver on load 481207 to get the ETA to Riverbend plant. No ETA arrived after our requests.".
  6. 0:49 8:30 AM: Nothing is left to do: the escalation chain is finished.
  7. 0:55 1 model call in 5 steps; every step is asserted in client_a_load_481207_full_timeline.

Generated from the scenario export at build time. Download MP4 (1.4 MB) with captions (WebVTT). Narration and sound effects generated with ElevenLabs.

1. What is the architecture of your agent? What are the alternatives?

Section titled “1. What is the architecture of your agent? What are the alternatives?”

A hybrid, event-driven pipeline, serial per load, where code owns state and procedure and the LLM is the exception (ADR 0001, design.md sections 3 and 4).

  • Every input (SMS, email, Slack, ping, timer, attachment, stop status, delivery outcome) and every decision is appended to a per-load log. State is a projection of that log, including open asks.
  • Per event: deterministic filter, trigger diff, SOP routing, triage (a model call only for a message that survives the filter), SOP executor or planner, action validation, then one transaction that commits the event, the decision record and the outbox rows. A dispatcher sends from the outbox with idempotency keys. Timers are durable and re-enter through the same queue.
  • The SOP is compiled at edit time into a state machine (ask, retries, wait, escalation chain). The runtime executes it without a model. The principle: the model answers typed questions about meaning, code owns values and actions. The built triage model is the TypeSafe Jev decision model (ADR 0005), which returns probabilities for closed questions and never text or values. Code finds and normalizes value candidates (llm/extract.py) and the model only confirms one. A DeepSeek planner proposes actions for situations no SOP covers, with no carrier-facing actions (a Hero task, a TMS note, or a Slack post to the broker).
AlternativeTrade-off
Single ReAct agent with tools, full history per eventSimplest and most flexible. Cost grows with history on every event, retry counting rests on model recall, and step-by-step testing is hard. Does not fit $1 per load at 50 to 100 messages plus pings.
Workflow engine with no LLMPredictable and cheap. Cannot read “What’s the PO number?” or draft text. We kept its good parts (durable timers, retries, idempotency) and chose Postgres over Temporal so the decision and its actions share one transaction (ADR 0003).
Multi-agentMore hand-offs for a problem whose core is one state machine per load. Cross-channel resolution needs shared state anyway.
Chat history or vector memoryProbabilistic and not replayable.
A generative model that reads each message and returns the valueOne call does everything, but it can invent a time and its confidence is uncalibrated. We use a decision model for meaning and code for values (ADR 0005).

Costs of ours: upfront modeling (events, asks, transitions), coverage depends on the SOP corpus (uncovered cases go to the planner or a Hero), and the outbox adds a dispatcher and a few seconds of delivery latency. Code: src/load_agent/agent/, domain/, store/.

Interactive Where the 91 runs of the benchmark load were decided
  1. Eventappended to the load log; the run projects the state from it
  2. Filter Pings with no band change, stale pings, acknowledgements with no open ask 36 runs, 39.6% code only, $0
  3. Trigger and SOP A risk band change, a stage change or a timer runs a compiled SOP 8 runs, 8.8% code only, $0
  4. Triage The decision model reads one message against the open asks 35 runs, 38.5% model call
  5. Hero on doubt Unknown sender, or triage not sure enough: Hero task 9 runs, 9.9% 7 of 9 called a model
  6. Planner No SOP covers it: a bounded plan, never carrier-facing 3 runs, 3.3% model call
  7. Validate, commit, dispatchevery action: contact on the load, channel supported, role allowed, idempotency key unused; then one transaction with the outbox

46 of 91 runs (50.5%) never called a model. Of the 87 inputs, 45 (51.7%) reached one; no timer did.

Table view
Stage where the run was decidedRunsShareRuns with a model call
Filter3639.6%0 of 36
Trigger and SOP88.8%0 of 8
Triage3538.5%35 of 35
Hero on doubt99.9%7 of 9
Planner33.3%3 of 3
Total91100%45 of 91

Source: bench/results/2026-10-08-typesafe-deepseek.json (87 inputs and 4 timers, seed 1). Full architecture.

Film, 1:05 How triage decides

Four recorded answers of the decision model, and why each became a value or a Hero task.

Scene list (text version)
  1. 0:00 Code finds candidate values; the decision model answers yes or no questions with probabilities; code applies the thresholds; unsure goes to a Hero.
  2. 0:10 "ETA 9:30am": main 0.82 (at least 0.60), highest guard 0.08 (at most 0.30): accepted, ETA 09:30.
  3. 0:22 "9:30am": main 0.59 (at least 0.60), highest guard 0.28 (at most 0.30): Hero task.
  4. 0:29 "tow eta 9:30": main 0.74 (at least 0.60), highest guard 0.68 (at most 0.30): Hero task.
  5. 0:43 "eta was 9:30": main 0.63 (at least 0.60), highest guard 0.90 (at most 0.30): Hero task.
  6. 0:54 0 wrong values in 190 phrasings over two recordings; 62/74 and 61/74 positives accepted.

Generated from the scenario export at build time. Download MP4 (1.3 MB) with captions (WebVTT). Narration and sound effects generated with ElevenLabs.

2. How does the agent read documents, find their type and get their key values?

Section titled “2. How does the agent read documents, find their type and get their key values?”

Classify, extract with a schema per type, cross-validate in code, escalate when unsure (design.md section 10).

  1. An AttachmentReceived event carries the file URI, media type and sender.
  2. One vision-model call (LLMClient.extract_document) returns the document type (bol, pod, rate_confirmation, receipt, unknown), the key fields and a confidence.
  3. Fields are per type: BOL (BOL number, PO number, seal number, pieces, weight, signature), POD (receiver, delivery time, signature, exceptions), rate confirmation (rate, stops, appointments, MC number, load id), receipt (amount, vendor, date, kind).
  4. Cross-validation is code, not a model: the extracted PO and BOL numbers are compared with LoadMetadata.reference_numbers. Designed additions: weight, load id, MC number, stop names, rate against the TMS. Photo downscaling and duplicate-file dropping before the call are also designed.
  5. Unknown type, confidence under 0.7, or any mismatch becomes a Hero task naming the field and both values, and no SOP step runs. A trusted document fires a DocumentReceived trigger, so the client’s SOP for that type decides (for example a TMS note that the BOL arrived).

Alternatives: OCR plus rules (cheap, breaks on photos and layout changes); a single prompt asking the model whether the document matches the load (the model can agree with a wrong document; an equality check cannot); auto-correct the load from the document (a misread digit would silently change a contract value). Trade-off: low-quality phone photos fall below the confidence gate and cost Hero time, never silent errors. Built: the seam, schemas, the 0.7 gate, PO and BOL cross-check, escalation (agent/reactions.py). Designed: the real vision client, per-type schemas, the wider cross-check set.

Interactive From attachment to action
Document Doubt
  1. Attachment received An AttachmentReceived event carries the file, media type and sender. Built
  2. Find the type One vision call (LLMClient.extract_document) returns the type: BOL, POD, rate confirmation, receipt or unknown. Built seam, designed client
  3. Extract fields A schema per type: BOL and PO numbers, seal, pieces, weight; POD receiver, time, exceptions; and so on. Designed
  4. Cross-check in code The PO and BOL numbers must equal the load reference numbers. An equality check, not a model opinion. Built
  5. Gate Unknown type, confidence under 0.7 or any mismatch: Hero task naming the field and both values. No SOP step runs. Built
  6. Trusted: route A DocumentReceived trigger routes to the client SOP for that type. Built

SOP Runs BOL received (standard/bol_received@v1)

Table view
DocumentTrusted (confidence at least 0.7, numbers match)Confidence under 0.7PO or BOL mismatch
Bill of ladingRuns BOL received (standard/bol_received@v1)Hero task: the document could not be identified with confidenceHero task naming the field and both values
Proof of deliveryRuns POD received (standard/pod_received@v1)Hero task: the document could not be identified with confidenceHero task naming the field and both values
Rate confirmationNo SOP for this type: recorded, nothing else to doHero task: the document could not be identified with confidenceHero task naming the field and both values
Unknown typeHero task: the type could not be identifiedHero task: the document could not be identified with confidenceHero task naming the field and both values

Source: agent/reactions.py (gate 0.7), routing from sops/ for Client A. Built means in the code and tested.

1. How do you organize the SOPs of many clients? The standard procedure applies to all clients until a client changes it.

Section titled “1. How do you organize the SOPs of many clients? The standard procedure applies to all clients until a client changes it.”

Layered files per topic, most specific wins (ADR 0002):

sops/standard/delay_risk_high.md
sops/clients/client_a/delay_risk_high.md # overrides standard for client_a
sops/clients/client_b/reefer/departed_pickup.md # client + load type

Order: client + load_type over client over standard. The winning file replaces the topic as a whole. A new client starts with zero files and gets the standard for every topic (the 60 to 70% shared part); it overrides only what differs. Each Markdown file has a compiled, versioned sibling (delay_risk_high.v1.compiled.yaml), and templates.v1.yaml files layer the same way for wording.

Alternative: section-level merge, where a client patches part of a standard topic. Smaller client files, but editors cannot predict “which step 2 applies”. Whole-topic override is easy to preview and diff; the cost is duplicated text when a client changes one line. The compiler can flag drift from standard as a review hint (designed). Built: FileSopRepository (sop/repository.py).

Interactive Which file wins: standard, client, client plus load type
  1. Client + load type ... ...
  2. Client ... ...
  3. Standard ... ...

Client + load type client_b/reefer/departed_pickup@v1

When the load enters the stage 'to delivery' from 'at pickup':
  1. Write a TMS note (record that the driver left the pickup).
  2. Text the driver (ask for the trailer temperature and the seal number). Wait for the trailer temperature and seal number.
     The request stays open. There is no follow-up.
  The request is answered as soon as we receive the trailer temperature and seal number from any contact of the load on any channel.

Behavior change compared with standard/departed_pickup@v1:

- Now sends a request and waits for an answer; it did not before.
- May contact the driver again.
- May contact the dispatcher again.
- Guidance (affects wording only, not what happens): added 'Wording'.
Table view
Topics that a client or a client and load type override; every other topic resolves to the standard file.
LayerClientLoad typeTopicFile
Client Client A any Delay risk high sops/clients/client_a/delay_risk_high.md
Client Client B any Delay risk high sops/clients/client_b/delay_risk_high.md
Client + load type Client B reefer Departed pickup sops/clients/client_b/reefer/departed_pickup.md

Source: the sops/ tree (11 compiled files) resolved by FileSopRepository; text from load-agent sop preview.

2. How does the agent find the correct SOP files for the current event? What are the advantages and disadvantages of your method?

Section titled “2. How does the agent find the correct SOP files for the current event? What are the advantages and disadvantages of your method?”

Deterministic routing, no retrieval on the load path. The event becomes a trigger by diffing projected state: DelayRiskBecame(high, stop), StageEntered(to_delivery, from at_pickup), DocumentReceived(bol). The compiler builds a routing index from triggers to topics. SopRepository.topics_for(trigger, client_id, load_type) is a lookup that returns the winning layer. If the winning spec has conditions (for example no_open_ask: eta_request), the executor checks them against state. Messages are not routed to SOPs: they are matched to open asks, which already carry the SOP version they run under.

For a situation with no topic the pipeline goes to triage, then the planner (no carrier-facing actions (a Hero task, a TMS note, or a Slack post to the broker), at most 3 steps, confidence at least 0.8), and the default is a Hero task.

AdvantagesDisadvantages
Zero cost and milliseconds; testable per triggerA new situation needs a new topic and a compile
A client override can never be shadowed by a similar chunk from another layerFallback coverage depends on the planner and on Hero capacity
The same input always routes the same way, so evals are exactThe closed vocabulary of trigger kinds needs a code change to extend

Alternative: embed SOP chunks and retrieve per event. It handles unseen phrasing, but it is nondeterministic, can mix layers, and costs an embedding plus a longer prompt per event. We keep it only as a possible fallback to suggest a topic to a Hero. Built: sop/routing.py, agent/triggers.py.

Interactive Route an event: a lookup, not a search
  1. Trigger from the state diff...
  2. Routing index lookup...
  3. DecisionSOP ...

An inbound message is never routed to an SOP: it is matched to the open asks, which already carry the SOP version they run under.

Table view
EventClient A, dry vanClient A, reeferClient B, dry vanClient B, reeferAny other client, dry vanAny other client, reefer
Delay risk for the pickup becomes highDelay risk high, client layer (client_a/delay_risk_high@v1)Delay risk high, client layer (client_a/delay_risk_high@v1)Delay risk high, client layer (client_b/delay_risk_high@v1)Delay risk high, client layer (client_b/delay_risk_high@v1)Delay risk high, standard layer (standard/delay_risk_high@v1)Delay risk high, standard layer (standard/delay_risk_high@v1)
Delay risk for the pickup becomes mediumNo SOP: noopNo SOP: noopNo SOP: noopNo SOP: noopNo SOP: noopNo SOP: noop
Driver arrives at the pickupArrived at pickup, standard layer (standard/arrived_at_pickup@v1)Arrived at pickup, standard layer (standard/arrived_at_pickup@v1)Arrived at pickup, standard layer (standard/arrived_at_pickup@v1)Arrived at pickup, standard layer (standard/arrived_at_pickup@v1)Arrived at pickup, standard layer (standard/arrived_at_pickup@v1)Arrived at pickup, standard layer (standard/arrived_at_pickup@v1)
Driver departs the pickupDeparted pickup, standard layer (standard/departed_pickup@v1)Departed pickup, standard layer (standard/departed_pickup@v1)Departed pickup, standard layer (standard/departed_pickup@v1)Departed pickup, client + load type layer (client_b/reefer/departed_pickup@v1)Departed pickup, standard layer (standard/departed_pickup@v1)Departed pickup, standard layer (standard/departed_pickup@v1)
Driver arrives at the deliveryArrived at delivery, standard layer (standard/arrived_at_delivery@v1)Arrived at delivery, standard layer (standard/arrived_at_delivery@v1)Arrived at delivery, standard layer (standard/arrived_at_delivery@v1)Arrived at delivery, standard layer (standard/arrived_at_delivery@v1)Arrived at delivery, standard layer (standard/arrived_at_delivery@v1)Arrived at delivery, standard layer (standard/arrived_at_delivery@v1)
Load marked deliveredNo SOP: noopNo SOP: noopNo SOP: noopNo SOP: noopNo SOP: noopNo SOP: noop
A bill of lading is receivedBOL received, standard layer (standard/bol_received@v1)BOL received, standard layer (standard/bol_received@v1)BOL received, standard layer (standard/bol_received@v1)BOL received, standard layer (standard/bol_received@v1)BOL received, standard layer (standard/bol_received@v1)BOL received, standard layer (standard/bol_received@v1)
A proof of delivery is receivedPOD received, standard layer (standard/pod_received@v1)POD received, standard layer (standard/pod_received@v1)POD received, standard layer (standard/pod_received@v1)POD received, standard layer (standard/pod_received@v1)POD received, standard layer (standard/pod_received@v1)POD received, standard layer (standard/pod_received@v1)
A rate confirmation is receivedNo SOP: noopNo SOP: noopNo SOP: noopNo SOP: noopNo SOP: noopNo SOP: noop
A document of unknown type arrivesHero taskHero taskHero taskHero taskHero taskHero task

Source: route_trigger in sop/routing.py over the sops/ tree, for each client and load type.

3. People without a technical background write and change the SOPs. How does your design make this safe? You do not have to build an editor or a platform. Describe the design.

Section titled “3. People without a technical background write and change the SOPs. How does your design make this safe? You do not have to build an editor or a platform. Describe the design.”

The editor writes plain English and approves what the agent understood, not YAML, behind a pipeline of gates (design.md section 8, ADR 0002):

  1. Compile on save. A strong model extracts the procedure; code validates it. Ambiguity is rejected with a reason the editor can act on (“if the driver doesn’t answer soon” is ambiguous_step), as are unknown roles, a step that contacts a role the same SOP forbids, and waits or counts not grounded in the text.
  2. Plain-English preview, rendered by code from the compiled spec, not written by a model. What the editor approves is exactly what runs. Real output is in design.md section 8 (uv run load-agent sop preview client_a delay_risk_high).
  3. Behavior diff (designed on scenario suites; built between specs), not a text diff: which decisions change on the scenario suites (“at 08:01 the agent now posts to Slack instead of emailing the dispatcher”).
  4. Eval gates: standard suite plus this client’s suite must pass.
  5. Shadow mode on live loads: the new version decides, does not act, and differences with the current version are reviewed.
  6. Publish creates an immutable version and refuses a stale version number when two editors publish at once. Rollback republishes an old spec as the newest version.
  7. Version pinning: an ask started under v1 finishes under v1, so an edit cannot change a wait or an escalation target in the middle of a chain.
  8. Guidance that is not procedure (“keep texts short”) stays text and cannot change what happens. Message wording lives in templates, versioned separately.

Alternatives: editors write YAML or a DSL (exact, but not for non-technical people); the model reads raw Markdown at runtime (no compile step, but behavior shifts with phrasing and model version, and nothing can be diffed). Trade-off: the compile step adds latency to an edit (seconds) and a closed vocabulary; the compile heuristics fail safe by rejecting. Built: preview, behavior diff, validation rules, stale-source detection. Designed: the editor, scenario-based diff, shadow mode, the publish service.

Interactive From an editor's sentence to a published version

Designed: editorBuilt: Markdown files The editor writes plain English. No YAML, no DSL.

sops/clients/client_a/delay_risk_high.md

1. Send an SMS to the driver to get the ETA.
2. If there is no ETA after 30 minutes, send the SMS again.
3. If there is still no ETA 30 minutes after the second SMS, send an email to the dispatcher in the main email thread.
4. Then create an urgent Hero task for a Hero to call the driver.
Table view
  1. Edit (Designed: editor; Built: Markdown files): The editor writes plain English. No YAML, no DSL.
  2. Compile (Built): A strong model extracts the procedure, code validates it. Rejections the editor can act on: ambiguous_step, unknown_contact_role, forbidden_role_used, unsupported_action, missing_trigger, conflicting_steps, forbidden_role_not_declared, schema_invalid.
  3. Preview (Built): Rendered by code from the compiled spec, not by a model: what the editor approves is what runs.
  4. Behavior diff (Built: between specs; Designed: on scenario suites): What changes compared with standard/delay_risk_high@v1, in behavior, not text.
  5. Eval gates (Built: all scenarios run in CI on every commit; Designed: per-client suite selection, publish gate): The standard suite plus this client's suite must pass. Today: 30 scenario files (27 for Client A) and 1503 tests in CI.
  6. Shadow (Designed): On live loads the new version decides but does not act; differences with the current version are reviewed.
  7. Publish and pin (Designed: publish service; Built: version pinning): Publish creates an immutable version and refuses a stale version number. An ask started under client_a/delay_risk_high@v1 finishes under it. Rollback republishes an old spec as the newest.

Source: sops/clients/client_a/delay_risk_high.md, sops/clients/client_a/delay_risk_high.v1.compiled.yaml, sop/preview.py, evals/scenarios/.

1. How does the agent know its past actions and the past messages of the load?

Section titled “1. How does the agent know its past actions and the past messages of the load?”

From the per-load event log; state is a projection of it (ADR 0001). Inbound events, DecisionRecorded entries (which hold the actions taken, the SOP version, ask transitions and the rationale) and delivery outcomes are in one ordered log. project_state(records) folds it into LoadState: stage, delay risk per stop, open asks, processed event ids, committed action keys. The agent reads the log at the start of each run and never relies on a model’s recall of earlier turns.

The model sees only what a call needs: the one message, the open asks (id, type, what they await, target role), the stop timezone. Not the history.

Alternatives: chat history in the prompt (cost grows, recall errors), a vector store of messages (probabilistic), a mutable state row (no audit, no replay). Trade-off: modeling effort, and stored models must remain readable across schema changes. Built: store/, domain/.

2. At 07:13, how does the agent remember that the ETA request is still open?

Section titled “2. At 07:13, how does the agent remember that the ETA request is still open?”

It is a row of state, not a memory. At 07:01 the decision for ping ping_0700 committed an AskOpened transition for ask:eta_request:ping_0700 with attempts: 1, target drv_20931, next_timer: 07:31, resolved_by: null, pinned to client_a/delay_risk_high@v1. At 07:12 the driver asks “What’s the PO number?”. The run projects the log, sees one open ask awaiting an ETA, and triage returns “question about the PO number, no ETA”. satisfies(resolution, extracted, ...) is false because no ETA was extracted, so nothing resolves the ask. The agent answers “Load 481207: the PO number is 55821.” from reference_numbers.po_number (parsed from instructions at registration when the field is absent; see design.md, Load metadata; with no PO found it escalates to a Hero and never invents one) and commits no ask transition. The ask, its attempt count and the 07:31 timer are untouched, so at 07:31 the follow-up fires.

The PO reply takes no ask step, so its idempotency key comes from its triggering event and cannot collide with the 07:01 key. Eval: client_a_load_481207_full_timeline (and unrelated_driver_question_is_answered_and_ask_stays_open). Alternative: ask the model “did we already request the ETA?” over the chat; that costs a call and can be wrong. Also see ok_with_open_eta_ask_keeps_ask_open: an “ok” with the ask open is a noop that leaves the ask open.

Animated, scrubbable Load 481207: the log, and the state projected from it

7:13 AM Driver asks for the PO: answer it from metadata, the ETA ask stays open

Event log (append only)

  1. 1 7:00 AM Location ping ping_0700 delay risk high for stop 1 (Riverbend plant)
  2. 2 7:01 AM Decision recorded path sop, client_a/delay_risk_high@v1 SMS to drv_20931: "Load 481207: what is your ETA to Riverbend plant? Please reply with the time."Timer set for 7:31 AMAskOpened ask:eta_request:ping_0700, attempt 1
  3. 3 7:12 AM SMS received sms_0712 from drv_20931 (driver)"What's the PO number?"
  4. 4 7:13 AM Decision recorded path triage SMS to drv_20931: "Load 481207: the PO number is 55821."no ask transition1 model call: triage
  5. 5 7:31 AM Timer fired evt:tmr:481207:ask:eta_request:ping_0700:1 ask:eta_request:ping_0700, attempt 1
  6. 6 7:31 AM Decision recorded path sop, client_a/delay_risk_high@v1 SMS to drv_20931: "Load 481207: what is your ETA to Riverbend plant? Please reply with the time."Timer set for 8:01 AMAsk ask:eta_request:ping_0700: attempt 1 to 2
  7. 7 8:01 AM Timer fired evt:tmr:481207:ask:eta_request:ping_0700:2 ask:eta_request:ping_0700, attempt 2
  8. 8 8:01 AM Decision recorded path sop, client_a/delay_risk_high@v1 Email to dsp_4410: "The driver on load 481207 has not answered our texts. Please send the ETA to Riverbend plant."Hero task (urgent): "Load 481207: Call driver for eta: Call the driver on load 481207 to get the ETA to Riverbend plant. No ETA arrived after our requests."Ask ask:eta_request:ping_0700: escalations 0 to 2, status open to escalated
  9. 9 8:30 AM Clock check no record no event is due, nothing is appended

Projected state (project_state over the log)

Stage
to pickup
Delay risk, Riverbend plant
high
Delay risk, Valley DC
low
Open asks
1

Open ask:eta_request:ping_0700

Awaits
ETA from drv_20931 (driver), Riverbend plant
Attempt
1 of 2
Next timer
7:31 AM
Escalations done
0
Resolved by
null
Pinned SOP
client_a/delay_risk_high@v1

unchanged: no ask transition was committed

Table view
#TimeRecordDetailsETA ask after the record
1 7:00 AM Location ping (ping_0700) delay risk high for stop 1 (Riverbend plant) none
2 7:01 AM Decision recorded (path sop, client_a/delay_risk_high@v1) SMS to drv_20931: "Load 481207: what is your ETA to Riverbend plant? Please reply with the time."; Timer set for 7:31 AM; AskOpened ask:eta_request:ping_0700, attempt 1 Open, attempt 1 of 2, next timer 7:31 AM
3 7:12 AM SMS received (sms_0712) from drv_20931 (driver); "What's the PO number?" Open, attempt 1 of 2, next timer 7:31 AM
4 7:13 AM Decision recorded (path triage) SMS to drv_20931: "Load 481207: the PO number is 55821."; no ask transition; 1 model call: triage Open, attempt 1 of 2, next timer 7:31 AM
5 7:31 AM Timer fired (evt:tmr:481207:ask:eta_request:ping_0700:1) ask:eta_request:ping_0700, attempt 1 Open, attempt 1 of 2, next timer 7:31 AM
6 7:31 AM Decision recorded (path sop, client_a/delay_risk_high@v1) SMS to drv_20931: "Load 481207: what is your ETA to Riverbend plant? Please reply with the time."; Timer set for 8:01 AM; Ask ask:eta_request:ping_0700: attempt 1 to 2 Open, attempt 2 of 2, next timer 8:01 AM
7 8:01 AM Timer fired (evt:tmr:481207:ask:eta_request:ping_0700:2) ask:eta_request:ping_0700, attempt 2 Open, attempt 2 of 2, next timer 8:01 AM
8 8:01 AM Decision recorded (path sop, client_a/delay_risk_high@v1) Email to dsp_4410: "The driver on load 481207 has not answered our texts. Please send the ETA to Riverbend plant."; Hero task (urgent): "Load 481207: Call driver for eta: Call the driver on load 481207 to get the ETA to Riverbend plant. No ETA arrived after our requests."; Ask ask:eta_request:ping_0700: escalations 0 to 2, status open to escalated Escalated, attempt 2 of 2, next timer none
9 8:30 AM Clock check (no record) no event is due, nothing is appended Escalated, attempt 2 of 2, next timer none

Source: the scenario export of client_a_load_481207_full_timeline (open in the explorer).

3. How does a follow-up timer work? What does the agent see when the timer starts it?

Section titled “3. How does a follow-up timer work? What does the agent see when the timer starts it?”

A timer is a durable claim that something may be due; state decides whether it is.

  1. set_timer is an outbox action committed in the same transaction as the decision that took the step. The dispatcher writes it to the timer store with timer_id = tmr:{load_id}:{ask_id}:{attempt}. fire_at = decided_at + retry_after.
  2. The timer pump polls (at least every 30 s) for due timers, publishes TimerFired(event_id derived from timer_id, payload {load_id, ask_id, attempt}) to the load queue, then marks it fired. A crash between publish and mark publishes twice; the event store drops the second copy by event id.
  3. The agent sees a normal event on the load’s queue, projects the state, and checks ask.is_timer_current(payload): the ask must be open and payload.attempt must equal the ask’s current step. Otherwise it is a noop with no model call: the ask was resolved (ETA arrived by email at 07:20), the load arrived, the step moved on, or the timer is a duplicate.
  4. If current, the executor continues the ask under its pinned SOP version: 07:31 sends the second SMS and sets the 08:01 timer; 08:01 runs the escalation chain (email the dispatcher in the main thread, urgent Hero task).

The timer also closes an ask when the risk eased: if the stop’s delay risk was read LOW after the ask’s last text, the current timer closes the ask with no text and no escalation; MEDIUM keeps the chain (risk_falls_back_cancels_the_eta_chain). Cancellation is housekeeping; correctness never depends on it. The timer survives deploys because it lives in the database; a late timer is harmless because state decides. Evals: timer_after_eta_received_is_noop, duplicate_timer_delivery_is_noop, driver_arrives_before_follow_up_cancels_eta_ask.

Alternatives: Temporal timers (strong, but a second source of truth next to our log); EventBridge Scheduler (per-schedule cost and quotas at our volume, cancellation can fail; kept as a drop-in TimerStore); in-process sleeps (lost on deploy). Trade-off: a pump to operate and monitor (lag alert at 2 minutes). Built: store/timers.py, agent/timer_pump.py.

Interactive timeline Timers: set, fired, stale and duplicated
  • Ask opened, timer set
  • Ask resolved
  • Timer fired and acted
  • Noop
  • Timer pending

timer_after_eta_received_is_noop A timer delivered after the ETA arrived is a filtered noop with no LLM call.

duplicate_timer_delivery_is_noop The same timer delivered twice produces one follow-up SMS; the copy is dropped.

7:25 AM. A timer for attempt 1 is delivered although the ask is resolved: noop, no LLM. Path filtered. timer tmr:481207:ask:eta_request:ping_0700:1 is stale: the ask is closed or already past that step. No action. Model calls: 0.

Table view
Scenario#TimeWhat happenedDecision
timer_after_eta_received_is_noop17:01 AMDelay risk high: SMS the driver for the ETAPath sop. followed client_a/delay_risk_high@v1. Actions: send sms, timer set for 7:31 AM. Model calls: 0.
timer_after_eta_received_is_noop27:20 AMDispatcher emails the ETA: ask resolvedPath triage. ask:eta_request:ping_0700 answered by email_0720. Actions: cancel timer, tms note. Model calls: 1.
timer_after_eta_received_is_noop37:25 AMA timer for attempt 1 is delivered although the ask is resolved: noop, no LLMPath filtered. timer tmr:481207:ask:eta_request:ping_0700:1 is stale: the ask is closed or already past that step. No action. Model calls: 0.
duplicate_timer_delivery_is_noop17:01 AMDelay risk high: SMS the driver for the ETAPath sop. followed client_a/delay_risk_high@v1. Actions: send sms, timer set for 7:31 AM. Model calls: 0.
duplicate_timer_delivery_is_noop27:31 AMThe pump delivers the timer: second SMSPath sop. ask:eta_request:ping_0700 unanswered at step 1, continued client_a/delay_risk_high@v1. Actions: send sms, timer set for 8:01 AM. Model calls: 0.
duplicate_timer_delivery_is_noop37:31 AMThe same timer is delivered again: dropped as a duplicatePath filtered. event evt:tmr:481207:ask:eta_request:ping_0700:1 was already processed. No action. Model calls: 0.

Source: the scenario exports of timer_after_eta_received_is_noop and duplicate_timer_delivery_is_noop.

Film, 1:08 How an open ask lives

Open, follow up, escalate or resolve, and the timers that must do nothing, from three scenario runs.

Scene list (text version)
  1. 0:00 An ask is state projected from the log; timers only say that something may be due.
  2. 0:06 7:01 AM (client_a_load_481207_full_timeline): The SOP opens an ask: a row of state with a target, an attempt count, a timer and its SOP version.
  3. 0:15 7:13 AM (client_a_load_481207_full_timeline): An unrelated question gets its answer. No ask transition: the ask, its attempt and its timer stay as they were.
  4. 0:22 7:31 AM (client_a_load_481207_full_timeline): The timer is current (same attempt, ask still open), so the SOP takes the next step.
  5. 0:29 8:01 AM (client_a_load_481207_full_timeline): Attempts are used up: the escalation chain runs in one decision.
  6. 0:36 7:20 AM (timer_after_eta_received_is_noop): In another run the dispatcher emails the ETA. Any contact on any channel resolves the ask; its timer is cancelled.
  7. 0:45 7:25 AM (timer_after_eta_received_is_noop): The old timer is delivered anyway. State says the ask is closed: noop, no model call.
  8. 0:53 7:31 AM (duplicate_timer_delivery_is_noop): The same timer delivered twice: the event id is already in the log, the copy is dropped.
  9. 1:01 A timer acts only if its ask is open and on the same attempt; late, stale or duplicated timers are noops with no model call.

Generated from the scenario export at build time. Download MP4 (1.6 MB) with captions (WebVTT). Narration and sound effects generated with ElevenLabs.

1. How do you make sure that the system works correctly in production?

Section titled “1. How do you make sure that the system works correctly in production?”

Layers, from before release to after (design.md section 12):

  1. Before release: unit and contract tests, scenario evals, CI on every PR (ruff, format, mypy strict, pytest with a Postgres service).
  2. Validation at run time: every action is checked before commit (contact belongs to the load, channel supported, role not forbidden by the client’s SOPs for the whole load, idempotency key unused). Rejected actions become a Hero task. The agent can do harm only through actions that passed these checks.
  3. Decision records: every run records path, SOP version, actions, rationale and LLM usage, so any action is explainable and replayable.
  4. Metrics and alerts: path mix, escalation rate per client and topic, Hero override rate (primary quality metric), run latency against 5 minutes, outbox age, dead rows, timer pump lag, cost per load, triage confidence distribution, noop share.
  5. Human review loop: Heroes mark a task as “agent was wrong” or reverse an action; a weekly sample of decisions per client is reviewed; every confirmed error becomes a new scenario eval.
  6. Staged rollout: shadow mode for SOP and model changes (decide, do not act, compare), then a canary by client or by share of loads.

Alternative: LLM-as-judge on production traffic. Useful for tone, but structural checks and Hero overrides are cheaper and less ambiguous; we use a judge only where structure cannot decide. Built: decision records, validation, evals, CI. Designed: dashboards, alerts, review loop, shadow and canary.

2. A client changes its procedure. How do you make sure that other behavior does not break?

Section titled “2. A client changes its procedure. How do you make sure that other behavior does not break?”
  1. Scope is structural: a client file overrides one topic for one client. Other clients resolve to their own file or standard, so they cannot change.
  2. Suites run on every SOP change: the standard suite (it must stay green because the change must not leak) plus the client’s suite.
  3. Behavior diff: the old and new versions run on the same scenarios and the editor sees which decisions changed, with an explanation, not a text diff. An unexpected change in an unrelated scenario blocks publish. Today render_behavior_diff compares specs (built, shown by sop preview); the scenario-based diff is designed.
  4. Compile validation rejects ambiguous or contradictory text before it becomes a spec.
  5. Shadow mode on live loads for the client: compare decisions with the current version before publish.
  6. Version pinning: asks in flight finish under their version; only new triggers use the new one.
  7. One-click rollback republishes the previous spec.

Trade-off: a client with no scenarios has nothing to regress against, so onboarding includes writing its first scenarios (the 481207 file is the template), and a change to a topic with no coverage is flagged as unprotected (designed). Alternative: rely on shadow mode alone; slower feedback and only covers traffic that happens to occur.

3. You change the LLM. How do you make sure that the system does not break?

Section titled “3. You change the LLM. How do you make sure that the system does not break?”

Two kinds of tests, because they answer different questions:

  • Pipeline logic (scripted FakeLLM): the suite runs offline with the model’s outputs scripted per scenario. It proves the pipeline is correct for given model outputs and runs on every commit with no API key. An unscripted call fails the test, so “this step must not call a model” is enforced (llm_calls: 0). It does not test the model.
  • Model behavior (replay): a frozen set of real inbound messages with labels (human-checked or Hero-confirmed decisions from production) is run through the candidate model for each task. Compared per task: triage category and extracted value accuracy (ETA parsing with timezones is the sharp edge), confidence calibration, planner action agreement, document field accuracy, schema-validity rate, cost and latency. A change must meet or beat the current model on every gate; thresholds (PRESENT_MIN, GUARD_MAX) are re-fitted on the replay set, not copied.

Built for triage: tests/llm/test_gate.py replays recorded decision-model answers (two samples of 190 phrasings) through the real client and fails on any wrong value; python -m tests.llm.calibrate refits PRESENT_MIN and GUARD_MAX from those recordings; .github/workflows/live.yml runs the live smoke tests daily and on demand. The recorded phrasings were written by us, so this gate does not replace a production replay set. Production must also pin the calibrated model version (the default jev-latest floats, and an upgrade invalidates the thresholds).

Then shadow (candidate runs beside the current model on live loads, only the current one acts, disagreements are reviewed) and canary (a small share of loads or one client, with the Hero override rate and escalation rate as rollback triggers). Model ids are per task in LLMSettings, so each task can change on its own, and the decision record stores the model id of every call, so a regression is attributable.

Alternative: compare only final actions end to end. Simple, but a coincidence can hide a worse classifier. Trade-off: the replay set needs curation and refresh; it is the main cost of changing models safely. Built: the scripted suite and the typed failure handling (an invalid output becomes a Hero task, never a guess). Designed: the production replay set, shadow and canary.

Interactive The triage replay gate: wrong values first, then acceptance

Fit 70% 51 positives, 81 negatives

Sample 1 43/51 (84%) 0 wrong
Sample 2 44/51 (86%) 0 wrong

Held-out 30% 23 positives, 35 negatives

Sample 1 19/23 (83%) 0 wrong
Sample 2 17/23 (74%) 0 wrong

All 74 positives, 116 negatives

Sample 1 62/74 (84%) 0 wrong
Sample 2 61/74 (82%) 0 wrong

0 wrong values across both samples; a positive the gate is unsure of goes to a Hero, never to a guess. The two recordings give the same outcome on 180 of 190 phrasings.

Try the ETA rule

Value Accepted: the code's reading of the time resolves the ask.

The presets load the recorded answers of jev-1.13.0 (sample 1); the sliders are yours to move. The thresholds are the calibrated constants: main at least 0.60, guards at most 0.30 (ETA), 0.30 (temperature), 0.20 (seal).

Table view
SetSample 1: accepted, wrongSample 2: accepted, wrong
Fit 70% (51 positives, 81 negatives)43/51, 0 wrong44/51, 0 wrong
Held-out 30% (23 positives, 35 negatives)19/23, 0 wrong17/23, 0 wrong
All (74 positives, 116 negatives)62/74, 0 wrong61/74, 0 wrong

Source: tests/llm/calibrate.py over tests/llm/fixtures/jev_gate.json and jev_gate_sample2.json (190 phrasings, two recordings), thresholds from llm/typesafe.py.

4. Clients have different procedures and different complexity. How do you organize the evals?

Section titled “4. Clients have different procedures and different complexity. How do you organize the evals?”

By layer, mirroring the SOP layers, plus layers of test for the machinery:

LayerWhat it provesWhere
Unit and contract testsDomain rules, projection, store/outbox/timer contracts (SQLite and Postgres run the same suite), validation, templates, compiler rulestests/
Standard scenariosBehavior every client gets unless it overrides the topicdesigned layout evals/standard/
Client scenariosOverrides: Client A escalation chain, Client B Slack alert and “never the dispatcher”, Client B reefer temperature and sealdesigned layout evals/clients/<id>/
Cross-cutting scenariosNoop cases, cross-channel resolution, duplicates, stale pings, unknown situationsscenario files
Replay of production decisionsModel changes and driftdesigned

A client change runs standard plus that client. A change to a standard topic runs every client that does not override it (derivable from the SOP tree). Each scenario is a YAML timeline replayed with FakeClock through the real worker, dispatcher and timer pump; each step asserts the exact set of actions (kind, contact, channel, ask, attempt, key) or an explicit noop, and optionally path, number of LLM calls, state and whether the log grew. Message content is checked structurally (the PO number is present). A judge is used only where structure cannot decide.

Today the scenarios sit in evals/scenarios/, named by client where relevant; the split into standard/ and clients/<id>/ is a file move plus a runner that selects by SOP tree, not a redesign. Complexity scales by adding scenarios per client, not by editing a shared test: a complex client gets many files, a simple one gets the standard suite for free. Alternative: one parameterized suite across all clients; compact, but a failure does not say which client’s procedure broke and per-client edge cases get lost.

Do-nothing coverage: ok_with_no_open_ask_is_noop, ok_with_open_eta_ask_keeps_ask_open, timer_after_eta_received_is_noop, duplicate_timer_delivery_is_noop, stale_ping_is_ignored, high_ping_while_already_high_is_not_a_trigger.

Interactive Scenario evals by client and tag

30scenario files, every step asserted

15include a "do nothing" step

1503tests: 1402 offline, 101 against Postgres in CI

Show
Scenario files and their tags. A filled mark means the scenario has the tag.
Scenario Client Do nothing 15Timers 9Cross-channel 2Idempotent 1Escalation 11Model call 15
Broker stand down cancels the ETA chain Client A no no no
Driver arrives before follow up cancels ETA ask Client A no no no no
Duplicate timer delivery is noop Client A no no no
High ping while already high is not a trigger Client A no no no no no
Ok with no open ask is noop Client A no no no no no
Ok with open ETA ask keeps ask open Client A no no no no no
Repeated answer within ten minutes is sent once Client A no no no no
Risk falls back cancels the ETA chain Client A no no no no
Risk flapping through low keeps the chain on schedule Client A no no no
Risk flapping through medium keeps the chain on schedule Client A no no no
Stale ping is ignored Client A no no no no no
Thumbs up with no open ask is noop Client A no no no no no
Thumbs up with open ETA ask keeps ask open Client A no no no no no
Timer after ETA received is noop Client A no no
Unknown sender gets one hero task per half hour Client A no no no no
Address question names the stop it is about Client A no no no no
Client A load 481207 full timeline Client A no no no
Client A PO from instructions Client A no no no no no
Delivery risk high while pickup ETA ask is open asks for the delivery ETA Client A no no no no no
Departing without arriving closes the pickup ask Client A no no no no no
Dispatcher email ETA resolves driver ask Client A no no no no
Email outside the main thread is escalated not answered Client A no no no no
Eta reply resolves only the nearest stop ask Client A no no no no no
Named stop ETA resolves only that stops ask Client A no no no no no
Po question without PO number becomes hero task Client A no no no no
Unknown situation escalates to hero task Client A no no no no
Unrelated driver question is answered and ask stays open Client A no no no no no
Client B delay risk sms then slack never dispatcher Client B no no no no
Client B dispatcher PO email goes to a hero with or without an open ask Client B no no no no
Client B reefer departed pickup asks temperature and seal Client B, reefer no no no no no no

Source: the scenario export index (30 files in evals/scenarios/, all replayed and matching in this build) and the default pytest run.

1. How does your design keep each agent run at 5 minutes or less?

Section titled “1. How does your design keep each agent run at 5 minutes or less?”

Most runs are milliseconds because most events never reach a model (design.md section 13).

StepTypicalBound
Queue and ingestunder 1 sFIFO group per load, no cross-load blocking
Read and projectunder 50 msabout 1,300 rows per load with pings; snapshot row every 200 records beyond about 2,000 rows (designed)
Filter, trigger diff, SOP executor, templatesmillisecondspure code
Triage call (only if needed)about 0.2 s (measured, ADR 0005)30 s timeout, 2 retries, capped by the run deadline (LLMSettings)
Planner call (rare)3 to 10 s30 s timeout, 2 retries, capped by the run deadline
Commitunder 20 msoptimistic concurrency, up to 3 re-decides (max_conflict_retries)
DispatchsecondsRetryPolicy.budget (150 s) from the first send; permanent failure escalates
Timer pumpunder 1 minute after fire_atpoll every 30 s; alert on lag over 2 minutes

Why it holds: serial per load with optimistic concurrency means no waiting on locks; outbox rows are claimed by many dispatchers in parallel; the pump scales by partition. Failure paths are bounded: an LLM call gets at most 2 retries inside the client and then becomes a Hero task, and a failed send is a failed outcome the agent escalates.

Bound on the model path (built, agent/deadline.py, RetryPolicy.budget): a 120 s run deadline spans all conflict re-decides. Once it is spent, no model call starts and the run commits an urgent Hero task. An in-flight HTTP attempt is capped at the time left. Event to last action is at most 270 s, including the 150 s dispatch budget. Residual: httpx timeouts are per phase (connect, write, read), so the bound assumes small non-streamed responses. Alternative: per-call timeouts alone, which bound one call but not the sum of a triage call, a planner call and the conflict re-decides.

Interactive Event to last action: at most 270 s against 5 minutes
Axis
Bound
Run deadline 120 s Dispatch 150 s 30 s
Measured
p50, all runs: 17 ms p95, all runs: 1.56 s p95, runs with a model call: 7.01 s slowest (bench:1:053): 21.2 s

The slowest run, 21.2 s (triage 192 ms + plan 21.0 s), used 18% of the run deadline. Once the deadline is spent no model call starts and the run commits an urgent Hero task, so a slow provider costs a Hero task, not the 5 minutes.

Table view
PartSeconds
Run deadline (all model calls and conflict re-decides)120
Dispatch retry budget, from the first send150
Bound, event to last action270 of 300
Measured p50, all runs0.02
Measured p95, all runs1.56
Measured p95, runs with a model call7.01
Measured slowest (bench:1:053)21.21

Source: agent/deadline.py (RUN_DEADLINE), agent/dispatcher.py (RetryPolicy.budget), bench/results/2026-10-08-typesafe-deepseek.json (91 runs).

2. The volume increases ten times. What must change?

Section titled “2. The volume increases ten times. What must change?”

Today’s design has no per-process state, so the worker tier scales by adding workers; what changes is the data and external limits (design.md section 13):

PressureChange
Postgres: about 500 inserts per second average and 1,500 at peak with pings in the log, 14B log rows per yearPartition log and outbox by load_id hash, archive closed loads to object storage after 30 days, read replicas for dashboards, shard by client group if one cluster saturates
QueueSQS FIFO with many message groups; Kafka keyed by load if per-group limits or cost bite (trade-offs in ADR 0003)
Timer pump and dispatcherMore instances, partitioned by fire_at bucket and hash, SKIP LOCKED; per-service rate limits and circuit breakers
LLM provider limitsReserved throughput, a fallback provider, a concurrency cap on the queue side
Pings (the largest input, 45M per month at 1x on the assumption of one per 15 minutes)Filter at ingestion against a cached band per stop, or keep only band-changing pings in the log and the rest in a telemetry table (designed)
CostTriage is the largest line and already runs on the decision model. Next levers: debounce consecutive SMS into one call, and a distilled in-house classifier validated by the recorded-answer gate (both designed)
HumansEscalations scale with volume; watch tasks per load and fix SOPs where Heroes repeat the same manual step

Alternative: move now to Kafka plus a separate state store. It scales writes further but splits the event, decision and outbox commit across two systems, which ADR 0003 rejects until Postgres partitioning saturates (it says to reconsider Kafka if write throughput becomes the bottleneck). Trade-off of ours: a bigger Postgres operation (partitions, archive, replicas) in exchange for keeping one transaction.

What does not change: the contracts. EventStore, Outbox, TimerStore, LoadQueue and LLMClient are protocols, so a different queue or store plugs in behind them.

Interactive Ten times the volume: what moves
Volume
Loads per month
100K
Inbound messages per month
5M to 10M
Model calls per month
4.4M to 8.8M
Model cost per month
$9.8K to $19.6K
  • Workers Same at 10x No per-process state: add workers. Unchanged contracts (EventStore, Outbox, TimerStore, LoadQueue, LLMClient).
  • Postgres No change at 1x
  • Queue No change at 1x
  • Timer pump and dispatcher No change at 1x
  • LLM provider limits No change at 1x
  • Pings No change at 1x
  • Cost No change at 1x
  • Humans No change at 1x
Table view

The changes are the table in the answer above. Volume at 1x: loads per month 100K; inbound messages per month 5M to 10M; model calls per month 4.4M to 8.8M; model cost per month $9.8K to $19.6K. At 10x: loads per month 1M; inbound messages per month 50M to 100M; model calls per month 43.8M to 87.5M; model cost per month $98.2K to $196.5K.

Source: volume from docs/assignment.md; 0.88 model calls and $0.0020 per inbound message measured in bench/results/2026-10-08-typesafe-deepseek.json (56 messages); changes from the table above.

3. How do you keep the LLM cost at approximately $1 or less for each load? You do not have to implement this.

Section titled “3. How do you keep the LLM cost at approximately $1 or less for each load? You do not have to implement this.”

A funnel in which the model is the last resort (ADR 0001, ADR 0004; arithmetic in design.md section 13).

Volume: 100,000 loads, 50 to 100 messages each, 5M to 10M messages per month. $1 over 75 messages is about $0.013 per message. A model call on every message at that price would spend the whole budget on triage and leave nothing for the planner and documents, and any call above that price breaks it. The funnel:

StageEventsCost
Deterministic filter: duplicates, acknowledgements with no open ask, unknown senders (assumed 25% to 40% of inbound messages; basis: driver SMS is short and often “ok” or “thanks”, to be measured from the path mix)25% to 40% of messages$0
Timers that are stale or resolved, delivery outcomes, pings with no band change (system events, not inbound messages)all of them$0
SOP executor: trigger fired, timer continues an askthe procedural core$0
Outbound messages: templates filled in code40 to 70 per load$0
Triage: decision model (TypeSafe Jev), cached system prompt, short request (one message plus open asks)30 to 75 calls$0.0005 to $0.003 each, so $0.015 to $0.23
Planner: strong model, no carrier-facing actions (a Hero task, a TMS note, or a Slack post to the broker)0 to 2$0.02 to $0.05 each, so up to $0.10
Documents: vision2 to 6$0.01 to $0.02 each, so $0.02 to $0.12
Template-less drafts0 to 1up to $0.01
SOP compile, amortized (100 clients times 20 saves per month at $0.10)per edit, not per eventabout $0.002 per load

Total about $0.04 to $0.46 per load, which is 2x to 25x of headroom under $1. The prices are assumptions to replace with contract prices; the structure is what matters. Load 481207 makes one model call in the whole timeline.

Guards, so the average stays low: prompt caching for the system prompt and SOP guidance; a per-load cost counter from the LLMUsage records, with a warning at $0.60 and a ceiling at $2 per load (twice the budget, so only a runaway load such as a chatty driver or a looping SOP reaches it) that disables the planner and escalates to Hero tasks instead (designed); debounce of consecutive SMS into one triage call (designed); path-mix monitoring, since a rising planner share is the early signal of missing SOP coverage; a regression gate in evals (llm_calls per step), so a change that adds hidden calls fails CI.

Alternatives: a cheaper model for everything (accuracy drops on the language-heavy cases and the cost of a wrong action exceeds the saving); an LLM draft for every message (adds a call per outbound message and nondeterministic text, and verification catches a lost fact but not an invented claim); caching responses by message text (works for acknowledgements and is covered by the filter, but unsafe for messages whose meaning depends on open asks). Trade-off of ours: templates are less adaptive than model text, and the planner path is cheap only while SOP coverage stays good.

Interactive Measured: $0.110 of model cost per load against $1.00
  • Triage on jev-1.13.0: $0.090, 45 calls, at an assumed $0.002 per call
  • Planner on deepseek-v4-pro: $0.020, 4 calls, measured

9.1x headroom under the budget. The decision model's price is an assumption: the cost reaches $1.00 at $0.0218 per triage call. Change the price and the volume in the calculator.

Table view
PurposeModelCallsCost per loadBasis
Triagejev-1.13.045$0.090assumed price
Plannerdeepseek-v4-pro4$0.020measured
Total49$0.110budget $1.00

Source: bench/results/2026-10-08-typesafe-deepseek.json, one load with 56 inbound messages. Benchmark and cost calculator.

Prepared for Freight Hero by Marcus Caum Source on GitHub