Answers to the assignment questions
Same order and headings as assignment.md. The full system is in design.md; decisions are in docs/adr/. “Built” means it is in the code and tested; “designed” means described only.
The 5 steps of the assignment timeline as the agent processed them, with the ETA ask beside each one.
The player could not load. The rendered MP4 plays below, with the same captions.
Scene list (text version)
- 0:00 Load 481207: Client A, Riverbend Foods, dry van, Riverbend plant to Valley DC.
- 0:12 7:01 AM: Ping turns delay risk high: SMS the driver for the ETA and set the 7:31 AM follow-up. Actions: send sms "Load 481207: what is your ETA to Riverbend plant? Please reply with the time."; timer for 7:31 AM.
- 0:21 7:13 AM: Driver asks for the PO: answer it from metadata, the ETA ask stays open. Actions: send sms "Load 481207: the PO number is 55821.".
- 0:32 7:31 AM: Timer: no ETA yet, ask again and set the 8:01 AM follow-up. Actions: send sms "Load 481207: what is your ETA to Riverbend plant? Please reply with the time."; timer for 8:01 AM.
- 0:40 8:01 AM: Timer: two requests unanswered, email the dispatcher in the main thread and create an urgent Hero task. Actions: send email "The driver on load 481207 has not answered our texts. Please send the ETA to Riverbend plant."; create hero task "Load 481207: Call driver for eta: Call the driver on load 481207 to get the ETA to Riverbend plant. No ETA arrived after our requests.".
- 0:49 8:30 AM: Nothing is left to do: the escalation chain is finished.
- 0:55 1 model call in 5 steps; every step is asserted in client_a_load_481207_full_timeline.
Generated from the scenario export at build time. Download MP4 (1.4 MB) with captions (WebVTT). Narration and sound effects generated with ElevenLabs.
Agent architecture
Section titled “Agent architecture”1. What is the architecture of your agent? What are the alternatives?
Section titled “1. What is the architecture of your agent? What are the alternatives?”A hybrid, event-driven pipeline, serial per load, where code owns state and procedure and the LLM is the exception (ADR 0001, design.md sections 3 and 4).
- Every input (SMS, email, Slack, ping, timer, attachment, stop status, delivery outcome) and every decision is appended to a per-load log. State is a projection of that log, including open asks.
- Per event: deterministic filter, trigger diff, SOP routing, triage (a model call only for a message that survives the filter), SOP executor or planner, action validation, then one transaction that commits the event, the decision record and the outbox rows. A dispatcher sends from the outbox with idempotency keys. Timers are durable and re-enter through the same queue.
- The SOP is compiled at edit time into a state machine (ask, retries, wait, escalation chain). The runtime executes it without a model. The principle: the model answers typed questions about meaning, code owns values and actions. The built triage model is the TypeSafe Jev decision model (ADR 0005), which returns probabilities for closed questions and never text or values. Code finds and normalizes value candidates (
llm/extract.py) and the model only confirms one. A DeepSeek planner proposes actions for situations no SOP covers, with no carrier-facing actions (a Hero task, a TMS note, or a Slack post to the broker).
| Alternative | Trade-off |
|---|---|
| Single ReAct agent with tools, full history per event | Simplest and most flexible. Cost grows with history on every event, retry counting rests on model recall, and step-by-step testing is hard. Does not fit $1 per load at 50 to 100 messages plus pings. |
| Workflow engine with no LLM | Predictable and cheap. Cannot read “What’s the PO number?” or draft text. We kept its good parts (durable timers, retries, idempotency) and chose Postgres over Temporal so the decision and its actions share one transaction (ADR 0003). |
| Multi-agent | More hand-offs for a problem whose core is one state machine per load. Cross-channel resolution needs shared state anyway. |
| Chat history or vector memory | Probabilistic and not replayable. |
| A generative model that reads each message and returns the value | One call does everything, but it can invent a time and its confidence is uncalibrated. We use a decision model for meaning and code for values (ADR 0005). |
Costs of ours: upfront modeling (events, asks, transitions), coverage depends on the SOP corpus (uncovered cases go to the planner or a Hero), and the outbox adds a dispatcher and a few seconds of delivery latency. Code: src/load_agent/agent/, domain/, store/.
- Eventappended to the load log; the run projects the state from it
- Filter Pings with no band change, stale pings, acknowledgements with no open ask code only, $0
- Trigger and SOP A risk band change, a stage change or a timer runs a compiled SOP code only, $0
- Triage The decision model reads one message against the open asks model call
- Hero on doubt Unknown sender, or triage not sure enough: Hero task 7 of 9 called a model
- Planner No SOP covers it: a bounded plan, never carrier-facing model call
- Validate, commit, dispatchevery action: contact on the load, channel supported, role allowed, idempotency key unused; then one transaction with the outbox
46 of 91 runs (50.5%) never called a model. Of the 87 inputs, 45 (51.7%) reached one; no timer did.
Table view
| Stage where the run was decided | Runs | Share | Runs with a model call |
|---|---|---|---|
| Filter | 36 | 39.6% | 0 of 36 |
| Trigger and SOP | 8 | 8.8% | 0 of 8 |
| Triage | 35 | 38.5% | 35 of 35 |
| Hero on doubt | 9 | 9.9% | 7 of 9 |
| Planner | 3 | 3.3% | 3 of 3 |
| Total | 91 | 100% | 45 of 91 |
Source: bench/results/2026-10-08-typesafe-deepseek.json (87 inputs and 4 timers, seed 1). Full architecture.
Four recorded answers of the decision model, and why each became a value or a Hero task.
The player could not load. The rendered MP4 plays below, with the same captions.
Scene list (text version)
- 0:00 Code finds candidate values; the decision model answers yes or no questions with probabilities; code applies the thresholds; unsure goes to a Hero.
- 0:10 "ETA 9:30am": main 0.82 (at least 0.60), highest guard 0.08 (at most 0.30): accepted, ETA 09:30.
- 0:22 "9:30am": main 0.59 (at least 0.60), highest guard 0.28 (at most 0.30): Hero task.
- 0:29 "tow eta 9:30": main 0.74 (at least 0.60), highest guard 0.68 (at most 0.30): Hero task.
- 0:43 "eta was 9:30": main 0.63 (at least 0.60), highest guard 0.90 (at most 0.30): Hero task.
- 0:54 0 wrong values in 190 phrasings over two recordings; 62/74 and 61/74 positives accepted.
Generated from the scenario export at build time. Download MP4 (1.3 MB) with captions (WebVTT). Narration and sound effects generated with ElevenLabs.
2. How does the agent read documents, find their type and get their key values?
Section titled “2. How does the agent read documents, find their type and get their key values?”Classify, extract with a schema per type, cross-validate in code, escalate when unsure (design.md section 10).
- An
AttachmentReceivedevent carries the file URI, media type and sender. - One vision-model call (
LLMClient.extract_document) returns the document type (bol,pod,rate_confirmation,receipt,unknown), the key fields and a confidence. - Fields are per type: BOL (BOL number, PO number, seal number, pieces, weight, signature), POD (receiver, delivery time, signature, exceptions), rate confirmation (rate, stops, appointments, MC number, load id), receipt (amount, vendor, date, kind).
- Cross-validation is code, not a model: the extracted PO and BOL numbers are compared with
LoadMetadata.reference_numbers. Designed additions: weight, load id, MC number, stop names, rate against the TMS. Photo downscaling and duplicate-file dropping before the call are also designed. - Unknown type, confidence under 0.7, or any mismatch becomes a Hero task naming the field and both values, and no SOP step runs. A trusted document fires a
DocumentReceivedtrigger, so the client’s SOP for that type decides (for example a TMS note that the BOL arrived).
Alternatives: OCR plus rules (cheap, breaks on photos and layout changes); a single prompt asking the model whether the document matches the load (the model can agree with a wrong document; an equality check cannot); auto-correct the load from the document (a misread digit would silently change a contract value). Trade-off: low-quality phone photos fall below the confidence gate and cost Hero time, never silent errors. Built: the seam, schemas, the 0.7 gate, PO and BOL cross-check, escalation (agent/reactions.py). Designed: the real vision client, per-type schemas, the wider cross-check set.
- Attachment received An AttachmentReceived event carries the file, media type and sender. Built
- Find the type One vision call (LLMClient.extract_document) returns the type: BOL, POD, rate confirmation, receipt or unknown. Built seam, designed client
- Extract fields A schema per type: BOL and PO numbers, seal, pieces, weight; POD receiver, time, exceptions; and so on. Designed
- Cross-check in code The PO and BOL numbers must equal the load reference numbers. An equality check, not a model opinion. Built
- Gate Unknown type, confidence under 0.7 or any mismatch: Hero task naming the field and both values. No SOP step runs. Built
- Trusted: route A DocumentReceived trigger routes to the client SOP for that type. Built
SOP Runs BOL received (standard/bol_received@v1)
Table view
| Document | Trusted (confidence at least 0.7, numbers match) | Confidence under 0.7 | PO or BOL mismatch |
|---|---|---|---|
| Bill of lading | Runs BOL received (standard/bol_received@v1) | Hero task: the document could not be identified with confidence | Hero task naming the field and both values |
| Proof of delivery | Runs POD received (standard/pod_received@v1) | Hero task: the document could not be identified with confidence | Hero task naming the field and both values |
| Rate confirmation | No SOP for this type: recorded, nothing else to do | Hero task: the document could not be identified with confidence | Hero task naming the field and both values |
| Unknown type | Hero task: the type could not be identified | Hero task: the document could not be identified with confidence | Hero task naming the field and both values |
Source: agent/reactions.py (gate 0.7), routing from sops/ for Client A. Built means in the code and tested.
Instructions
Section titled “Instructions”1. How do you organize the SOPs of many clients? The standard procedure applies to all clients until a client changes it.
Section titled “1. How do you organize the SOPs of many clients? The standard procedure applies to all clients until a client changes it.”Layered files per topic, most specific wins (ADR 0002):
sops/standard/delay_risk_high.mdsops/clients/client_a/delay_risk_high.md # overrides standard for client_asops/clients/client_b/reefer/departed_pickup.md # client + load typeOrder: client + load_type over client over standard. The winning file replaces the topic as a whole. A new client starts with zero files and gets the standard for every topic (the 60 to 70% shared part); it overrides only what differs. Each Markdown file has a compiled, versioned sibling (delay_risk_high.v1.compiled.yaml), and templates.v1.yaml files layer the same way for wording.
Alternative: section-level merge, where a client patches part of a standard topic. Smaller client files, but editors cannot predict “which step 2 applies”. Whole-topic override is easy to preview and diff; the cost is duplicated text when a client changes one line. The compiler can flag drift from standard as a review hint (designed). Built: FileSopRepository (sop/repository.py).
- Client + load type
...... - Client
...... - Standard
......
Client + load type client_b/reefer/departed_pickup@v1
When the load enters the stage 'to delivery' from 'at pickup':
1. Write a TMS note (record that the driver left the pickup).
2. Text the driver (ask for the trailer temperature and the seal number). Wait for the trailer temperature and seal number.
The request stays open. There is no follow-up.
The request is answered as soon as we receive the trailer temperature and seal number from any contact of the load on any channel. Behavior change compared with standard/departed_pickup@v1:
- Now sends a request and waits for an answer; it did not before. - May contact the driver again. - May contact the dispatcher again. - Guidance (affects wording only, not what happens): added 'Wording'.
Table view
| Layer | Client | Load type | Topic | File |
|---|---|---|---|---|
| Client | Client A | any | Delay risk high | sops/clients/client_a/delay_risk_high.md |
| Client | Client B | any | Delay risk high | sops/clients/client_b/delay_risk_high.md |
| Client + load type | Client B | reefer | Departed pickup | sops/clients/client_b/reefer/departed_pickup.md |
Source: the sops/ tree (11 compiled files) resolved by FileSopRepository; text from load-agent sop preview.
2. How does the agent find the correct SOP files for the current event? What are the advantages and disadvantages of your method?
Section titled “2. How does the agent find the correct SOP files for the current event? What are the advantages and disadvantages of your method?”Deterministic routing, no retrieval on the load path. The event becomes a trigger by diffing projected state: DelayRiskBecame(high, stop), StageEntered(to_delivery, from at_pickup), DocumentReceived(bol). The compiler builds a routing index from triggers to topics. SopRepository.topics_for(trigger, client_id, load_type) is a lookup that returns the winning layer. If the winning spec has conditions (for example no_open_ask: eta_request), the executor checks them against state. Messages are not routed to SOPs: they are matched to open asks, which already carry the SOP version they run under.
For a situation with no topic the pipeline goes to triage, then the planner (no carrier-facing actions (a Hero task, a TMS note, or a Slack post to the broker), at most 3 steps, confidence at least 0.8), and the default is a Hero task.
| Advantages | Disadvantages |
|---|---|
| Zero cost and milliseconds; testable per trigger | A new situation needs a new topic and a compile |
| A client override can never be shadowed by a similar chunk from another layer | Fallback coverage depends on the planner and on Hero capacity |
| The same input always routes the same way, so evals are exact | The closed vocabulary of trigger kinds needs a code change to extend |
Alternative: embed SOP chunks and retrieve per event. It handles unseen phrasing, but it is nondeterministic, can mix layers, and costs an embedding plus a longer prompt per event. We keep it only as a possible fallback to suggest a topic to a Hero. Built: sop/routing.py, agent/triggers.py.
- Trigger from the state diff
... - Routing index lookup...
- DecisionSOP ...
An inbound message is never routed to an SOP: it is matched to the open asks, which already carry the SOP version they run under.
Table view
| Event | Client A, dry van | Client A, reefer | Client B, dry van | Client B, reefer | Any other client, dry van | Any other client, reefer |
|---|---|---|---|---|---|---|
| Delay risk for the pickup becomes high | Delay risk high, client layer (client_a/delay_risk_high@v1) | Delay risk high, client layer (client_a/delay_risk_high@v1) | Delay risk high, client layer (client_b/delay_risk_high@v1) | Delay risk high, client layer (client_b/delay_risk_high@v1) | Delay risk high, standard layer (standard/delay_risk_high@v1) | Delay risk high, standard layer (standard/delay_risk_high@v1) |
| Delay risk for the pickup becomes medium | No SOP: noop | No SOP: noop | No SOP: noop | No SOP: noop | No SOP: noop | No SOP: noop |
| Driver arrives at the pickup | Arrived at pickup, standard layer (standard/arrived_at_pickup@v1) | Arrived at pickup, standard layer (standard/arrived_at_pickup@v1) | Arrived at pickup, standard layer (standard/arrived_at_pickup@v1) | Arrived at pickup, standard layer (standard/arrived_at_pickup@v1) | Arrived at pickup, standard layer (standard/arrived_at_pickup@v1) | Arrived at pickup, standard layer (standard/arrived_at_pickup@v1) |
| Driver departs the pickup | Departed pickup, standard layer (standard/departed_pickup@v1) | Departed pickup, standard layer (standard/departed_pickup@v1) | Departed pickup, standard layer (standard/departed_pickup@v1) | Departed pickup, client + load type layer (client_b/reefer/departed_pickup@v1) | Departed pickup, standard layer (standard/departed_pickup@v1) | Departed pickup, standard layer (standard/departed_pickup@v1) |
| Driver arrives at the delivery | Arrived at delivery, standard layer (standard/arrived_at_delivery@v1) | Arrived at delivery, standard layer (standard/arrived_at_delivery@v1) | Arrived at delivery, standard layer (standard/arrived_at_delivery@v1) | Arrived at delivery, standard layer (standard/arrived_at_delivery@v1) | Arrived at delivery, standard layer (standard/arrived_at_delivery@v1) | Arrived at delivery, standard layer (standard/arrived_at_delivery@v1) |
| Load marked delivered | No SOP: noop | No SOP: noop | No SOP: noop | No SOP: noop | No SOP: noop | No SOP: noop |
| A bill of lading is received | BOL received, standard layer (standard/bol_received@v1) | BOL received, standard layer (standard/bol_received@v1) | BOL received, standard layer (standard/bol_received@v1) | BOL received, standard layer (standard/bol_received@v1) | BOL received, standard layer (standard/bol_received@v1) | BOL received, standard layer (standard/bol_received@v1) |
| A proof of delivery is received | POD received, standard layer (standard/pod_received@v1) | POD received, standard layer (standard/pod_received@v1) | POD received, standard layer (standard/pod_received@v1) | POD received, standard layer (standard/pod_received@v1) | POD received, standard layer (standard/pod_received@v1) | POD received, standard layer (standard/pod_received@v1) |
| A rate confirmation is received | No SOP: noop | No SOP: noop | No SOP: noop | No SOP: noop | No SOP: noop | No SOP: noop |
| A document of unknown type arrives | Hero task | Hero task | Hero task | Hero task | Hero task | Hero task |
Source: route_trigger in sop/routing.py over the sops/ tree, for each client and load type.
3. People without a technical background write and change the SOPs. How does your design make this safe? You do not have to build an editor or a platform. Describe the design.
Section titled “3. People without a technical background write and change the SOPs. How does your design make this safe? You do not have to build an editor or a platform. Describe the design.”The editor writes plain English and approves what the agent understood, not YAML, behind a pipeline of gates (design.md section 8, ADR 0002):
- Compile on save. A strong model extracts the procedure; code validates it. Ambiguity is rejected with a reason the editor can act on (“if the driver doesn’t answer soon” is
ambiguous_step), as are unknown roles, a step that contacts a role the same SOP forbids, and waits or counts not grounded in the text. - Plain-English preview, rendered by code from the compiled spec, not written by a model. What the editor approves is exactly what runs. Real output is in design.md section 8 (
uv run load-agent sop preview client_a delay_risk_high). - Behavior diff (designed on scenario suites; built between specs), not a text diff: which decisions change on the scenario suites (“at 08:01 the agent now posts to Slack instead of emailing the dispatcher”).
- Eval gates: standard suite plus this client’s suite must pass.
- Shadow mode on live loads: the new version decides, does not act, and differences with the current version are reviewed.
- Publish creates an immutable version and refuses a stale version number when two editors publish at once. Rollback republishes an old spec as the newest version.
- Version pinning: an ask started under v1 finishes under v1, so an edit cannot change a wait or an escalation target in the middle of a chain.
- Guidance that is not procedure (“keep texts short”) stays text and cannot change what happens. Message wording lives in templates, versioned separately.
Alternatives: editors write YAML or a DSL (exact, but not for non-technical people); the model reads raw Markdown at runtime (no compile step, but behavior shifts with phrasing and model version, and nothing can be diffed). Trade-off: the compile step adds latency to an edit (seconds) and a closed vocabulary; the compile heuristics fail safe by rejecting. Built: preview, behavior diff, validation rules, stale-source detection. Designed: the editor, scenario-based diff, shadow mode, the publish service.
Designed: editorBuilt: Markdown files The editor writes plain English. No YAML, no DSL.
sops/clients/client_a/delay_risk_high.md
1. Send an SMS to the driver to get the ETA. 2. If there is no ETA after 30 minutes, send the SMS again. 3. If there is still no ETA 30 minutes after the second SMS, send an email to the dispatcher in the main email thread. 4. Then create an urgent Hero task for a Hero to call the driver.
Built A strong model extracts the procedure, code validates it. Rejections the editor can act on: ambiguous_step, unknown_contact_role, forbidden_role_used, unsupported_action, missing_trigger, conflicting_steps, forbidden_role_not_declared, schema_invalid.
sops/clients/client_a/delay_risk_high.v1.compiled.yaml
ask:
type: eta_request
target_role: driver
request:
action: send_sms
intent: request_eta
target_role: driver
channel: sms
urgent: false
wait_after: PT0S
max_attempts: 2
retry_after: PT30M
resolution:
requires:
- eta Built Rendered by code from the compiled spec, not by a model: what the editor approves is what runs.
load-agent sop preview client_a delay_risk_high
When delay risk becomes high:
Only if no ETA request is already open.
1. Text the driver (ask for the ETA). Wait for the ETA.
2. If there is no ETA within 30 minutes, repeat step 1 (attempt 2 of 2).
3. If there is still no ETA 30 minutes after the last attempt, in order:
a. Email the dispatcher in the main thread (say the driver does not answer and ask for the ETA).
b. Create an urgent Hero task (call the driver to get the ETA).
The request is answered as soon as we receive the ETA from any contact of the load on any channel. The remaining steps are then skipped. Built: between specsDesigned: on scenario suites What changes compared with standard/delay_risk_high@v1, in behavior, not text.
render_behavior_diff
- Attempts at the request: 2 (was 1). - Escalation: now also does: email the dispatcher in the main thread (say the driver does not answer and ask for the ETA). - Escalation: the Hero task is now urgent. - Guidance (affects wording only, not what happens): added 'Email to the dispatcher'; changed 'Tone'.
Built: all scenarios run in CI on every commitDesigned: per-client suite selection, publish gate The standard suite plus this client's suite must pass. Today: 30 scenario files (27 for Client A) and 1503 tests in CI.
evals/scenarios/
address_question_names_the_stop_it_is_about broker_stand_down_cancels_the_eta_chain client_a_load_481207_full_timeline client_a_po_from_instructions delivery_risk_high_while_pickup_eta_ask_is_open_asks_for_the_delivery_eta departing_without_arriving_closes_the_pickup_ask ... 21 more
Designed On live loads the new version decides but does not act; differences with the current version are reviewed.
Designed: publish serviceBuilt: version pinning Publish creates an immutable version and refuses a stale version number. An ask started under client_a/delay_risk_high@v1 finishes under it. Rollback republishes an old spec as the newest.
Table view
- Edit (Designed: editor; Built: Markdown files): The editor writes plain English. No YAML, no DSL.
- Compile (Built): A strong model extracts the procedure, code validates it. Rejections the editor can act on: ambiguous_step, unknown_contact_role, forbidden_role_used, unsupported_action, missing_trigger, conflicting_steps, forbidden_role_not_declared, schema_invalid.
- Preview (Built): Rendered by code from the compiled spec, not by a model: what the editor approves is what runs.
- Behavior diff (Built: between specs; Designed: on scenario suites): What changes compared with standard/delay_risk_high@v1, in behavior, not text.
- Eval gates (Built: all scenarios run in CI on every commit; Designed: per-client suite selection, publish gate): The standard suite plus this client's suite must pass. Today: 30 scenario files (27 for Client A) and 1503 tests in CI.
- Shadow (Designed): On live loads the new version decides but does not act; differences with the current version are reviewed.
- Publish and pin (Designed: publish service; Built: version pinning): Publish creates an immutable version and refuses a stale version number. An ask started under client_a/delay_risk_high@v1 finishes under it. Rollback republishes an old spec as the newest.
Source: sops/clients/client_a/delay_risk_high.md, sops/clients/client_a/delay_risk_high.v1.compiled.yaml, sop/preview.py, evals/scenarios/.
Memory
Section titled “Memory”1. How does the agent know its past actions and the past messages of the load?
Section titled “1. How does the agent know its past actions and the past messages of the load?”From the per-load event log; state is a projection of it (ADR 0001). Inbound events, DecisionRecorded entries (which hold the actions taken, the SOP version, ask transitions and the rationale) and delivery outcomes are in one ordered log. project_state(records) folds it into LoadState: stage, delay risk per stop, open asks, processed event ids, committed action keys. The agent reads the log at the start of each run and never relies on a model’s recall of earlier turns.
The model sees only what a call needs: the one message, the open asks (id, type, what they await, target role), the stop timezone. Not the history.
Alternatives: chat history in the prompt (cost grows, recall errors), a vector store of messages (probabilistic), a mutable state row (no audit, no replay). Trade-off: modeling effort, and stored models must remain readable across schema changes. Built: store/, domain/.
2. At 07:13, how does the agent remember that the ETA request is still open?
Section titled “2. At 07:13, how does the agent remember that the ETA request is still open?”It is a row of state, not a memory. At 07:01 the decision for ping ping_0700 committed an AskOpened transition for ask:eta_request:ping_0700 with attempts: 1, target drv_20931, next_timer: 07:31, resolved_by: null, pinned to client_a/delay_risk_high@v1. At 07:12 the driver asks “What’s the PO number?”. The run projects the log, sees one open ask awaiting an ETA, and triage returns “question about the PO number, no ETA”. satisfies(resolution, extracted, ...) is false because no ETA was extracted, so nothing resolves the ask. The agent answers “Load 481207: the PO number is 55821.” from reference_numbers.po_number (parsed from instructions at registration when the field is absent; see design.md, Load metadata; with no PO found it escalates to a Hero and never invents one) and commits no ask transition. The ask, its attempt count and the 07:31 timer are untouched, so at 07:31 the follow-up fires.
The PO reply takes no ask step, so its idempotency key comes from its triggering event and cannot collide with the 07:01 key. Eval: client_a_load_481207_full_timeline (and unrelated_driver_question_is_answered_and_ask_stays_open). Alternative: ask the model “did we already request the ETA?” over the chat; that costs a call and can be wrong. Also see ok_with_open_eta_ask_keeps_ask_open: an “ok” with the ask open is a noop that leaves the ask open.
7:00 AM Appended to the load log. The run reads the log and projects the state before it decides.7:01 AM Ping turns delay risk high: SMS the driver for the ETA and set the 7:31 AM follow-up7:12 AM Appended to the load log. The run reads the log and projects the state before it decides.7:13 AM Driver asks for the PO: answer it from metadata, the ETA ask stays open7:31 AM Appended to the load log. The run reads the log and projects the state before it decides.7:31 AM Timer: no ETA yet, ask again and set the 8:01 AM follow-up8:01 AM Appended to the load log. The run reads the log and projects the state before it decides.8:01 AM Timer: two requests unanswered, email the dispatcher in the main thread and create an urgent Hero task8:30 AM Nothing is left to do: the escalation chain is finished
Event log (append only)
- 1 7:00 AM Location ping
ping_0700delay risk high for stop 1 (Riverbend plant) - 2 7:01 AM Decision recorded
path sop, client_a/delay_risk_high@v1SMS to drv_20931: "Load 481207: what is your ETA to Riverbend plant? Please reply with the time."Timer set for 7:31 AMAskOpened ask:eta_request:ping_0700, attempt 1 - 3 7:12 AM SMS received
sms_0712from drv_20931 (driver)"What's the PO number?" - 4 7:13 AM Decision recorded
path triageSMS to drv_20931: "Load 481207: the PO number is 55821."no ask transition1 model call: triage - 5 7:31 AM Timer fired
evt:tmr:481207:ask:eta_request:ping_0700:1ask:eta_request:ping_0700, attempt 1 - 6 7:31 AM Decision recorded
path sop, client_a/delay_risk_high@v1SMS to drv_20931: "Load 481207: what is your ETA to Riverbend plant? Please reply with the time."Timer set for 8:01 AMAsk ask:eta_request:ping_0700: attempt 1 to 2 - 7 8:01 AM Timer fired
evt:tmr:481207:ask:eta_request:ping_0700:2ask:eta_request:ping_0700, attempt 2 - 8 8:01 AM Decision recorded
path sop, client_a/delay_risk_high@v1Email to dsp_4410: "The driver on load 481207 has not answered our texts. Please send the ETA to Riverbend plant."Hero task (urgent): "Load 481207: Call driver for eta: Call the driver on load 481207 to get the ETA to Riverbend plant. No ETA arrived after our requests."Ask ask:eta_request:ping_0700: escalations 0 to 2, status open to escalated - 9 8:30 AM Clock check
no recordno event is due, nothing is appended
Projected state (project_state over the log)
- Stage
- to pickup
- Delay risk, Riverbend plant
- high
- Delay risk, Valley DC
- low
- Open asks
- 0
No open ask yet.
- Stage
- to pickup
- Delay risk, Riverbend plant
- high
- Delay risk, Valley DC
- low
- Open asks
- 1
Open ask:eta_request:ping_0700
- Awaits
- ETA from drv_20931 (driver), Riverbend plant
- Attempt
- 1 of 2
- Next timer
- 7:31 AM
- Escalations done
- 0
- Resolved by
- null
- Pinned SOP
client_a/delay_risk_high@v1
- Stage
- to pickup
- Delay risk, Riverbend plant
- high
- Delay risk, Valley DC
- low
- Open asks
- 1
Open ask:eta_request:ping_0700
- Awaits
- ETA from drv_20931 (driver), Riverbend plant
- Attempt
- 1 of 2
- Next timer
- 7:31 AM
- Escalations done
- 0
- Resolved by
- null
- Pinned SOP
client_a/delay_risk_high@v1
unchanged: an inbound event alone never moves an ask
- Stage
- to pickup
- Delay risk, Riverbend plant
- high
- Delay risk, Valley DC
- low
- Open asks
- 1
Open ask:eta_request:ping_0700
- Awaits
- ETA from drv_20931 (driver), Riverbend plant
- Attempt
- 1 of 2
- Next timer
- 7:31 AM
- Escalations done
- 0
- Resolved by
- null
- Pinned SOP
client_a/delay_risk_high@v1
unchanged: no ask transition was committed
- Stage
- to pickup
- Delay risk, Riverbend plant
- high
- Delay risk, Valley DC
- low
- Open asks
- 1
Open ask:eta_request:ping_0700
- Awaits
- ETA from drv_20931 (driver), Riverbend plant
- Attempt
- 1 of 2
- Next timer
- 7:31 AM
- Escalations done
- 0
- Resolved by
- null
- Pinned SOP
client_a/delay_risk_high@v1
unchanged: an inbound event alone never moves an ask
- Stage
- to pickup
- Delay risk, Riverbend plant
- high
- Delay risk, Valley DC
- low
- Open asks
- 1
Open ask:eta_request:ping_0700
- Awaits
- ETA from drv_20931 (driver), Riverbend plant
- Attempt
- 2 of 2
- Next timer
- 8:01 AM
- Escalations done
- 0
- Resolved by
- null
- Pinned SOP
client_a/delay_risk_high@v1
- Stage
- to pickup
- Delay risk, Riverbend plant
- high
- Delay risk, Valley DC
- low
- Open asks
- 1
Open ask:eta_request:ping_0700
- Awaits
- ETA from drv_20931 (driver), Riverbend plant
- Attempt
- 2 of 2
- Next timer
- 8:01 AM
- Escalations done
- 0
- Resolved by
- null
- Pinned SOP
client_a/delay_risk_high@v1
unchanged: an inbound event alone never moves an ask
- Stage
- to pickup
- Delay risk, Riverbend plant
- high
- Delay risk, Valley DC
- low
- Open asks
- 1
Escalated ask:eta_request:ping_0700
- Awaits
- ETA from drv_20931 (driver), Riverbend plant
- Attempt
- 2 of 2
- Next timer
- none
- Escalations done
- 2
- Resolved by
- null
- Pinned SOP
client_a/delay_risk_high@v1
- Stage
- to pickup
- Delay risk, Riverbend plant
- high
- Delay risk, Valley DC
- low
- Open asks
- 1
Escalated ask:eta_request:ping_0700
- Awaits
- ETA from drv_20931 (driver), Riverbend plant
- Attempt
- 2 of 2
- Next timer
- none
- Escalations done
- 2
- Resolved by
- null
- Pinned SOP
client_a/delay_risk_high@v1
unchanged
Table view
| # | Time | Record | Details | ETA ask after the record |
|---|---|---|---|---|
| 1 | 7:00 AM | Location ping (ping_0700) | delay risk high for stop 1 (Riverbend plant) | none |
| 2 | 7:01 AM | Decision recorded (path sop, client_a/delay_risk_high@v1) | SMS to drv_20931: "Load 481207: what is your ETA to Riverbend plant? Please reply with the time."; Timer set for 7:31 AM; AskOpened ask:eta_request:ping_0700, attempt 1 | Open, attempt 1 of 2, next timer 7:31 AM |
| 3 | 7:12 AM | SMS received (sms_0712) | from drv_20931 (driver); "What's the PO number?" | Open, attempt 1 of 2, next timer 7:31 AM |
| 4 | 7:13 AM | Decision recorded (path triage) | SMS to drv_20931: "Load 481207: the PO number is 55821."; no ask transition; 1 model call: triage | Open, attempt 1 of 2, next timer 7:31 AM |
| 5 | 7:31 AM | Timer fired (evt:tmr:481207:ask:eta_request:ping_0700:1) | ask:eta_request:ping_0700, attempt 1 | Open, attempt 1 of 2, next timer 7:31 AM |
| 6 | 7:31 AM | Decision recorded (path sop, client_a/delay_risk_high@v1) | SMS to drv_20931: "Load 481207: what is your ETA to Riverbend plant? Please reply with the time."; Timer set for 8:01 AM; Ask ask:eta_request:ping_0700: attempt 1 to 2 | Open, attempt 2 of 2, next timer 8:01 AM |
| 7 | 8:01 AM | Timer fired (evt:tmr:481207:ask:eta_request:ping_0700:2) | ask:eta_request:ping_0700, attempt 2 | Open, attempt 2 of 2, next timer 8:01 AM |
| 8 | 8:01 AM | Decision recorded (path sop, client_a/delay_risk_high@v1) | Email to dsp_4410: "The driver on load 481207 has not answered our texts. Please send the ETA to Riverbend plant."; Hero task (urgent): "Load 481207: Call driver for eta: Call the driver on load 481207 to get the ETA to Riverbend plant. No ETA arrived after our requests."; Ask ask:eta_request:ping_0700: escalations 0 to 2, status open to escalated | Escalated, attempt 2 of 2, next timer none |
| 9 | 8:30 AM | Clock check (no record) | no event is due, nothing is appended | Escalated, attempt 2 of 2, next timer none |
Source: the scenario export of client_a_load_481207_full_timeline (open in the explorer).
3. How does a follow-up timer work? What does the agent see when the timer starts it?
Section titled “3. How does a follow-up timer work? What does the agent see when the timer starts it?”A timer is a durable claim that something may be due; state decides whether it is.
set_timeris an outbox action committed in the same transaction as the decision that took the step. The dispatcher writes it to the timer store withtimer_id = tmr:{load_id}:{ask_id}:{attempt}.fire_at = decided_at + retry_after.- The timer pump polls (at least every 30 s) for due timers, publishes
TimerFired(event_id derived from timer_id, payload {load_id, ask_id, attempt})to the load queue, then marks it fired. A crash between publish and mark publishes twice; the event store drops the second copy by event id. - The agent sees a normal event on the load’s queue, projects the state, and checks
ask.is_timer_current(payload): the ask must be open andpayload.attemptmust equal the ask’s current step. Otherwise it is a noop with no model call: the ask was resolved (ETA arrived by email at 07:20), the load arrived, the step moved on, or the timer is a duplicate. - If current, the executor continues the ask under its pinned SOP version: 07:31 sends the second SMS and sets the 08:01 timer; 08:01 runs the escalation chain (email the dispatcher in the main thread, urgent Hero task).
The timer also closes an ask when the risk eased: if the stop’s delay risk was read LOW after the ask’s last text, the current timer closes the ask with no text and no escalation; MEDIUM keeps the chain (risk_falls_back_cancels_the_eta_chain). Cancellation is housekeeping; correctness never depends on it. The timer survives deploys because it lives in the database; a late timer is harmless because state decides. Evals: timer_after_eta_received_is_noop, duplicate_timer_delivery_is_noop, driver_arrives_before_follow_up_cancels_eta_ask.
Alternatives: Temporal timers (strong, but a second source of truth next to our log); EventBridge Scheduler (per-schedule cost and quotas at our volume, cancellation can fail; kept as a drop-in TimerStore); in-process sleeps (lost on deploy). Trade-off: a pump to operate and monitor (lag alert at 2 minutes). Built: store/timers.py, agent/timer_pump.py.
- Ask opened, timer set
- Ask resolved
- Timer fired and acted
- Noop
- Timer pending
timer_after_eta_received_is_noop A timer delivered after the ETA arrived is a filtered noop with no LLM call.
duplicate_timer_delivery_is_noop The same timer delivered twice produces one follow-up SMS; the copy is dropped.
7:25 AM. A timer for attempt 1 is delivered although the ask is resolved: noop, no LLM. Path filtered. timer tmr:481207:ask:eta_request:ping_0700:1 is stale: the ask is closed or already past that step. No action. Model calls: 0.
Table view
| Scenario | # | Time | What happened | Decision |
|---|---|---|---|---|
timer_after_eta_received_is_noop | 1 | 7:01 AM | Delay risk high: SMS the driver for the ETA | Path sop. followed client_a/delay_risk_high@v1. Actions: send sms, timer set for 7:31 AM. Model calls: 0. |
timer_after_eta_received_is_noop | 2 | 7:20 AM | Dispatcher emails the ETA: ask resolved | Path triage. ask:eta_request:ping_0700 answered by email_0720. Actions: cancel timer, tms note. Model calls: 1. |
timer_after_eta_received_is_noop | 3 | 7:25 AM | A timer for attempt 1 is delivered although the ask is resolved: noop, no LLM | Path filtered. timer tmr:481207:ask:eta_request:ping_0700:1 is stale: the ask is closed or already past that step. No action. Model calls: 0. |
duplicate_timer_delivery_is_noop | 1 | 7:01 AM | Delay risk high: SMS the driver for the ETA | Path sop. followed client_a/delay_risk_high@v1. Actions: send sms, timer set for 7:31 AM. Model calls: 0. |
duplicate_timer_delivery_is_noop | 2 | 7:31 AM | The pump delivers the timer: second SMS | Path sop. ask:eta_request:ping_0700 unanswered at step 1, continued client_a/delay_risk_high@v1. Actions: send sms, timer set for 8:01 AM. Model calls: 0. |
duplicate_timer_delivery_is_noop | 3 | 7:31 AM | The same timer is delivered again: dropped as a duplicate | Path filtered. event evt:tmr:481207:ask:eta_request:ping_0700:1 was already processed. No action. Model calls: 0. |
Source: the scenario exports of timer_after_eta_received_is_noop and duplicate_timer_delivery_is_noop.
Open, follow up, escalate or resolve, and the timers that must do nothing, from three scenario runs.
The player could not load. The rendered MP4 plays below, with the same captions.
Scene list (text version)
- 0:00 An ask is state projected from the log; timers only say that something may be due.
- 0:06 7:01 AM (client_a_load_481207_full_timeline): The SOP opens an ask: a row of state with a target, an attempt count, a timer and its SOP version.
- 0:15 7:13 AM (client_a_load_481207_full_timeline): An unrelated question gets its answer. No ask transition: the ask, its attempt and its timer stay as they were.
- 0:22 7:31 AM (client_a_load_481207_full_timeline): The timer is current (same attempt, ask still open), so the SOP takes the next step.
- 0:29 8:01 AM (client_a_load_481207_full_timeline): Attempts are used up: the escalation chain runs in one decision.
- 0:36 7:20 AM (timer_after_eta_received_is_noop): In another run the dispatcher emails the ETA. Any contact on any channel resolves the ask; its timer is cancelled.
- 0:45 7:25 AM (timer_after_eta_received_is_noop): The old timer is delivered anyway. State says the ask is closed: noop, no model call.
- 0:53 7:31 AM (duplicate_timer_delivery_is_noop): The same timer delivered twice: the event id is already in the log, the copy is dropped.
- 1:01 A timer acts only if its ask is open and on the same attempt; late, stale or duplicated timers are noops with no model call.
Generated from the scenario export at build time. Download MP4 (1.6 MB) with captions (WebVTT). Narration and sound effects generated with ElevenLabs.
1. How do you make sure that the system works correctly in production?
Section titled “1. How do you make sure that the system works correctly in production?”Layers, from before release to after (design.md section 12):
- Before release: unit and contract tests, scenario evals, CI on every PR (ruff, format, mypy strict, pytest with a Postgres service).
- Validation at run time: every action is checked before commit (contact belongs to the load, channel supported, role not forbidden by the client’s SOPs for the whole load, idempotency key unused). Rejected actions become a Hero task. The agent can do harm only through actions that passed these checks.
- Decision records: every run records path, SOP version, actions, rationale and LLM usage, so any action is explainable and replayable.
- Metrics and alerts: path mix, escalation rate per client and topic, Hero override rate (primary quality metric), run latency against 5 minutes, outbox age, dead rows, timer pump lag, cost per load, triage confidence distribution, noop share.
- Human review loop: Heroes mark a task as “agent was wrong” or reverse an action; a weekly sample of decisions per client is reviewed; every confirmed error becomes a new scenario eval.
- Staged rollout: shadow mode for SOP and model changes (decide, do not act, compare), then a canary by client or by share of loads.
Alternative: LLM-as-judge on production traffic. Useful for tone, but structural checks and Hero overrides are cheaper and less ambiguous; we use a judge only where structure cannot decide. Built: decision records, validation, evals, CI. Designed: dashboards, alerts, review loop, shadow and canary.
2. A client changes its procedure. How do you make sure that other behavior does not break?
Section titled “2. A client changes its procedure. How do you make sure that other behavior does not break?”- Scope is structural: a client file overrides one topic for one client. Other clients resolve to their own file or standard, so they cannot change.
- Suites run on every SOP change: the standard suite (it must stay green because the change must not leak) plus the client’s suite.
- Behavior diff: the old and new versions run on the same scenarios and the editor sees which decisions changed, with an explanation, not a text diff. An unexpected change in an unrelated scenario blocks publish. Today
render_behavior_diffcompares specs (built, shown bysop preview); the scenario-based diff is designed. - Compile validation rejects ambiguous or contradictory text before it becomes a spec.
- Shadow mode on live loads for the client: compare decisions with the current version before publish.
- Version pinning: asks in flight finish under their version; only new triggers use the new one.
- One-click rollback republishes the previous spec.
Trade-off: a client with no scenarios has nothing to regress against, so onboarding includes writing its first scenarios (the 481207 file is the template), and a change to a topic with no coverage is flagged as unprotected (designed). Alternative: rely on shadow mode alone; slower feedback and only covers traffic that happens to occur.
3. You change the LLM. How do you make sure that the system does not break?
Section titled “3. You change the LLM. How do you make sure that the system does not break?”Two kinds of tests, because they answer different questions:
- Pipeline logic (scripted
FakeLLM): the suite runs offline with the model’s outputs scripted per scenario. It proves the pipeline is correct for given model outputs and runs on every commit with no API key. An unscripted call fails the test, so “this step must not call a model” is enforced (llm_calls: 0). It does not test the model. - Model behavior (replay): a frozen set of real inbound messages with labels (human-checked or Hero-confirmed decisions from production) is run through the candidate model for each task. Compared per task: triage category and extracted value accuracy (ETA parsing with timezones is the sharp edge), confidence calibration, planner action agreement, document field accuracy, schema-validity rate, cost and latency. A change must meet or beat the current model on every gate; thresholds (
PRESENT_MIN,GUARD_MAX) are re-fitted on the replay set, not copied.
Built for triage: tests/llm/test_gate.py replays recorded decision-model answers (two samples of 190 phrasings) through the real client and fails on any wrong value; python -m tests.llm.calibrate refits PRESENT_MIN and GUARD_MAX from those recordings; .github/workflows/live.yml runs the live smoke tests daily and on demand. The recorded phrasings were written by us, so this gate does not replace a production replay set. Production must also pin the calibrated model version (the default jev-latest floats, and an upgrade invalidates the thresholds).
Then shadow (candidate runs beside the current model on live loads, only the current one acts, disagreements are reviewed) and canary (a small share of loads or one client, with the Hero override rate and escalation rate as rollback triggers). Model ids are per task in LLMSettings, so each task can change on its own, and the decision record stores the model id of every call, so a regression is attributable.
Alternative: compare only final actions end to end. Simple, but a coincidence can hide a worse classifier. Trade-off: the replay set needs curation and refresh; it is the main cost of changing models safely. Built: the scripted suite and the typed failure handling (an invalid output becomes a Hero task, never a guess). Designed: the production replay set, shadow and canary.
Try the ETA rule
Value Accepted: the code's reading of the time resolves the ask.
The presets load the recorded answers of jev-1.13.0 (sample 1); the sliders are yours to move. The thresholds are the calibrated constants: main at least 0.60, guards at most 0.30 (ETA), 0.30 (temperature), 0.20 (seal).
Table view
| Set | Sample 1: accepted, wrong | Sample 2: accepted, wrong |
|---|---|---|
| Fit 70% (51 positives, 81 negatives) | 43/51, 0 wrong | 44/51, 0 wrong |
| Held-out 30% (23 positives, 35 negatives) | 19/23, 0 wrong | 17/23, 0 wrong |
| All (74 positives, 116 negatives) | 62/74, 0 wrong | 61/74, 0 wrong |
Source: tests/llm/calibrate.py over tests/llm/fixtures/jev_gate.json and jev_gate_sample2.json (190 phrasings, two recordings), thresholds from llm/typesafe.py.
4. Clients have different procedures and different complexity. How do you organize the evals?
Section titled “4. Clients have different procedures and different complexity. How do you organize the evals?”By layer, mirroring the SOP layers, plus layers of test for the machinery:
| Layer | What it proves | Where |
|---|---|---|
| Unit and contract tests | Domain rules, projection, store/outbox/timer contracts (SQLite and Postgres run the same suite), validation, templates, compiler rules | tests/ |
| Standard scenarios | Behavior every client gets unless it overrides the topic | designed layout evals/standard/ |
| Client scenarios | Overrides: Client A escalation chain, Client B Slack alert and “never the dispatcher”, Client B reefer temperature and seal | designed layout evals/clients/<id>/ |
| Cross-cutting scenarios | Noop cases, cross-channel resolution, duplicates, stale pings, unknown situations | scenario files |
| Replay of production decisions | Model changes and drift | designed |
A client change runs standard plus that client. A change to a standard topic runs every client that does not override it (derivable from the SOP tree). Each scenario is a YAML timeline replayed with FakeClock through the real worker, dispatcher and timer pump; each step asserts the exact set of actions (kind, contact, channel, ask, attempt, key) or an explicit noop, and optionally path, number of LLM calls, state and whether the log grew. Message content is checked structurally (the PO number is present). A judge is used only where structure cannot decide.
Today the scenarios sit in evals/scenarios/, named by client where relevant; the split into standard/ and clients/<id>/ is a file move plus a runner that selects by SOP tree, not a redesign. Complexity scales by adding scenarios per client, not by editing a shared test: a complex client gets many files, a simple one gets the standard suite for free. Alternative: one parameterized suite across all clients; compact, but a failure does not say which client’s procedure broke and per-client edge cases get lost.
Do-nothing coverage: ok_with_no_open_ask_is_noop, ok_with_open_eta_ask_keeps_ask_open, timer_after_eta_received_is_noop, duplicate_timer_delivery_is_noop, stale_ping_is_ignored, high_ping_while_already_high_is_not_a_trigger.
30scenario files, every step asserted
15include a "do nothing" step
1503tests: 1402 offline, 101 against Postgres in CI
Source: the scenario export index (30 files in evals/scenarios/, all replayed and matching in this build) and the default pytest run.
Scale, response time and cost
Section titled “Scale, response time and cost”1. How does your design keep each agent run at 5 minutes or less?
Section titled “1. How does your design keep each agent run at 5 minutes or less?”Most runs are milliseconds because most events never reach a model (design.md section 13).
| Step | Typical | Bound |
|---|---|---|
| Queue and ingest | under 1 s | FIFO group per load, no cross-load blocking |
| Read and project | under 50 ms | about 1,300 rows per load with pings; snapshot row every 200 records beyond about 2,000 rows (designed) |
| Filter, trigger diff, SOP executor, templates | milliseconds | pure code |
| Triage call (only if needed) | about 0.2 s (measured, ADR 0005) | 30 s timeout, 2 retries, capped by the run deadline (LLMSettings) |
| Planner call (rare) | 3 to 10 s | 30 s timeout, 2 retries, capped by the run deadline |
| Commit | under 20 ms | optimistic concurrency, up to 3 re-decides (max_conflict_retries) |
| Dispatch | seconds | RetryPolicy.budget (150 s) from the first send; permanent failure escalates |
| Timer pump | under 1 minute after fire_at | poll every 30 s; alert on lag over 2 minutes |
Why it holds: serial per load with optimistic concurrency means no waiting on locks; outbox rows are claimed by many dispatchers in parallel; the pump scales by partition. Failure paths are bounded: an LLM call gets at most 2 retries inside the client and then becomes a Hero task, and a failed send is a failed outcome the agent escalates.
Bound on the model path (built, agent/deadline.py, RetryPolicy.budget): a 120 s run deadline spans all conflict re-decides. Once it is spent, no model call starts and the run commits an urgent Hero task. An in-flight HTTP attempt is capped at the time left. Event to last action is at most 270 s, including the 150 s dispatch budget. Residual: httpx timeouts are per phase (connect, write, read), so the bound assumes small non-streamed responses. Alternative: per-call timeouts alone, which bound one call but not the sum of a triage call, a planner call and the conflict re-decides.
The slowest run, 21.2 s (triage 192 ms + plan 21.0 s), used 18% of the run deadline. Once the deadline is spent no model call starts and the run commits an urgent Hero task, so a slow provider costs a Hero task, not the 5 minutes.
Table view
| Part | Seconds |
|---|---|
| Run deadline (all model calls and conflict re-decides) | 120 |
| Dispatch retry budget, from the first send | 150 |
| Bound, event to last action | 270 of 300 |
| Measured p50, all runs | 0.02 |
| Measured p95, all runs | 1.56 |
| Measured p95, runs with a model call | 7.01 |
| Measured slowest (bench:1:053) | 21.21 |
Source: agent/deadline.py (RUN_DEADLINE), agent/dispatcher.py (RetryPolicy.budget), bench/results/2026-10-08-typesafe-deepseek.json (91 runs).
2. The volume increases ten times. What must change?
Section titled “2. The volume increases ten times. What must change?”Today’s design has no per-process state, so the worker tier scales by adding workers; what changes is the data and external limits (design.md section 13):
| Pressure | Change |
|---|---|
| Postgres: about 500 inserts per second average and 1,500 at peak with pings in the log, 14B log rows per year | Partition log and outbox by load_id hash, archive closed loads to object storage after 30 days, read replicas for dashboards, shard by client group if one cluster saturates |
| Queue | SQS FIFO with many message groups; Kafka keyed by load if per-group limits or cost bite (trade-offs in ADR 0003) |
| Timer pump and dispatcher | More instances, partitioned by fire_at bucket and hash, SKIP LOCKED; per-service rate limits and circuit breakers |
| LLM provider limits | Reserved throughput, a fallback provider, a concurrency cap on the queue side |
| Pings (the largest input, 45M per month at 1x on the assumption of one per 15 minutes) | Filter at ingestion against a cached band per stop, or keep only band-changing pings in the log and the rest in a telemetry table (designed) |
| Cost | Triage is the largest line and already runs on the decision model. Next levers: debounce consecutive SMS into one call, and a distilled in-house classifier validated by the recorded-answer gate (both designed) |
| Humans | Escalations scale with volume; watch tasks per load and fix SOPs where Heroes repeat the same manual step |
Alternative: move now to Kafka plus a separate state store. It scales writes further but splits the event, decision and outbox commit across two systems, which ADR 0003 rejects until Postgres partitioning saturates (it says to reconsider Kafka if write throughput becomes the bottleneck). Trade-off of ours: a bigger Postgres operation (partitions, archive, replicas) in exchange for keeping one transaction.
What does not change: the contracts. EventStore, Outbox, TimerStore, LoadQueue and LLMClient are protocols, so a different queue or store plugs in behind them.
- Loads per month
- 100K
- Inbound messages per month
- 5M to 10M
- Model calls per month
- 4.4M to 8.8M
- Model cost per month
- $9.8K to $19.6K
- Loads per month
- 1M
- Inbound messages per month
- 50M to 100M
- Model calls per month
- 43.8M to 87.5M
- Model cost per month
- $98.2K to $196.5K
- Workers Same at 10x No per-process state: add workers. Unchanged contracts (
EventStore,Outbox,TimerStore,LoadQueue,LLMClient). - Postgres No change at 1x Changes at 10xPartition log and outbox by
load_idhash, archive closed loads to object storage after 30 days, read replicas for dashboards, shard by client group if one cluster saturates - Queue No change at 1x Changes at 10xSQS FIFO with many message groups; Kafka keyed by load if per-group limits or cost bite (trade-offs in ADR 0003)
- Timer pump and dispatcher No change at 1x Changes at 10xMore instances, partitioned by
fire_atbucket and hash,SKIP LOCKED; per-service rate limits and circuit breakers - LLM provider limits No change at 1x Changes at 10xReserved throughput, a fallback provider, a concurrency cap on the queue side
- Pings No change at 1x Changes at 10xFilter at ingestion against a cached band per stop, or keep only band-changing pings in the log and the rest in a telemetry table (designed)
- Cost No change at 1x Changes at 10xTriage is the largest line and already runs on the decision model. Next levers: debounce consecutive SMS into one call, and a distilled in-house classifier validated by the recorded-answer gate (both designed)
- Humans No change at 1x Changes at 10xEscalations scale with volume; watch tasks per load and fix SOPs where Heroes repeat the same manual step
Table view
The changes are the table in the answer above. Volume at 1x: loads per month 100K; inbound messages per month 5M to 10M; model calls per month 4.4M to 8.8M; model cost per month $9.8K to $19.6K. At 10x: loads per month 1M; inbound messages per month 50M to 100M; model calls per month 43.8M to 87.5M; model cost per month $98.2K to $196.5K.
Source: volume from docs/assignment.md; 0.88 model calls and $0.0020 per inbound message measured in bench/results/2026-10-08-typesafe-deepseek.json (56 messages); changes from the table above.
3. How do you keep the LLM cost at approximately $1 or less for each load? You do not have to implement this.
Section titled “3. How do you keep the LLM cost at approximately $1 or less for each load? You do not have to implement this.”A funnel in which the model is the last resort (ADR 0001, ADR 0004; arithmetic in design.md section 13).
Volume: 100,000 loads, 50 to 100 messages each, 5M to 10M messages per month. $1 over 75 messages is about $0.013 per message. A model call on every message at that price would spend the whole budget on triage and leave nothing for the planner and documents, and any call above that price breaks it. The funnel:
| Stage | Events | Cost |
|---|---|---|
| Deterministic filter: duplicates, acknowledgements with no open ask, unknown senders (assumed 25% to 40% of inbound messages; basis: driver SMS is short and often “ok” or “thanks”, to be measured from the path mix) | 25% to 40% of messages | $0 |
| Timers that are stale or resolved, delivery outcomes, pings with no band change (system events, not inbound messages) | all of them | $0 |
| SOP executor: trigger fired, timer continues an ask | the procedural core | $0 |
| Outbound messages: templates filled in code | 40 to 70 per load | $0 |
| Triage: decision model (TypeSafe Jev), cached system prompt, short request (one message plus open asks) | 30 to 75 calls | $0.0005 to $0.003 each, so $0.015 to $0.23 |
| Planner: strong model, no carrier-facing actions (a Hero task, a TMS note, or a Slack post to the broker) | 0 to 2 | $0.02 to $0.05 each, so up to $0.10 |
| Documents: vision | 2 to 6 | $0.01 to $0.02 each, so $0.02 to $0.12 |
| Template-less drafts | 0 to 1 | up to $0.01 |
| SOP compile, amortized (100 clients times 20 saves per month at $0.10) | per edit, not per event | about $0.002 per load |
Total about $0.04 to $0.46 per load, which is 2x to 25x of headroom under $1. The prices are assumptions to replace with contract prices; the structure is what matters. Load 481207 makes one model call in the whole timeline.
Guards, so the average stays low: prompt caching for the system prompt and SOP guidance; a per-load cost counter from the LLMUsage records, with a warning at $0.60 and a ceiling at $2 per load (twice the budget, so only a runaway load such as a chatty driver or a looping SOP reaches it) that disables the planner and escalates to Hero tasks instead (designed); debounce of consecutive SMS into one triage call (designed); path-mix monitoring, since a rising planner share is the early signal of missing SOP coverage; a regression gate in evals (llm_calls per step), so a change that adds hidden calls fails CI.
Alternatives: a cheaper model for everything (accuracy drops on the language-heavy cases and the cost of a wrong action exceeds the saving); an LLM draft for every message (adds a call per outbound message and nondeterministic text, and verification catches a lost fact but not an invented claim); caching responses by message text (works for acknowledgements and is covered by the filter, but unsafe for messages whose meaning depends on open asks). Trade-off of ours: templates are less adaptive than model text, and the planner path is cheap only while SOP coverage stays good.
- Triage on
jev-1.13.0: $0.090, 45 calls, at an assumed $0.002 per call - Planner on
deepseek-v4-pro: $0.020, 4 calls, measured
9.1x headroom under the budget. The decision model's price is an assumption: the cost reaches $1.00 at $0.0218 per triage call. Change the price and the volume in the calculator.
Table view
| Purpose | Model | Calls | Cost per load | Basis |
|---|---|---|---|---|
| Triage | jev-1.13.0 | 45 | $0.090 | assumed price |
| Planner | deepseek-v4-pro | 4 | $0.020 | measured |
| Total | 49 | $0.110 | budget $1.00 |
Source: bench/results/2026-10-08-typesafe-deepseek.json, one load with 56 inbound messages. Benchmark and cost calculator.