Skip to content

Executive summary

For an engineering leader, three minutes. Detail: design.md, answers.md, rollout-and-metrics.md, benchmark.md.

Bottom line. A working agent replays load #481207 end to end with one model call, passes 30 scenario evals offline with no API key, and escalates to a Hero only for the 08:01 call. Production readiness depends on a labelled production triage set and Hero feedback tooling, both designed and not built. Measured with the real providers on a synthetic day of 87 inputs (benchmark.md): $0.110 of LLM per load against the $1 budget (TypeSafe at an assumed per-call price; break-even $0.0218 per call), slowest run 21 s against the 5-minute limit, and triage matching the expected outcome on 30 of 45 calls. Every disagreement but one fails toward a Hero task or a TMS note; the one exception is an unhelpful reply to the driver.

A load needs someone to notice trouble and follow the client’s SOP for it: text the driver, wait 30 minutes, text again, email the dispatcher, then call. Done by hand across 100,000 loads a month, that is Hero time spent on waiting and repeating; done by one large model reading chat history, it is unpredictable and costs more than the budget of about $1 per load.

DecisionRejected alternativeWhy
The event log is the memory. Every input and action is appended per load; state, including open asks (an ETA request with its attempts, timer and resolved_by), is a projection. (ADR 0001)Chat history or a vector store in the prompt; a mutable state row“Is the ETA request still open?” is answered by state, replay is exact, and there is an audit trail. Cost: stored models must stay readable across schema changes.
The LLM is the exception. Filters, state transitions, timers, attempt counts and message text come from code and templates; a model classifies language and handles uncovered cases. (ADR 0001, ADR 0004)One ReAct agent with tools reading the history on every eventCost stays flat as history grows and behavior is testable step by step. Cost: new situations need a new topic.
SOPs are compiled and layered, with a preview. Plain-English Markdown is compiled into a structured spec, layered standard, client, client + load type; an editor approves a plain-English preview (built) and a behavior diff (approval flow designed). (ADR 0002)Raw Markdown to a model on every event; embedding retrievalNon-engineers can change an SOP safely and the runtime is deterministic. Cost: whole-topic override duplicates text.
Triage uses a decision model; code owns the values. Calibrated probabilities per typed question feed thresholds; ETAs, temperatures and seal numbers are parsed by code; doubt escalates. (ADR 0005)A generative model that returns the valueA model that cannot write a value cannot invent one. Cost: more escalations (rates and their limits in the ADR).
Transactional outbox and durable timers in Postgres. Event, decision and actions commit in one transaction; a dispatcher sends with idempotency keys; timers survive restarts. (ADR 0003)Temporal or EventBridge for timers; sending inside the runNo lost or duplicate sends on a crash, one source of truth. Cost: a pump and a dispatcher to operate.

Built and tested offline with uv run pytest (no API key): the log and projection (SQLite and Postgres), open asks, outbox, dispatcher, timers, layered SOPs with sop preview, the pipeline, validation, 30 scenario evals including load #481207 and the noop cases, optional real clients for triage and generation, and the service itself (FastAPI ingestion, Postgres queue, worker, docker compose up, proven end to end in CI). Designed only: provider signature verification, SQS and autoscaling, real channel integrations, SOP editor and publish flow, shadow and canary tooling, the kill switch, the observability stack. Section 1 of design.md has the full table.

SkippedWhy
Web UI, SOP editor, dashboardsGraded behavior is the decision core; a UI would not change a decision. Previews and decision records carry what a UI would show.
Real SMS, email, Slack, TMS integrationsThey hide the logic behind vendor setup. Typed clients and contract-tested fakes keep the seam honest.
SQS, autoscaling, provider signature checksThe Postgres queue implements the same contract an SQS FIFO adapter would; auth is a shared secret.
Tuning on production dataNone exists. Triage thresholds come from 190 phrasings we wrote.
  • Triage calibration is on written phrasings, not real driver SMS; a labelled production replay set is required before launch, and the model version must be pinned (ADR 0005, design.md section 14).
  • Cost and path-mix figures in the design are estimates from assumptions until measured; prices default to 0 in decision records.
  • Vision extraction on real BOLs and PODs is unverified.
  • The metrics that judge quality (wrong-action and silent-miss rates) need Hero feedback tooling that is designed, not built.

Offline replay, shadow mode, canary by action type, then full rollout, each with a go/no-go gate and a per-client, per-action kill switch (designed). Seven metrics, with their definitions, sources and which fields exist today: rollout-and-metrics.md. Measured numbers for the triage and cost claims: benchmark.md.

Prepared for Freight Hero by Marcus Caum Source on GitHub