Run it locally
An event-driven agent that monitors a truckload freight load, follows each client’s SOP, acts on its own when safe, escalates to the Hero team or the broker when not, and does nothing when nothing is needed.
Every input and every agent action is appended to a per-load event log; load state, including open asks such as a pending ETA request, is a projection of that log. Deterministic code handles filtering, timers, attempt counts and escalation chains compiled from plain-English SOPs. A language model is used only to classify messages that pass the filter and for rare situations no SOP covers. Messages to drivers and dispatchers come from templates filled in code; the only model-written text is a fallback draft whose named facts are checked verbatim, and a planner post to the broker’s Slack.
Short on time? Read the executive summary (three minutes). How success is measured and how it ships: rollout and metrics.
Run it
Section titled “Run it”Requires Python 3.12+ and uv.
uv syncuv run pytestThat runs the unit tests and all scenario evals offline, with a scripted fake LLM. No API key, no database and no network are needed. The Postgres contract tests are skipped unless LOAD_AGENT_TEST_PG_URL is set (CI runs them against Postgres 16).
Watch the agent handle load #481207 step by step:
uv run load-agent run evals/scenarios/client_a_load_481207_full_timeline.yaml7:01 AM ping turns delay risk high: SMS the driver for the ETA and set the 7:31 AM follow-up7:13 AM driver asks for the PO: answer it from metadata, the ETA ask stays open7:31 AM timer: no ETA yet, ask again and set the 8:01 AM follow-up8:01 AM timer: two requests unanswered, email the dispatcher in the main thread and create an urgent Hero taskAny file in evals/scenarios/ runs the same way, for example the “do nothing” case:
uv run load-agent run evals/scenarios/ok_with_no_open_ask_is_noop.yamlExport every scenario as JSON for the docs site explorer (index.json plus one file per scenario):
uv run load-agent export --out build/scenarios-jsonSee what the agent understood from a client’s SOP and how it differs from the layer it overrides:
uv run load-agent sop preview client_a delay_risk_highuv run load-agent sop preview client_b delay_risk_highuv run load-agent sop preview client_b departed_pickup --load-type reeferRunning with real models (optional)
Section titled “Running with real models (optional)”The default run never calls a model. To use the real clients (triage on the TypeSafe AI decision model, planner, SOP compilation, documents and fallback drafts on DeepSeek), set:
export LOAD_AGENT_LLM_PROVIDER=typesafe+deepseekexport TYPESAFE_AI_API_KEY=...export DEEPSEEK_API_KEY=...uv run load-agent sop compile sops/clients/client_a/delay_risk_high.md # prints the preview, add --write to saveuv run pytest -m live -s # live smoke tests: opt-in, excluded from the default runMeasure LLM cost and run latency for one synthetic load day (87 inputs) against the real providers; the result and the report are in docs/benchmark.md:
uv run load-agent bench --dry-run # no network: expected call countsuv run load-agent bench # real providers, capped at 150 model callsThe same live tests run in GitHub Actions (.github/workflows/live.yml) daily, on demand, and on pushes to main that touch the LLM clients, using the TYPESAFE_AI_API_KEY and DEEPSEEK_API_KEY repository secrets. They stay out of the per-push CI so a provider outage never blocks a merge.
Optional overrides: LOAD_AGENT_MODEL_TRIAGE (default jev-latest, which floats: the triage thresholds were calibrated on one version, so production must pin it, because a provider upgrade invalidates them), LOAD_AGENT_MODEL_DRAFT and _DOCUMENT (deepseek-flash), LOAD_AGENT_MODEL_PLANNER and _SOP (deepseek-v4-pro), LOAD_AGENT_LLM_TIMEOUT_SECONDS (30; sop compile defaults to 300 because the strong model takes about a minute), LOAD_AGENT_LLM_MAX_RETRIES (2), LOAD_AGENT_ATTACHMENTS_DIR (local document files are read only from inside it; without it only http(s) and data URLs are accepted), and prices for the cost estimate in the decision record (LOAD_AGENT_TYPESAFE_USD_PER_CALL, LOAD_AGENT_DEEPSEEK_USD_PER_MTOK_INPUT, _OUTPUT and _CACHED). Prices default to 0, so a $0 cost in a decision record is a placeholder until you set them. LOAD_AGENT_LLM_PROVIDER=fake builds an unscripted fake. Keys are read from the environment only and never logged.
Each scenario in evals/scenarios/ replays a timeline with a fake clock and asserts, at every step, the exact set of actions (type, target contact, channel, ask and attempt) or an explicit noop, plus the state of the open asks. They cover:
- the full load #481207 timeline for Client A (07:00 to 08:02), and the PO answer with the assignment’s literal load JSON;
- “ok” with no open question (do nothing, no LLM call) and “ok” while the ETA ask is open;
- an unrelated driver question answered while the ETA ask stays open;
- the ETA arriving from the dispatcher by email (cross-channel resolution);
- a timer firing after the ask was resolved, and the same timer delivered twice;
- the driver arriving before the follow-up;
- Client B (one SMS, then a broker Slack alert, never the dispatcher) and Client B reefer (temperature and seal number);
- a PO question when the load has no PO, stale and repeated pings, and an unknown situation (Hero task).
- a thumbs-up reply with or without an open ask (noop, no model call);
- delay risk that flaps (high, low, high or through medium) keeps one chain on schedule, and risk that stays low closes it at the next follow-up;
- a reply that names a stop resolves only that stop’s ask, and asks for stops already passed close;
- a broker “stand down” message reaches a Hero, whose override ends the chain; an unknown sender gets one Hero task per distinct message per 30 minutes.
Repository map
Section titled “Repository map”| Path | Contents |
|---|---|
docs/assignment.md | The assignment |
docs/executive-summary.md | One-page summary: decisions, built versus designed, what was skipped, risks |
docs/rollout-and-metrics.md | Success metrics, staged rollout with gates, kill switch, ownership |
docs/design.md | Full system design, built versus designed, diagrams |
docs/answers.md | Answers to the assignment questions |
docs/adr/ | Architecture decision records |
src/load_agent/ | Agent code: domain models, event store and projection, SOP repository and compiler, pipeline, dispatcher, timers |
sops/ | SOPs in Markdown with their compiled specs and message templates, layered: standard, client, client + load type |
evals/ | Scenario files, fixtures and scenario tests |
src/load_agent/api/, src/load_agent/service/ | HTTP ingestion API, worker loop and Postgres wiring |
Dockerfile, docker-compose.yml, scripts/demo.sh | Run as a service and the demo |
tests/ | Unit and contract tests |