Skip to content

Run it locally

An event-driven agent that monitors a truckload freight load, follows each client’s SOP, acts on its own when safe, escalates to the Hero team or the broker when not, and does nothing when nothing is needed.

Every input and every agent action is appended to a per-load event log; load state, including open asks such as a pending ETA request, is a projection of that log. Deterministic code handles filtering, timers, attempt counts and escalation chains compiled from plain-English SOPs. A language model is used only to classify messages that pass the filter and for rare situations no SOP covers. Messages to drivers and dispatchers come from templates filled in code; the only model-written text is a fallback draft whose named facts are checked verbatim, and a planner post to the broker’s Slack.

Short on time? Read the executive summary (three minutes). How success is measured and how it ships: rollout and metrics.

Requires Python 3.12+ and uv.

Terminal window
uv sync
uv run pytest

That runs the unit tests and all scenario evals offline, with a scripted fake LLM. No API key, no database and no network are needed. The Postgres contract tests are skipped unless LOAD_AGENT_TEST_PG_URL is set (CI runs them against Postgres 16).

Watch the agent handle load #481207 step by step:

Terminal window
uv run load-agent run evals/scenarios/client_a_load_481207_full_timeline.yaml
7:01 AM ping turns delay risk high: SMS the driver for the ETA and set the 7:31 AM follow-up
7:13 AM driver asks for the PO: answer it from metadata, the ETA ask stays open
7:31 AM timer: no ETA yet, ask again and set the 8:01 AM follow-up
8:01 AM timer: two requests unanswered, email the dispatcher in the main thread and create an urgent Hero task

Any file in evals/scenarios/ runs the same way, for example the “do nothing” case:

Terminal window
uv run load-agent run evals/scenarios/ok_with_no_open_ask_is_noop.yaml

Export every scenario as JSON for the docs site explorer (index.json plus one file per scenario):

Terminal window
uv run load-agent export --out build/scenarios-json

See what the agent understood from a client’s SOP and how it differs from the layer it overrides:

Terminal window
uv run load-agent sop preview client_a delay_risk_high
uv run load-agent sop preview client_b delay_risk_high
uv run load-agent sop preview client_b departed_pickup --load-type reefer

The default run never calls a model. To use the real clients (triage on the TypeSafe AI decision model, planner, SOP compilation, documents and fallback drafts on DeepSeek), set:

Terminal window
export LOAD_AGENT_LLM_PROVIDER=typesafe+deepseek
export TYPESAFE_AI_API_KEY=...
export DEEPSEEK_API_KEY=...
uv run load-agent sop compile sops/clients/client_a/delay_risk_high.md # prints the preview, add --write to save
uv run pytest -m live -s # live smoke tests: opt-in, excluded from the default run

Measure LLM cost and run latency for one synthetic load day (87 inputs) against the real providers; the result and the report are in docs/benchmark.md:

Terminal window
uv run load-agent bench --dry-run # no network: expected call counts
uv run load-agent bench # real providers, capped at 150 model calls

The same live tests run in GitHub Actions (.github/workflows/live.yml) daily, on demand, and on pushes to main that touch the LLM clients, using the TYPESAFE_AI_API_KEY and DEEPSEEK_API_KEY repository secrets. They stay out of the per-push CI so a provider outage never blocks a merge.

Optional overrides: LOAD_AGENT_MODEL_TRIAGE (default jev-latest, which floats: the triage thresholds were calibrated on one version, so production must pin it, because a provider upgrade invalidates them), LOAD_AGENT_MODEL_DRAFT and _DOCUMENT (deepseek-flash), LOAD_AGENT_MODEL_PLANNER and _SOP (deepseek-v4-pro), LOAD_AGENT_LLM_TIMEOUT_SECONDS (30; sop compile defaults to 300 because the strong model takes about a minute), LOAD_AGENT_LLM_MAX_RETRIES (2), LOAD_AGENT_ATTACHMENTS_DIR (local document files are read only from inside it; without it only http(s) and data URLs are accepted), and prices for the cost estimate in the decision record (LOAD_AGENT_TYPESAFE_USD_PER_CALL, LOAD_AGENT_DEEPSEEK_USD_PER_MTOK_INPUT, _OUTPUT and _CACHED). Prices default to 0, so a $0 cost in a decision record is a placeholder until you set them. LOAD_AGENT_LLM_PROVIDER=fake builds an unscripted fake. Keys are read from the environment only and never logged.

Each scenario in evals/scenarios/ replays a timeline with a fake clock and asserts, at every step, the exact set of actions (type, target contact, channel, ask and attempt) or an explicit noop, plus the state of the open asks. They cover:

  • the full load #481207 timeline for Client A (07:00 to 08:02), and the PO answer with the assignment’s literal load JSON;
  • “ok” with no open question (do nothing, no LLM call) and “ok” while the ETA ask is open;
  • an unrelated driver question answered while the ETA ask stays open;
  • the ETA arriving from the dispatcher by email (cross-channel resolution);
  • a timer firing after the ask was resolved, and the same timer delivered twice;
  • the driver arriving before the follow-up;
  • Client B (one SMS, then a broker Slack alert, never the dispatcher) and Client B reefer (temperature and seal number);
  • a PO question when the load has no PO, stale and repeated pings, and an unknown situation (Hero task).
  • a thumbs-up reply with or without an open ask (noop, no model call);
  • delay risk that flaps (high, low, high or through medium) keeps one chain on schedule, and risk that stays low closes it at the next follow-up;
  • a reply that names a stop resolves only that stop’s ask, and asks for stops already passed close;
  • a broker “stand down” message reaches a Hero, whose override ends the chain; an unknown sender gets one Hero task per distinct message per 30 minutes.
PathContents
docs/assignment.mdThe assignment
docs/executive-summary.mdOne-page summary: decisions, built versus designed, what was skipped, risks
docs/rollout-and-metrics.mdSuccess metrics, staged rollout with gates, kill switch, ownership
docs/design.mdFull system design, built versus designed, diagrams
docs/answers.mdAnswers to the assignment questions
docs/adr/Architecture decision records
src/load_agent/Agent code: domain models, event store and projection, SOP repository and compiler, pipeline, dispatcher, timers
sops/SOPs in Markdown with their compiled specs and message templates, layered: standard, client, client + load type
evals/Scenario files, fixtures and scenario tests
src/load_agent/api/, src/load_agent/service/HTTP ingestion API, worker loop and Postgres wiring
Dockerfile, docker-compose.yml, scripts/demo.shRun as a service and the demo
tests/Unit and contract tests
Prepared for Freight Hero by Marcus Caum Source on GitHub