Design notes (brainstorm summary)
Superseded. Pre-implementation brainstorm, kept for history.
docs/design.mdand ADRs 0001 to 0005 replace it wherever they differ: timers live in Postgres (not Temporal or EventBridge), triage runs on the TypeSafe decision model with deterministic value extraction (not a small generative model or rules), and evals are in a flatevals/scenarios/folder.
Decisions and reasoning agreed before implementation. These became the ADRs and docs/design.md. If something here conflicts with the assignment, the assignment wins.
1. The cost math drives the architecture
Section titled “1. The cost math drives the architecture”- 100,000 loads per month times 50 to 100 inbound messages is about 7.5M messages per month, plus location pings (many more).
- Budget is about $1 of LLM inference per load. If every message hit an LLM, that is about $0.013 per event. A single large agent reading full history on every event does not fit.
- Therefore a funnel:
- Deterministic filter (zero cost): duplicate events, pings with no delay-risk band change, timers for resolved asks, stale events.
- Cheap triage (small model, short cached prompt, or rules): does this message resolve an open ask, ask a question, carry new info, or is it noise like “ok”?
- Planner (stronger model, rare): only when a decision is not covered by the compiled SOP.
- This answers both the cost question and the 5 minute latency question: most events finish in milliseconds.
2. Separate “what the SOP says” from “judgment”: compile SOPs at edit time
Section titled “2. Separate “what the SOP says” from “judgment”: compile SOPs at edit time”- Client A’s procedure (retry after 30 min, count attempts, escalate in sequence) is a state machine, not judgment. Letting an LLM “remember” it sent two SMS is fragile and expensive.
- SOP editors (non technical) write plain English Markdown, as the assignment requires.
- On save, a strong model compiles the Markdown into a structured spec: trigger, conditions, steps, timers, max attempts, escalation chain, resolution condition.
- The editor sees a plain English preview of what was understood (“When delay risk becomes high: 1) SMS driver asking for ETA; 2) if no ETA in 30 min, repeat; …”) and approves it.
- At runtime the spec executes deterministically. The LLM only classifies messages, drafts text, and handles what the spec does not cover.
- Alternative (document it): LLM reads raw Markdown on every event. More flexible, less predictable, more expensive, harder to test.
3. Memory: append-only event log plus open asks
Section titled “3. Memory: append-only event log plus open asks”- Event log per load is the source of truth: every input (sms, email, slack, ping, timer, attachment) and every agent action.
- Load state is a projection of the log. Key part is
open_asks:
open_asks: - id: ask_eta_01 type: eta_request target: drv_20931 reason: delay_risk_high sop: client_a/delay_risk_high@v3 attempts: 1 max_attempts: 2 next_timer: tmr_88 at 07:31 resolved_by: null- 07:12 / 07:13: driver asks for the PO. Triage says “question, no ETA”. The agent answers from load metadata (
instructionsfield has PO 55821).ask_eta_01stays open because nothing resolved it. Memory is state, not model recall. - Timers carry
{load_id, ask_id, attempt}. When one fires the agent sees the current state plus that ask. Ifresolved_byis set (for example the dispatcher emailed an ETA at 07:20), the timer is a noop. Cross-channel resolution is a good extra eval. - “ok” with no open ask: nothing to match, noop. This is the required “do nothing” eval and falls out of the model naturally.
4. SOP organization and retrieval
Section titled “4. SOP organization and retrieval”Layered like CSS cascade, per topic:
sops/standard/delay_risk_high.mdsops/clients/client_a/delay_risk_high.md # overrides standardsops/clients/client_b/reefer/departed_pickup.md # client + load typeResolution: client + load_type > client > standard.
Retrieval is deterministic routing first (event type + stage + load attributes to topic), not vector search. The routing index is generated by the compile step, so editors never touch frontmatter. Semantic fallback only for unmapped situations, and the conservative default there is escalate to a Hero, not improvise.
Pros: predictable, cheap, testable. Cons: a new situation needs a new topic; the fallback only partly covers it.
5. Safety for non technical SOP editors
Section titled “5. Safety for non technical SOP editors”Draft, compile, behavior diff (not just text diff), run that client’s evals plus standard evals, shadow mode on live loads (agent decides but does not act, compare), publish. Versioned, one click rollback.
Open question to discuss in docs: SOP changes mid load. Proposal: pin the SOP version per open ask (an ask started under v3 finishes under v3); new triggers use the current version.
6. Overall architecture
Section titled “6. Overall architecture”Inbound (SMS / email / Slack / ping / timer) -> normalizer -> queue partitioned by load_id (serial per load) -> append to event log -> update state -> deterministic filter -> triage -> SOP executor / planner -> action validator (contacts of this load only, rate limit per contact, dedupe) -> executors (mocks) + action loggedTimers: durable scheduler (Temporal, or SQS delay / EventBridge Scheduler on AWS)Alternatives to document: single ReAct agent with tools (simple, expensive, unpredictable); pure workflow engine with no LLM (cannot handle language); multi agent (overkill here). Ours is the hybrid.
Details worth showing: debounce (driver sends three SMS in a row, process together), idempotency on every action, serial processing per load (FIFO message group = load_id).
7. Documents (BOL, POD, rate confirmation, receipts)
Section titled “7. Documents (BOL, POD, rate confirmation, receipts)”Classify with a cheap vision model, extract with a schema per document type, cross-validate against load metadata (PO matches? weight matches? load id?). Low confidence or mismatch becomes a Hero task. Cross-validation is what makes extraction trustworthy.
8. Evals
Section titled “8. Evals”- Scenario replay with a fake clock. Each timeline event is a step; assert expected actions (type, recipient, channel) or explicit noop.
- Message content checked structurally (contains the PO?). LLM-as-judge only where needed.
- Organization:
evals/standard/runs for every client that did not override a topic;evals/clients/<id>/for overrides. A client SOP change runs standard plus that client. - Model change: full suite plus replay of a production sample, compare decisions. Canary rollout.
- Production: traces, escalation rate, Hero override rate, anomaly alerts.
- Extra scenarios: ETA via dispatcher, driver gives an ETA that is still late, “ok” with no question, out of order message, duplicate timer.
9. Language and stack
Section titled “9. Language and stack”Python. Freight Hero’s AI Engineer posting lists Python with Pydantic and FastAPI, so matching their stack lowers friction for reviewers. TypeScript was considered (better typing, Zod) but rejected: off stack and weaker LLM/eval tooling.
LLM behind an interface with a recorded/fake mode so evals run without an API key.