ADR 0004: Templated messages, layered like SOPs
Status: accepted (issue #5)
Context
Section titled “Context”- A compiled SOP stores an
intentper step, never a body. Something must turn “request_eta” into text for a driver, an email for a dispatcher, a Hero task body. - Messages reach drivers, dispatchers and brokers on behalf of a client. A wrong PO number, an invented time or a made-up address in one of them is a business incident, not a quality blemish.
- Budget: about $1 of LLM inference per load. Most outbound messages repeat the same sentence with a few load facts filled in.
- SOP editors are not engineers. Wording must change without a code release, and differ per client and per load type, the way procedure does (ADR 0002).
Decision
Section titled “Decision”Every intent the SOPs and the pipeline use has a template: a sentence with {placeholders} ({load_id}, {stop_name}, {po_number}, …) that code fills from load metadata and the event. The LLM drafts a message only for an intent with no template, and then every fact named for verification must appear verbatim in the draft, or no message is sent and a Hero task is created.
flowchart LR
S[SOP step: intent] --> T{Template for intent?}
T -- yes --> R[Fill facts in code]
T -- no --> D[LLM draft with facts and guidance]
D --> V{Facts verbatim?}
V -- no --> H[Hero task, nothing sent]
V -- yes --> A[Action]
R --> A
Templates are data deployed with the SOP tree, not code in the package:
sops/standard/templates.v1.yamlsops/clients/<client>/templates.v1.yamlsops/clients/<client>/<load_type>/templates.v1.yamlThey layer per intent, like SOPs: client + load_type > client > standard. A file lists only the intents (and urgency entries for agent escalations) it overrides. Each file is versioned (templates.v<N>.yaml, latest wins), and the repository is read when the worker starts, as for SOPs, so changing wording is a reviewed YAML change published with the SOP tree and needs no package release. A missing fact in a template fails loudly (TemplateError, which becomes a Hero task) instead of sending a message with a hole in it.
Agent escalations (a Hero task for an unclear message, a failed delivery) use code-built text with the offending message quoted verbatim, because they are the situations where the model failed or cannot be trusted. Only their urgency comes from the template files.
Alternatives considered
Section titled “Alternatives considered”| Alternative | Why not |
|---|---|
| LLM drafts every message, facts verified afterwards | Flexible wording, but a model call per message (cost and latency), nondeterministic text in evals, and verification catches a lost fact but not an added claim (“your load is already late”). |
| Templates only, no LLM fallback | Safest, but a client adding an intent in an SOP would need an engineer before the SOP works. The fallback keeps SOP editing self-service, with verification and a Hero task as the safety net. |
| Templates in the Python package | Simplest to load. Wording changes would need a release, and per-client wording would need code. |
| Bodies stored in the compiled SOP | Couples wording to procedure versions: fixing a typo creates a new SOP version, and open asks pin versions. Templates version independently. |
| Template engine with logic (Jinja) | More expressive. Editors would write conditions, which is procedure and belongs in the SOP; plain {placeholders} keep wording reviewable by anyone. |
Trade-offs
Section titled “Trade-offs”- Gains: no invented values by construction for templated intents, zero LLM cost for most outbound messages (the 481207 timeline makes one LLM call, for classifying the driver’s question), deterministic evals, per-client wording as data.
- Templates resolve to the latest version at send time, while an open ask keeps the SOP version it started under (ADR 0002). So a client can edit a template mid-load and the next follow-up of an ask pinned to an older SOP version uses the new wording. This is acceptable because a template carries wording only: what is sent, to whom and when comes from the pinned SOP, the same facts are filled by code, and the intent name is unchanged. A wording change that should not reach open asks is a new intent name in a new SOP version. Alternative: pin the template version to the ask, which would also freeze typo fixes and make every wording edit a procedure release.
- Costs: wording is less adaptive than a model’s (no tone matching to the incoming message), and a template covers one language. Per-client overrides can drift from the standard text; the layering rule makes the winner predictable but reviewers must read the right file.
Consequences
Section titled “Consequences”AgentDeps.templatesis aTemplateRepositorybuilt from the samesops/directory as the SOP repository;for_load(client_id, load_type)returns the merged view.- Evals no longer script drafts for templated intents, so an accidental fall back to the model shows up as an unscripted LLM call.
- A test asserts every intent used by the shipped SOPs has a template.
- The planner (ADR 0001) may not write free text to drivers or dispatchers for the same reason: only code-filled or verified text reaches people outside the company.