ADR 0002: SOP layering per topic and compilation of the procedure at edit time
Status: accepted (issue #1)
Context
Section titled “Context”- Each client has about 100 pages of plain-English Markdown SOPs, one file per topic. 60 to 70% is shared (standard); the rest differs per client, sometimes per load type.
- Non-technical people (Heroes, customer teams) write and change SOPs. An AI engineer must not be needed for a change, and a change must not silently break other behavior.
- Clients change procedures often in their first weeks; new clients arrive frequently.
- Much of an SOP is procedure (Client A: SMS, wait 30 minutes, SMS again, email the dispatcher, Hero task). Much is soft guidance (“call 1 hr before arrival”, tone, what to mention) that does not fit a state machine.
- Cost and predictability limits from ADR 0001 apply: retry and attempt counting must not depend on the model.
Decision
Section titled “Decision”1. Layering per topic, most specific file wins
Section titled “1. Layering per topic, most specific file wins”sops/standard/delay_risk_high.mdsops/clients/client_a/delay_risk_high.md # overrides standard for client_asops/clients/client_b/reefer/departed_pickup.md # client + load typeFor a topic, resolution order is client + load_type > client > standard. The winning file replaces the topic as a whole; there is no merge of sections across layers.
Scope keys come from load metadata:
client_idisLoadMetadata.broker.id: the broker is our client. The assignment’s example metadata uses"broker": {"id": "northline", ...}; that is only an example value. Our scenario for load #481207 runs under Client A, so fixtures usebroker.id: client_a, and the SOP directory issops/clients/client_a/. A real deployment would use the broker’s stable id (northline) as both.load_typeis an explicit metadata field (dry_van,reefer, …), added to the assignment’s schema because routing must not depend on parsing the free-textcommodity.
SopRepository.resolve(topic, client_id, load_type) returns the winning CompiledSop with its SopRef (topic, layer, client, load type, version).
2. Compile the procedure, keep the guidance as text
Section titled “2. Compile the procedure, keep the guidance as text”On save, the SopCompiler (sop/compiler.py) compiles the Markdown file into a CompiledSop (domain/sop_spec.py). This is a production step, not a build script:
- LLM call (strong model, structured output against the
CompiledSopschema). - Schema validation; failures are
schema_invalid. - Semantic checks the schema cannot express. Ambiguity is rejected, never defaulted:
ambiguous_step(“if the driver doesn’t answer soon”),unknown_contact_role(“text the lumper”),forbidden_role_used(a step contacts a role the same SOP forbids),unsupported_action,missing_trigger,conflicting_steps,forbidden_role_not_declared(the text says never to contact a role but the spec does not record it). Grounding checks compare the spec with the Markdown: every wait must appear in the text, attempt counts above one must be stated in the Procedure section, and “never contact” sentences must be recorded. They are heuristics and fail safe: when in doubt the compile is rejected and the editor rewords; they do not prove the model read every sentence correctly, which is what the editor’s preview approval is for. - The result is a
CompileResult:CompileSucceeded(spec with the next immutable version,source_sha256, and a deterministic preview) orCompileRejected(one or moreCompileIssues with the heading and excerpt, written for the editor).
What the compiled spec contains:
- Procedure, compiled and executed deterministically:
trigger,conditions,ask(type, target role, request step,max_attempts,retry_after,resolution), orderedescalationsteps (action, target role, channel, email thread, Slack audience, intent, urgent,wait_after), one-shotsteps(alone, or run before the ask is opened), andforbidden_contact_roles. - Guidance, kept as text:
guidancesnippets withsource_file,headingand verbatimtext. They are never executed. The LLM receives the relevant snippets as context when drafting a message or handling a situation the procedure does not cover. A topic can be guidance only (no trigger).
This avoids forcing every sentence into a schema. Only what must be exact (waits, counts, order, who to contact, when an ask counts as answered) becomes structure; judgment stays as the editor wrote it. Message bodies are never stored in the spec, only an intent; the wording comes from the layered message templates for that intent, filled with load facts in code (ADR 0004). The LLM drafts a body, with guidance and verified facts, only for an intent that has no template.
Resolution is channel- and contact-agnostic by default: resolution: {requires: [eta]} means an ETA from any contact of the load on any channel resolves the ask (the dispatcher’s email resolves the driver SMS ask).
Routing is deterministic: the compiler builds a routing index (trigger to topics), so SopRepository.topics_for(trigger, client_id, load_type) is a lookup, not a vector search. Editors never write frontmatter. Unmapped situations go to triage and then the planner, whose conservative default is a Hero task.
3. Example: Client A, delay risk high (compiled)
Section titled “3. Example: Client A, delay risk high (compiled)”This is a condensed form of the shipped sops/clients/client_a/delay_risk_high.v1.compiled.yaml (same fields and values; defaults such as wait_after: PT0S and the summary text are shortened, and the second guidance snippet is omitted). Run uv run load-agent sop preview client_a delay_risk_high for the full rendering.
ref: {topic: delay_risk_high, layer: client, client_id: client_a, version: 1}summary: >- When delay risk becomes high: text the driver for an ETA. If no ETA in 30 minutes, text again. If still no ETA 30 minutes after the second text, email the dispatcher in the main thread, then create an urgent Hero task to call the driver.source_sha256: f032eba23224ee014e35ef6277814e9ebe81fc8dc3ea42df89768abb5fd47d56trigger: {kind: delay_risk_became, level: high}conditions: - {kind: no_open_ask, ask_type: eta_request}ask: type: eta_request target_role: driver request: {action: send_sms, target_role: driver, channel: sms, intent: request_eta} max_attempts: 2 retry_after: PT30M resolution: {requires: [eta]}escalation: - action: send_email target_role: dispatcher channel: email email_thread: main intent: driver_unresponsive_request_eta - {action: create_hero_task, intent: call_driver_for_eta, urgent: true}guidance: - source_file: sops/clients/client_a/delay_risk_high.md heading: Tone text: Keep texts to the driver short. Always include the load number.The urgent: true on the Client A Hero task is a deliberate interpretation: the assignment says only “create a Hero task”, and by that point the driver and the dispatcher have both failed to answer while the delay risk is high, so the task is marked urgent. The standard SOP uses normal priority.
Client B overrides the same topic: one SMS, then a Slack alert to the broker, never the dispatcher.
ref: {topic: delay_risk_high, layer: client, client_id: client_b, version: 1}summary: >- When delay risk becomes high: text the driver once for an ETA. If no ETA in 30 minutes, alert the broker in Slack. Never contact the dispatcher.source_sha256: fccdbc78f8ac7d5a7868f8515f0c22cbe64741a66f1c28bf41d224c954bd0491trigger: {kind: delay_risk_became, level: high}conditions: [{kind: no_open_ask, ask_type: eta_request}]ask: type: eta_request target_role: driver request: {action: send_sms, target_role: driver, channel: sms, intent: request_eta} max_attempts: 1 retry_after: PT30M resolution: {requires: [eta]}escalation: - {action: post_slack, slack_audience: broker, intent: eta_missing_alert}forbidden_contact_roles: [dispatcher]Client B reefer adds a topic at the client + load_type layer, fired by the stage diff at_pickup to to_delivery (see ADR 0001, trigger detection):
ref: {topic: departed_pickup, layer: client_load_type, client_id: client_b, load_type: reefer, version: 1}summary: When the driver leaves the pickup, ask for trailer temperature and seal number.source_sha256: 3b20b75db00a0a295c6a1f2d5c8624df1d8de788841b3427d8ec61da58514d43trigger: {kind: stage_entered, stage: to_delivery, from_stage: at_pickup}ask: type: temperature_and_seal_request target_role: driver request: action: send_sms target_role: driver channel: sms intent: request_temperature_and_seal max_attempts: 1 retry_after: null resolution: {requires: [trailer_temperature, seal_number]}4. Version pinning per open ask
Section titled “4. Version pinning per open ask”Every publish creates a new immutable version. An open ask stores the SopRef (with version) it was opened under, and its follow-ups and escalation use SopRepository.get(ref), not the latest version. An ask started under v3 finishes under v3; new triggers use the current version. This prevents a mid-load edit from, for example, shortening a wait that already started or changing who gets escalated in the middle of a chain.
5. Editor safety: designed, not built
Section titled “5. Editor safety: designed, not built”The editor UI is out of scope for the take-home. The design:
- Draft: the editor changes Markdown.
- Compile:
SopCompiler.compilereturnsCompileSucceededorCompileRejectedwith reasons the editor can act on. - Preview: the editor sees a plain-English rendering of the compiled procedure (“When delay risk becomes high: 1) text the driver for an ETA; 2) if no ETA in 30 minutes, text again; …”) and approves that, not the YAML. The preview is rendered deterministically from the spec (
sop.preview.render_preview), not written by the model, so what the editor approves is exactly what the runtime executes. - Behavior diff: the client’s scenario evals plus the standard evals run against old and new versions; the editor sees which decisions changed (“at 08:01 the agent now posts to Slack instead of emailing the dispatcher”), not a text diff.
- Shadow mode on live loads (decide, do not act, compare), then publish (
SopPublisher.publish, which refuses a stale version number when two editors publish at once). One-click rollback republishes an old spec as the new latest version.
A CLI command sop preview (plain-English rendering of the resolved, compiled spec for a client and load type, plus render_behavior_diff between versions) is built (load-agent sop preview <client> <topic> [--load-type T], issue #4) and is the working piece of this design.
Alternatives considered
Section titled “Alternatives considered”| Alternative | Trade-off |
|---|---|
| LLM reads the raw Markdown on every event | Most flexible, nothing to compile. Every event pays for long SOP context; behavior can change with phrasing or model version; retry and attempt counting by the model is fragile; hard to eval step by step. |
| Start with the LLM reading raw Markdown, compile later when cost bites | Fastest to ship and fine at low volume. At 100k loads per month and $1 per load it does not fit, and the fragile part (counting retries and attempts in the model) is exactly the part the assignment scenario exercises. Migrating later means re-validating every client’s behavior anyway. |
| Compile everything, including guidance, into a schema | Fully structured, but most sentences are judgment and do not map to fields; the schema would grow without end and the compiler would invent structure. |
| Editors write YAML or a DSL directly | No compile step, exact. Violates the requirement that non-technical people edit SOPs. |
| Vector retrieval of SOP chunks per event | Handles unseen phrasing. Non-deterministic, can mix chunks from different layers, and cannot guarantee that the client override wins over the standard. |
| Section-level merge across layers (client patches part of a standard topic) | Smaller client files. Merge semantics are hard for editors to predict (“which step 2 applies?”). Whole-topic override is easier to preview and to diff; the cost is duplicated text when a client changes one line. |
Consequences
Section titled “Consequences”- The runtime needs a compiled, approved spec per topic before a client goes live; a failing compile blocks publish, it never reaches production.
- Evals are organized by layer: standard evals run for every client that does not override a topic, client evals for overrides. A client change runs both.
- The compile step is a model call per SOP save, not per event; its cost is negligible against runtime volume.
- Closed vocabularies (
InfoField, step actions, trigger kinds) need a code change to extend. Topics, intents, ask types, load types and client ids are open strings, so editors can add them without code. - Whole-topic override duplicates shared text in client files; the compiler can flag drift from standard as a review hint.