Skip to content

ADR 0002: SOP layering per topic and compilation of the procedure at edit time

Status: accepted (issue #1)

  • Each client has about 100 pages of plain-English Markdown SOPs, one file per topic. 60 to 70% is shared (standard); the rest differs per client, sometimes per load type.
  • Non-technical people (Heroes, customer teams) write and change SOPs. An AI engineer must not be needed for a change, and a change must not silently break other behavior.
  • Clients change procedures often in their first weeks; new clients arrive frequently.
  • Much of an SOP is procedure (Client A: SMS, wait 30 minutes, SMS again, email the dispatcher, Hero task). Much is soft guidance (“call 1 hr before arrival”, tone, what to mention) that does not fit a state machine.
  • Cost and predictability limits from ADR 0001 apply: retry and attempt counting must not depend on the model.

1. Layering per topic, most specific file wins

Section titled “1. Layering per topic, most specific file wins”
sops/standard/delay_risk_high.md
sops/clients/client_a/delay_risk_high.md # overrides standard for client_a
sops/clients/client_b/reefer/departed_pickup.md # client + load type

For a topic, resolution order is client + load_type > client > standard. The winning file replaces the topic as a whole; there is no merge of sections across layers.

Scope keys come from load metadata:

  • client_id is LoadMetadata.broker.id: the broker is our client. The assignment’s example metadata uses "broker": {"id": "northline", ...}; that is only an example value. Our scenario for load #481207 runs under Client A, so fixtures use broker.id: client_a, and the SOP directory is sops/clients/client_a/. A real deployment would use the broker’s stable id (northline) as both.
  • load_type is an explicit metadata field (dry_van, reefer, …), added to the assignment’s schema because routing must not depend on parsing the free-text commodity.

SopRepository.resolve(topic, client_id, load_type) returns the winning CompiledSop with its SopRef (topic, layer, client, load type, version).

2. Compile the procedure, keep the guidance as text

Section titled “2. Compile the procedure, keep the guidance as text”

On save, the SopCompiler (sop/compiler.py) compiles the Markdown file into a CompiledSop (domain/sop_spec.py). This is a production step, not a build script:

  1. LLM call (strong model, structured output against the CompiledSop schema).
  2. Schema validation; failures are schema_invalid.
  3. Semantic checks the schema cannot express. Ambiguity is rejected, never defaulted: ambiguous_step (“if the driver doesn’t answer soon”), unknown_contact_role (“text the lumper”), forbidden_role_used (a step contacts a role the same SOP forbids), unsupported_action, missing_trigger, conflicting_steps, forbidden_role_not_declared (the text says never to contact a role but the spec does not record it). Grounding checks compare the spec with the Markdown: every wait must appear in the text, attempt counts above one must be stated in the Procedure section, and “never contact” sentences must be recorded. They are heuristics and fail safe: when in doubt the compile is rejected and the editor rewords; they do not prove the model read every sentence correctly, which is what the editor’s preview approval is for.
  4. The result is a CompileResult: CompileSucceeded (spec with the next immutable version, source_sha256, and a deterministic preview) or CompileRejected (one or more CompileIssues with the heading and excerpt, written for the editor).

What the compiled spec contains:

  • Procedure, compiled and executed deterministically: trigger, conditions, ask (type, target role, request step, max_attempts, retry_after, resolution), ordered escalation steps (action, target role, channel, email thread, Slack audience, intent, urgent, wait_after), one-shot steps (alone, or run before the ask is opened), and forbidden_contact_roles.
  • Guidance, kept as text: guidance snippets with source_file, heading and verbatim text. They are never executed. The LLM receives the relevant snippets as context when drafting a message or handling a situation the procedure does not cover. A topic can be guidance only (no trigger).

This avoids forcing every sentence into a schema. Only what must be exact (waits, counts, order, who to contact, when an ask counts as answered) becomes structure; judgment stays as the editor wrote it. Message bodies are never stored in the spec, only an intent; the wording comes from the layered message templates for that intent, filled with load facts in code (ADR 0004). The LLM drafts a body, with guidance and verified facts, only for an intent that has no template.

Resolution is channel- and contact-agnostic by default: resolution: {requires: [eta]} means an ETA from any contact of the load on any channel resolves the ask (the dispatcher’s email resolves the driver SMS ask).

Routing is deterministic: the compiler builds a routing index (trigger to topics), so SopRepository.topics_for(trigger, client_id, load_type) is a lookup, not a vector search. Editors never write frontmatter. Unmapped situations go to triage and then the planner, whose conservative default is a Hero task.

3. Example: Client A, delay risk high (compiled)

Section titled “3. Example: Client A, delay risk high (compiled)”

This is a condensed form of the shipped sops/clients/client_a/delay_risk_high.v1.compiled.yaml (same fields and values; defaults such as wait_after: PT0S and the summary text are shortened, and the second guidance snippet is omitted). Run uv run load-agent sop preview client_a delay_risk_high for the full rendering.

ref: {topic: delay_risk_high, layer: client, client_id: client_a, version: 1}
summary: >-
When delay risk becomes high: text the driver for an ETA. If no ETA in 30 minutes, text again.
If still no ETA 30 minutes after the second text, email the dispatcher in the main thread,
then create an urgent Hero task to call the driver.
source_sha256: f032eba23224ee014e35ef6277814e9ebe81fc8dc3ea42df89768abb5fd47d56
trigger: {kind: delay_risk_became, level: high}
conditions:
- {kind: no_open_ask, ask_type: eta_request}
ask:
type: eta_request
target_role: driver
request: {action: send_sms, target_role: driver, channel: sms, intent: request_eta}
max_attempts: 2
retry_after: PT30M
resolution: {requires: [eta]}
escalation:
- action: send_email
target_role: dispatcher
channel: email
email_thread: main
intent: driver_unresponsive_request_eta
- {action: create_hero_task, intent: call_driver_for_eta, urgent: true}
guidance:
- source_file: sops/clients/client_a/delay_risk_high.md
heading: Tone
text: Keep texts to the driver short. Always include the load number.

The urgent: true on the Client A Hero task is a deliberate interpretation: the assignment says only “create a Hero task”, and by that point the driver and the dispatcher have both failed to answer while the delay risk is high, so the task is marked urgent. The standard SOP uses normal priority.

Client B overrides the same topic: one SMS, then a Slack alert to the broker, never the dispatcher.

ref: {topic: delay_risk_high, layer: client, client_id: client_b, version: 1}
summary: >-
When delay risk becomes high: text the driver once for an ETA. If no ETA in 30 minutes,
alert the broker in Slack. Never contact the dispatcher.
source_sha256: fccdbc78f8ac7d5a7868f8515f0c22cbe64741a66f1c28bf41d224c954bd0491
trigger: {kind: delay_risk_became, level: high}
conditions: [{kind: no_open_ask, ask_type: eta_request}]
ask:
type: eta_request
target_role: driver
request: {action: send_sms, target_role: driver, channel: sms, intent: request_eta}
max_attempts: 1
retry_after: PT30M
resolution: {requires: [eta]}
escalation:
- {action: post_slack, slack_audience: broker, intent: eta_missing_alert}
forbidden_contact_roles: [dispatcher]

Client B reefer adds a topic at the client + load_type layer, fired by the stage diff at_pickup to to_delivery (see ADR 0001, trigger detection):

ref: {topic: departed_pickup, layer: client_load_type, client_id: client_b,
load_type: reefer, version: 1}
summary: When the driver leaves the pickup, ask for trailer temperature and seal number.
source_sha256: 3b20b75db00a0a295c6a1f2d5c8624df1d8de788841b3427d8ec61da58514d43
trigger: {kind: stage_entered, stage: to_delivery, from_stage: at_pickup}
ask:
type: temperature_and_seal_request
target_role: driver
request:
action: send_sms
target_role: driver
channel: sms
intent: request_temperature_and_seal
max_attempts: 1
retry_after: null
resolution: {requires: [trailer_temperature, seal_number]}

Every publish creates a new immutable version. An open ask stores the SopRef (with version) it was opened under, and its follow-ups and escalation use SopRepository.get(ref), not the latest version. An ask started under v3 finishes under v3; new triggers use the current version. This prevents a mid-load edit from, for example, shortening a wait that already started or changing who gets escalated in the middle of a chain.

The editor UI is out of scope for the take-home. The design:

  1. Draft: the editor changes Markdown.
  2. Compile: SopCompiler.compile returns CompileSucceeded or CompileRejected with reasons the editor can act on.
  3. Preview: the editor sees a plain-English rendering of the compiled procedure (“When delay risk becomes high: 1) text the driver for an ETA; 2) if no ETA in 30 minutes, text again; …”) and approves that, not the YAML. The preview is rendered deterministically from the spec (sop.preview.render_preview), not written by the model, so what the editor approves is exactly what the runtime executes.
  4. Behavior diff: the client’s scenario evals plus the standard evals run against old and new versions; the editor sees which decisions changed (“at 08:01 the agent now posts to Slack instead of emailing the dispatcher”), not a text diff.
  5. Shadow mode on live loads (decide, do not act, compare), then publish (SopPublisher.publish, which refuses a stale version number when two editors publish at once). One-click rollback republishes an old spec as the new latest version.

A CLI command sop preview (plain-English rendering of the resolved, compiled spec for a client and load type, plus render_behavior_diff between versions) is built (load-agent sop preview <client> <topic> [--load-type T], issue #4) and is the working piece of this design.

AlternativeTrade-off
LLM reads the raw Markdown on every eventMost flexible, nothing to compile. Every event pays for long SOP context; behavior can change with phrasing or model version; retry and attempt counting by the model is fragile; hard to eval step by step.
Start with the LLM reading raw Markdown, compile later when cost bitesFastest to ship and fine at low volume. At 100k loads per month and $1 per load it does not fit, and the fragile part (counting retries and attempts in the model) is exactly the part the assignment scenario exercises. Migrating later means re-validating every client’s behavior anyway.
Compile everything, including guidance, into a schemaFully structured, but most sentences are judgment and do not map to fields; the schema would grow without end and the compiler would invent structure.
Editors write YAML or a DSL directlyNo compile step, exact. Violates the requirement that non-technical people edit SOPs.
Vector retrieval of SOP chunks per eventHandles unseen phrasing. Non-deterministic, can mix chunks from different layers, and cannot guarantee that the client override wins over the standard.
Section-level merge across layers (client patches part of a standard topic)Smaller client files. Merge semantics are hard for editors to predict (“which step 2 applies?”). Whole-topic override is easier to preview and to diff; the cost is duplicated text when a client changes one line.
  • The runtime needs a compiled, approved spec per topic before a client goes live; a failing compile blocks publish, it never reaches production.
  • Evals are organized by layer: standard evals run for every client that does not override a topic, client evals for overrides. A client change runs both.
  • The compile step is a model call per SOP save, not per event; its cost is negligible against runtime volume.
  • Closed vocabularies (InfoField, step actions, trigger kinds) need a code change to extend. Topics, intents, ask types, load types and client ids are open strings, so editors can add them without code.
  • Whole-topic override duplicates shared text in client files; the compiler can flag drift from standard as a review hint.
Prepared for Freight Hero by Marcus Caum Source on GitHub