Success metrics and rollout
How we know the agent is worth running, and how it reaches production without hurting a client. Context (assumption): Freight Hero’s Hero team does what the agent cannot (assignment). The value claim we assume is Hero time saved per load and faster issue detection, so the metrics start there.
Rules for the numbers: a figure is either cited from the repository, labelled target (a goal we propose and Freight Hero should set), or labelled assumption. Nothing here is measured production data; there is none.
1. Success metrics
Section titled “1. Success metrics”Every metric is a query over the event log and the decision records (DecisionRecord: path, SOP ref, ask transitions, actions, llm_usage, run_latency_ms). Fields column: exists means the field is in the code today; designed means it needs work not in this repository.
| # | Metric | Definition | Computation | Fields | Target |
|---|---|---|---|---|---|
| 1 | Hero minutes per load | Sum of Hero handling time per load over Hero tasks and overrides, counting a task and the override that closes it as one touch | Count create_hero_task actions and AskOverride events per load; multiply by handling time (task opened to closed) | Touch counts: exists (CreateHeroTask, AskOverride). Handling time: designed (needs Hero tooling to report task open and close) | Below the pre-agent baseline for the same client, measured in shadow. The baseline is not known yet. |
| 2 | Escalation rate by reason | Hero tasks per 100 inbound events, split by reason | Group Hero-task actions by reason: unknown situation (path = escalated_unknown), low-confidence triage, failed delivery (ActionOutcome failed), document mismatch, SOP escalation step | Path and ActionOutcome: exists. A structured reason code on the task: designed (today the reason is free text in the task title and instructions) | No global target. Watch the trend per client and topic; a spike after an SOP publish is a stop signal (section 2). |
| 3 | Wrong-action rate | Share of agent actions a Hero reverses or overrides | Explicit “agent was wrong” marks, plus AskOverride events whose reason carries a structured agent_wrong code. A plain override after a Hero completes an escalation is a Hero touch (metric 1), not an error | AskOverride with source, actor, reason: exists (closes asks only). The “agent was wrong” mark, reversal of sent messages and a structured override reason code: designed | Target near zero for external actions (SMS, email); the stage gates in section 2 set the numeric limits. |
| 4 | Silent-miss rate | Issues a Hero found that the agent did not | Hero-created tasks, or issues a Hero reports, with no agent decision on that load and topic in the preceding window | Needs Hero-originated tasks in the log with a link to the load: designed. Ground truth comes from a stratified random audit: a fixed number of loads per client per week (proposed: 20), fully reviewed by a Hero against the SOP, with a minimum sample before any rate gates a stage. Report every metric per client and per topic, never only globally. Sampling audits of filtered decisions can start from existing records | Target: lower than the pre-agent rate of missed issues, measured in shadow. Highest-priority metric, because a silent miss is invisible to every other metric. |
| 5 | Time to detect a delay | Time from the first ping that carries the high risk band to the first committed action | decided_at of the first action-bearing decision minus the ping’s event time | Both exist. A stronger version (from the true moment the load became late) needs actual arrival times from stop_status events and is designed | Within the 5-minute run budget of the assignment, plus dispatch. |
| 6 | LLM cost per load | Sum of model cost over the load | Sum llm_cost_usd over the decision records of the load, by purpose and model | LLMUsage.cost_usd, tokens, purpose: exists. Prices default to 0 until configured (README), so a real number needs prices set | At or below $1 per load (assignment budget). Design estimate: about $0.04 to $0.46, an assumption-based range in design.md section 13; warn at $0.60, ceiling $2. Replace with the measured value from benchmark.md and replay. |
| 7 | p95 run latency | 95th percentile of run start to commit (run_latency_ms). Event received to run start (queue wait) and commit to delivery are designed additions | Percentile of run_latency_ms over runs, by path | Field per record: exists. Aggregation and export: designed | At or below 5 minutes (assignment). The design pages on p99 above 60 s. Both are targets. |
Supporting metrics, same sources: path mix (planner share rising means missing SOP coverage), noop share (a sudden drop means the filter broke), timer pump lag, outbox age. Listed in design.md section 12.
Trade-off: a metric from the log needs no extra instrumentation and cannot disagree with the audit trail. The alternative, a separate analytics pipeline, would let a dashboard drift from what the agent did. Cost: wrong-action and silent-miss depend on Hero feedback, which is human-labelled and sparse, so they need sampling audits to be trusted.
2. Rollout
Section titled “2. Rollout”flowchart LR A["1. Offline replay<br/>client history"] -->|"Gate 1"| B["2. Shadow<br/>agent logs, Heroes act"] B -->|"Gate 2"| C["3a. Canary: internal actions<br/>Hero task, TMS note, Slack"] C -->|"Gate 3"| D["3b. Canary: driver SMS"] D -->|"Gate 4"| E["3c. Canary: dispatcher email,<br/>broker Slack"] E -->|"Gate 5"| F["4. Full rollout"] C -.->|"gate fails"| B D -.->|"gate fails"| C E -.->|"gate fails"| D F -.->|"kill switch or SOP change"| B
Order of exposure follows the cost of a mistake: an internal task is read by a Hero before anything happens; a driver SMS reaches a person on the road; an email reaches the dispatcher in the main thread and is visible to the carrier. Alternative: canary by load percentage only. It is simpler but exposes every action type at once, so one bad template or threshold reaches drivers and dispatchers together. Cost of ours: slower, because there are three canary steps per client.
Gates that use metrics 1, 3 and 4 need the Hero feedback fields marked designed in section 1. Until those exist, substitute a manual weekly review of a fixed random sample of loads per client.
All thresholds below are proposed targets to be agreed with Freight Hero and the client. They are starting points, not results.
| Stage | What runs | Go criteria to leave the stage | No-go (stay or step back) |
|---|---|---|---|
| 1. Offline replay | The client’s historical loads replayed through the pipeline with FakeClock; triage answers recorded or live. Nothing is sent. | Every step of the replayed loads produces a valid decision; zero crashes; cost per load within the $1 budget on the replay, with the LOAD_AGENT_*_USD_* prices set from contract rates (a $0 replay does not satisfy this gate); triage accepts-and-agrees rate on a labelled sample meets the bar set from benchmark.md (a held-out production set is required first, see ADR 0005); every SOP topic the client uses has a compiled spec | A topic with no spec carrying real volume, or any wrong extracted value in the labelled sample |
| 2. Shadow, per client | Live events flow in; the agent decides and writes DecisionRecords but its actions are not dispatched. Heroes keep acting as today. The shadow agent writes to a separate shadow log and projection. Hero actions on the load are ingested as events, so shadow state reflects what was really sent, and agreement compares the shadow decision with the Hero’s action in the same window. | Agreement on action-bearing decisions at or above a target set per client (proposed: 95%), with every disagreement reviewed; zero decisions that would have contacted a person not on the load (validation already blocks these, so this checks the gate); silent-miss rate no worse than the Hero baseline; kill switch exercised end to end in production (flip, observe Hero task replacement, lift) | A disagreement class that repeats (fix the SOP or the threshold, restart the stage for that topic) |
| 3a. Canary, internal actions | Hero tasks, TMS notes, internal Slack posts are live on a subset of loads (proposed: 10% of the client’s loads, chosen by hash of load_id so a load never switches mid-way) | Wrong-action rate at or below a target for 2 weeks of loads (proposed: under 2%); Hero feedback shows tasks carry enough context to act | Hero tasks rejected as noise: raise thresholds or fix the SOP first |
| 3b. Canary, driver SMS | Driver SMS live on the same subset | Zero wrong-recipient or duplicate sends (both are structurally blocked, so any case is a bug); wrong-action rate under target, and on driver SMS under a separate external target (proposed: 0.5%, minimum 200 sends before the gate is read); driver replies resolve asks as expected | Any duplicate send, or any message with a wrong fact (stop, PO, ETA) |
| 3c. Canary, dispatcher email and broker Slack | Dispatcher email in the main thread and broker Slack posts live | Same as 3b; no client complaint attributable to the agent | Any off-thread email or contact the SOP forbids |
| 4. Full rollout | All loads of the client, all action types | Canary stable for the agreed period | n/a |
Kill switch
Section titled “Kill switch”Designed, not built. A switch per client and per action type (for example “client_a, send_sms: off”), checked by the action validator that already rejects bad actions (agent/validation.py), so it applies after the decision and before the outbox. A blocked action is replaced by a Hero task carrying the action the agent wanted to take (“send this SMS by hand”); the ask and its timers stay in state, so nothing is sent twice when the switch is lifted. The dispatcher checks the switch too, so rows committed before the flip are held and turned into Hero tasks. A blocked send does not advance the ask’s attempt count, but its follow-up timer still fires, so the Hero task carries the step and the chain stays on schedule. Alternative: a global pause. Simpler, but one client’s problem stops every client. A per-load “pause the agent” flag, also designed in design.md section 5, covers a single bad load. Rollback of a bad SOP is a republish of the previous version; asks pinned to a version keep finishing under it (ADR 0002).
Triggers that flip a switch without waiting for a meeting: any duplicate or wrong-recipient send, escalation rate for a client and topic doubling against its trailing average after a publish (proposed rule), outbox dead rows rising, or a provider incident.
A client changes its SOP
Section titled “A client changes its SOP”The publish flow (designed, design.md section 8) runs the behavior diff and that client’s evals. The changed topic then re-enters shadow for that client and topic while every other topic stays live: the new spec decides and logs, the previous version keeps acting, and agreement is reviewed. When the gate for that topic passes, the new version goes live; canary applies again if the change adds a new action type. Asks already open stay on the old version. Alternative: publish straight to live after evals pass. Faster, but evals only cover scenarios we thought of, and the SOP author thought of a case we did not. Cost: a change takes effect after a shadow period instead of at once.
3. Ownership and operations
Section titled “3. Ownership and operations”| Area | Owner and practice |
|---|---|
| SOP edits | Heroes and customer teams write Markdown; no engineer is needed to publish (assignment requirement). They approve the preview and the behavior diff. Engineers own the platform, standard templates and the compiler. |
| Evals | Engineers own the framework. Every published SOP change adds or updates a scenario for its topic. |
| On-call | Engineering on-call is paged by outbox age, dead rows, timer pump lag, run latency, and a kill-switch trigger. Hero operations lead owns quality signals (escalation rate, override rate). |
| Kill switch | Either on-call may flip it; the Hero lead decides when to lift it. |
What a Hero sees in a task (today the task has a title, instructions, urgency and the related contact; the context and reason fields below and the one-click override are designed):
- Context: load, stop, contacts, the open ask and its attempt count, the events and actions of the load in order (from the log).
- Reason: why the agent escalated (SOP step reached, unknown situation, failed delivery, low confidence) and the SOP version followed (
sop_ref). - Suggested next step: for example “Call the driver on load 481207 to get the ETA to Riverbend plant. No ETA arrived after our requests.” (real output of the scenario).
- One-click override: close the ask (“I called the driver, stand down”), which emits an
AskOverridewith the Hero’s id and a reason, and “agent was wrong”, which feeds metric 3.
Feedback loop. Every override and every “agent was wrong” mark is reviewed weekly (proposed cadence). Each confirmed case becomes a scenario in evals/scenarios/ with the real inputs (names and phone numbers removed) and the expected actions, written before the fix (the repository’s TDD rule). A triage misread is added to the recorded phrasing set that re-fits the thresholds. Alternative: retrain on feedback automatically. Cost of ours: it is manual and depends on Hero discipline; it gains that a change is reviewed and tested before it reaches a driver.
4. Worked example: load #481207
Section titled “4. Worked example: load #481207”This is the scenario from the assignment (design.md section 6), run offline with uv run load-agent run evals/scenarios/client_a_load_481207_full_timeline.yaml. It shows which steps need no Hero. It is not measured data. What a Hero would otherwise do at each step is an assumption about the current process.
| Time | Step | Agent action | Hero touch |
|---|---|---|---|
| 07:01 | Delay risk turns high | SMS to the driver asking for the ETA; follow-up timer set for 07:31 | None |
| 07:13 | Driver asks for the PO number | SMS with the PO from load metadata (55821); the ETA ask stays open | None |
| 07:31 | No ETA yet | Second SMS to the driver; timer set for 08:01 | None |
| 08:01 | Two requests unanswered | Email to the dispatcher in the main thread; urgent Hero task “call the driver” | First touch: the Hero reads the task and calls |
The Hero’s first touch is the call at 08:01, the one step that needs a human. The 07:01 text, the PO answer and the 07:31 follow-up (three of the four agent steps) happened without anyone. Had the dispatcher emailed the ETA at 07:20, the 07:31 timer would be a noop (timer_after_eta_received_is_noop) and no Hero would be involved at all. The same run produced one model call (the PO question) in the whole timeline.
What this does and does not show: it shows the hand-offs the design makes; it does not show how many minutes a Hero spends on such a load today, nor how often drivers reply. Those are the baseline measurements in stage 2.