Skip to content

Success metrics and rollout

How we know the agent is worth running, and how it reaches production without hurting a client. Context (assumption): Freight Hero’s Hero team does what the agent cannot (assignment). The value claim we assume is Hero time saved per load and faster issue detection, so the metrics start there.

Rules for the numbers: a figure is either cited from the repository, labelled target (a goal we propose and Freight Hero should set), or labelled assumption. Nothing here is measured production data; there is none.

Every metric is a query over the event log and the decision records (DecisionRecord: path, SOP ref, ask transitions, actions, llm_usage, run_latency_ms). Fields column: exists means the field is in the code today; designed means it needs work not in this repository.

#MetricDefinitionComputationFieldsTarget
1Hero minutes per loadSum of Hero handling time per load over Hero tasks and overrides, counting a task and the override that closes it as one touchCount create_hero_task actions and AskOverride events per load; multiply by handling time (task opened to closed)Touch counts: exists (CreateHeroTask, AskOverride). Handling time: designed (needs Hero tooling to report task open and close)Below the pre-agent baseline for the same client, measured in shadow. The baseline is not known yet.
2Escalation rate by reasonHero tasks per 100 inbound events, split by reasonGroup Hero-task actions by reason: unknown situation (path = escalated_unknown), low-confidence triage, failed delivery (ActionOutcome failed), document mismatch, SOP escalation stepPath and ActionOutcome: exists. A structured reason code on the task: designed (today the reason is free text in the task title and instructions)No global target. Watch the trend per client and topic; a spike after an SOP publish is a stop signal (section 2).
3Wrong-action rateShare of agent actions a Hero reverses or overridesExplicit “agent was wrong” marks, plus AskOverride events whose reason carries a structured agent_wrong code. A plain override after a Hero completes an escalation is a Hero touch (metric 1), not an errorAskOverride with source, actor, reason: exists (closes asks only). The “agent was wrong” mark, reversal of sent messages and a structured override reason code: designedTarget near zero for external actions (SMS, email); the stage gates in section 2 set the numeric limits.
4Silent-miss rateIssues a Hero found that the agent did notHero-created tasks, or issues a Hero reports, with no agent decision on that load and topic in the preceding windowNeeds Hero-originated tasks in the log with a link to the load: designed. Ground truth comes from a stratified random audit: a fixed number of loads per client per week (proposed: 20), fully reviewed by a Hero against the SOP, with a minimum sample before any rate gates a stage. Report every metric per client and per topic, never only globally. Sampling audits of filtered decisions can start from existing recordsTarget: lower than the pre-agent rate of missed issues, measured in shadow. Highest-priority metric, because a silent miss is invisible to every other metric.
5Time to detect a delayTime from the first ping that carries the high risk band to the first committed actiondecided_at of the first action-bearing decision minus the ping’s event timeBoth exist. A stronger version (from the true moment the load became late) needs actual arrival times from stop_status events and is designedWithin the 5-minute run budget of the assignment, plus dispatch.
6LLM cost per loadSum of model cost over the loadSum llm_cost_usd over the decision records of the load, by purpose and modelLLMUsage.cost_usd, tokens, purpose: exists. Prices default to 0 until configured (README), so a real number needs prices setAt or below $1 per load (assignment budget). Design estimate: about $0.04 to $0.46, an assumption-based range in design.md section 13; warn at $0.60, ceiling $2. Replace with the measured value from benchmark.md and replay.
7p95 run latency95th percentile of run start to commit (run_latency_ms). Event received to run start (queue wait) and commit to delivery are designed additionsPercentile of run_latency_ms over runs, by pathField per record: exists. Aggregation and export: designedAt or below 5 minutes (assignment). The design pages on p99 above 60 s. Both are targets.

Supporting metrics, same sources: path mix (planner share rising means missing SOP coverage), noop share (a sudden drop means the filter broke), timer pump lag, outbox age. Listed in design.md section 12.

Trade-off: a metric from the log needs no extra instrumentation and cannot disagree with the audit trail. The alternative, a separate analytics pipeline, would let a dashboard drift from what the agent did. Cost: wrong-action and silent-miss depend on Hero feedback, which is human-labelled and sparse, so they need sampling audits to be trusted.

flowchart LR
  A["1. Offline replay<br/>client history"] -->|"Gate 1"| B["2. Shadow<br/>agent logs, Heroes act"]
  B -->|"Gate 2"| C["3a. Canary: internal actions<br/>Hero task, TMS note, Slack"]
  C -->|"Gate 3"| D["3b. Canary: driver SMS"]
  D -->|"Gate 4"| E["3c. Canary: dispatcher email,<br/>broker Slack"]
  E -->|"Gate 5"| F["4. Full rollout"]
  C -.->|"gate fails"| B
  D -.->|"gate fails"| C
  E -.->|"gate fails"| D
  F -.->|"kill switch or SOP change"| B

Order of exposure follows the cost of a mistake: an internal task is read by a Hero before anything happens; a driver SMS reaches a person on the road; an email reaches the dispatcher in the main thread and is visible to the carrier. Alternative: canary by load percentage only. It is simpler but exposes every action type at once, so one bad template or threshold reaches drivers and dispatchers together. Cost of ours: slower, because there are three canary steps per client.

Gates that use metrics 1, 3 and 4 need the Hero feedback fields marked designed in section 1. Until those exist, substitute a manual weekly review of a fixed random sample of loads per client.

All thresholds below are proposed targets to be agreed with Freight Hero and the client. They are starting points, not results.

StageWhat runsGo criteria to leave the stageNo-go (stay or step back)
1. Offline replayThe client’s historical loads replayed through the pipeline with FakeClock; triage answers recorded or live. Nothing is sent.Every step of the replayed loads produces a valid decision; zero crashes; cost per load within the $1 budget on the replay, with the LOAD_AGENT_*_USD_* prices set from contract rates (a $0 replay does not satisfy this gate); triage accepts-and-agrees rate on a labelled sample meets the bar set from benchmark.md (a held-out production set is required first, see ADR 0005); every SOP topic the client uses has a compiled specA topic with no spec carrying real volume, or any wrong extracted value in the labelled sample
2. Shadow, per clientLive events flow in; the agent decides and writes DecisionRecords but its actions are not dispatched. Heroes keep acting as today. The shadow agent writes to a separate shadow log and projection. Hero actions on the load are ingested as events, so shadow state reflects what was really sent, and agreement compares the shadow decision with the Hero’s action in the same window.Agreement on action-bearing decisions at or above a target set per client (proposed: 95%), with every disagreement reviewed; zero decisions that would have contacted a person not on the load (validation already blocks these, so this checks the gate); silent-miss rate no worse than the Hero baseline; kill switch exercised end to end in production (flip, observe Hero task replacement, lift)A disagreement class that repeats (fix the SOP or the threshold, restart the stage for that topic)
3a. Canary, internal actionsHero tasks, TMS notes, internal Slack posts are live on a subset of loads (proposed: 10% of the client’s loads, chosen by hash of load_id so a load never switches mid-way)Wrong-action rate at or below a target for 2 weeks of loads (proposed: under 2%); Hero feedback shows tasks carry enough context to actHero tasks rejected as noise: raise thresholds or fix the SOP first
3b. Canary, driver SMSDriver SMS live on the same subsetZero wrong-recipient or duplicate sends (both are structurally blocked, so any case is a bug); wrong-action rate under target, and on driver SMS under a separate external target (proposed: 0.5%, minimum 200 sends before the gate is read); driver replies resolve asks as expectedAny duplicate send, or any message with a wrong fact (stop, PO, ETA)
3c. Canary, dispatcher email and broker SlackDispatcher email in the main thread and broker Slack posts liveSame as 3b; no client complaint attributable to the agentAny off-thread email or contact the SOP forbids
4. Full rolloutAll loads of the client, all action typesCanary stable for the agreed periodn/a

Designed, not built. A switch per client and per action type (for example “client_a, send_sms: off”), checked by the action validator that already rejects bad actions (agent/validation.py), so it applies after the decision and before the outbox. A blocked action is replaced by a Hero task carrying the action the agent wanted to take (“send this SMS by hand”); the ask and its timers stay in state, so nothing is sent twice when the switch is lifted. The dispatcher checks the switch too, so rows committed before the flip are held and turned into Hero tasks. A blocked send does not advance the ask’s attempt count, but its follow-up timer still fires, so the Hero task carries the step and the chain stays on schedule. Alternative: a global pause. Simpler, but one client’s problem stops every client. A per-load “pause the agent” flag, also designed in design.md section 5, covers a single bad load. Rollback of a bad SOP is a republish of the previous version; asks pinned to a version keep finishing under it (ADR 0002).

Triggers that flip a switch without waiting for a meeting: any duplicate or wrong-recipient send, escalation rate for a client and topic doubling against its trailing average after a publish (proposed rule), outbox dead rows rising, or a provider incident.

The publish flow (designed, design.md section 8) runs the behavior diff and that client’s evals. The changed topic then re-enters shadow for that client and topic while every other topic stays live: the new spec decides and logs, the previous version keeps acting, and agreement is reviewed. When the gate for that topic passes, the new version goes live; canary applies again if the change adds a new action type. Asks already open stay on the old version. Alternative: publish straight to live after evals pass. Faster, but evals only cover scenarios we thought of, and the SOP author thought of a case we did not. Cost: a change takes effect after a shadow period instead of at once.

AreaOwner and practice
SOP editsHeroes and customer teams write Markdown; no engineer is needed to publish (assignment requirement). They approve the preview and the behavior diff. Engineers own the platform, standard templates and the compiler.
EvalsEngineers own the framework. Every published SOP change adds or updates a scenario for its topic.
On-callEngineering on-call is paged by outbox age, dead rows, timer pump lag, run latency, and a kill-switch trigger. Hero operations lead owns quality signals (escalation rate, override rate).
Kill switchEither on-call may flip it; the Hero lead decides when to lift it.

What a Hero sees in a task (today the task has a title, instructions, urgency and the related contact; the context and reason fields below and the one-click override are designed):

  • Context: load, stop, contacts, the open ask and its attempt count, the events and actions of the load in order (from the log).
  • Reason: why the agent escalated (SOP step reached, unknown situation, failed delivery, low confidence) and the SOP version followed (sop_ref).
  • Suggested next step: for example “Call the driver on load 481207 to get the ETA to Riverbend plant. No ETA arrived after our requests.” (real output of the scenario).
  • One-click override: close the ask (“I called the driver, stand down”), which emits an AskOverride with the Hero’s id and a reason, and “agent was wrong”, which feeds metric 3.

Feedback loop. Every override and every “agent was wrong” mark is reviewed weekly (proposed cadence). Each confirmed case becomes a scenario in evals/scenarios/ with the real inputs (names and phone numbers removed) and the expected actions, written before the fix (the repository’s TDD rule). A triage misread is added to the recorded phrasing set that re-fits the thresholds. Alternative: retrain on feedback automatically. Cost of ours: it is manual and depends on Hero discipline; it gains that a change is reviewed and tested before it reaches a driver.

This is the scenario from the assignment (design.md section 6), run offline with uv run load-agent run evals/scenarios/client_a_load_481207_full_timeline.yaml. It shows which steps need no Hero. It is not measured data. What a Hero would otherwise do at each step is an assumption about the current process.

TimeStepAgent actionHero touch
07:01Delay risk turns highSMS to the driver asking for the ETA; follow-up timer set for 07:31None
07:13Driver asks for the PO numberSMS with the PO from load metadata (55821); the ETA ask stays openNone
07:31No ETA yetSecond SMS to the driver; timer set for 08:01None
08:01Two requests unansweredEmail to the dispatcher in the main thread; urgent Hero task “call the driver”First touch: the Hero reads the task and calls

The Hero’s first touch is the call at 08:01, the one step that needs a human. The 07:01 text, the PO answer and the 07:31 follow-up (three of the four agent steps) happened without anyone. Had the dispatcher emailed the ETA at 07:20, the 07:31 timer would be a noop (timer_after_eta_received_is_noop) and no Hero would be involved at all. The same run produced one model call (the PO question) in the whole timeline.

What this does and does not show: it shows the hand-offs the design makes; it does not show how many minutes a Hero spends on such a load today, nor how often drivers reply. Those are the baseline measurements in stage 2.

Prepared for Freight Hero by Marcus Caum Source on GitHub