Skip to content

Benchmark: cost and latency of one load over a realistic day

Generated from bench/results/2026-10-08-typesafe-deepseek.json by uv run load-agent bench --render. Do not edit by hand.

This measures the assignment’s budget: about $1 of LLM inference per load over 50 to 100 inbound messages, and at most 5 minutes per run.

  • Cost per load: $0.1100 against a budget of $1.00: within budget ($0.020019 DeepSeek, measured tokens at list prices; $0.0900 TypeSafe, 45 calls at an assumed $0.0020 per call).
  • TypeSafe break-even: the load reaches $1.00 at $0.0218 per TypeSafe call (45 calls).
  • Slowest run: 21.21 s against a limit of 300 s: within the limit. This is process_event only (no dispatch, queue wait or pump lag), and retries could push the worst case higher (see Method).
  • Inputs that reach a model: 51.7% of the 87 scripted inputs (49.5% of all 91 runs, timers included); the rest stop in deterministic code.

Time inside the agent is simulated (FakeClock): the 30 minute follow-ups fire at the right moment of the synthetic day without waiting. Channels are fakes, so no SMS or email was sent. The timer pump runs once per simulated minute.

Latency was measured from this cloud container with the network to TypeSafe and DeepSeek included. It is wall-clock time of each process_event call. Production latency depends on where the workers run and on provider load at the time, so treat it as one sample, not a guarantee.

The latency figures cover process_event only: they exclude the outbox dispatch, the queue wait and the timer pump lag. The default per-attempt timeout is 30 s with 2 retries inside a 120 s run deadline, so a slow provider could push a run well beyond the slowest observed here.

The call cap counts calls at the client interface. HTTP retries inside the clients are invisible to it, so the real ceiling is about --max-llm-calls x (1 + LOAD_AGENT_LLM_MAX_RETRIES) requests (default retries 2: up to 450 for a cap of 150). Whether a retried request is billed is up to the provider. Tokens and cost here count the final successful attempt of each call only.

The TypeSafe per-call price is unknown. The TypeSafe cost below is calls multiplied by an assumed price, not a quote. Replace typesafe.assumed_usd_per_call in bench/prices.yaml and re-render when the real price is known.

Provider: typesafe-deepseek. Seed: 1. Date: 2026-10-08.

87 scripted inputs plus 4 timer firings observed (timers are generated by the agent, not scripted).

CategoryCountWhat it is
ping_regular26location pings; most repeat the current risk band, a few change it. Pings barely move cost: the filter and the band diff drop them in code, so the count only adds volume
ping_stale3pings whose recorded_at is older than the last applied ping (filtered in code, no model call)
ack14ok, thanks, thumbs up and similar replies from the driver
driver_eta8driver ETA replies in varied phrasing (4 before pickup, 4 for delivery)
driver_question8PO, address, appointment, instructions, and 2 questions no load fact answers
dispatcher_email7dispatcher emails without an ETA (one outside the main thread)
dispatcher_email_eta2dispatcher emails that carry an ETA
broker_slack5broker Slack news
noise10off-topic driver chatter
unknown_sender2SMS from a number that is not a contact of the load
stage_change2arrived at and departed from the pickup

Documents (BOL, POD): 0 in this mix. A synthetic image would not measure the vision model, so document extraction cost is not measured here.

MeasureValue
Runs (scripted inputs plus timers fired)91
Runs that called a model45 (49.5%)
Model calls per load49
Failed model calls0
Run latency p50 / p95 / max (all runs)0.02 s / 1.56 s / 21.21 s
Run latency p50 / p95 (runs that called a model)0.20 s / 7.01 s
Slowest runbench:1:053 (noise), 21.21 s: 2 model call(s): triage on jev-1.13.0 (192 ms), plan on deepseek-v4-pro (20988 ms); decision path planner
Purpose / modelCallsInput tokensCachedOutput tokensCostp50 latencyMax latency
plan / deepseek-v4-pro4474946084957$0.0200196817 ms20988 ms
triage / jev-1.13.0453302605996$0.090000186 ms1546 ms

TypeSafe rows show the assumed per-call price in the cost column.

PathRuns
escalated_unknown9
filtered36
planner3
sop8
triage35

Triage agreement with the oracle: 30/45. The oracle is the same pipeline answered by the generator’s own idea of what each message means. A run agrees when its action kinds (timer cancellations aside) equal the oracle’s for the same input. Disagreements after the first can be consequences of earlier ones (an ask left open changes what later messages do).

InputCategoryExpectedActualTriage categories returned
bench:1:002noisetriage: noopescalated_unknown: create_hero_tasknot recorded
bench:1:013noisetriage: nooptriage: tms_notenot recorded
bench:1:028noisetriage: nooptriage: tms_notenot recorded
bench:1:048broker_slacktriage: tms_noteescalated_unknown: create_hero_tasknot recorded
bench:1:053noisetriage: noopplanner: create_hero_tasknot recorded
bench:1:060noisetriage: nooptriage: tms_notenot recorded
bench:1:063driver_questionescalated_unknown: create_hero_tasktriage: send_smsnot recorded
bench:1:064noisetriage: nooptriage: tms_notenot recorded
bench:1:067dispatcher_emailtriage: tms_noteescalated_unknown: create_hero_tasknot recorded
bench:1:070noisetriage: noopplanner: create_hero_tasknot recorded
bench:1:071dispatcher_email_etatriage: cancel_timer, tms_noteescalated_unknown: create_hero_tasknot recorded
bench:1:072broker_slacktriage: tms_noteescalated_unknown: create_hero_tasknot recorded
bench:1:076noisetriage: nooptriage: tms_notenot recorded
bench:1:080noisetriage: nooptriage: tms_notenot recorded
bench:1:085broker_slacktriage: tms_noteescalated_unknown: create_hero_tasknot recorded

Cross-channel case, bench:1:071. A dispatcher email with an ETA (“Our driver expects to be at Valley DC at 8:30 am tomorrow.”) arrived while the delivery ETA ask was open and went to a Hero task (“could not be classified safely”) instead of resolving the ask. Investigation: the extractor finds the right candidate (“8:30 am tomorrow” read as 15:30 UTC, which is 8:30 PDT). The request was rebuilt by hand and replayed against the real TypeSafe decision model twice (2 calls). Both times the message was accepted as answering the ask: the candidate probability was 0.66 and 0.67 against the 0.60 acceptance threshold, the question probability 0.26 against a 0.30 gray band, kind new_info at 0.97. So the case sits close to two calibrated thresholds, and the benchmark run fell on the conservative side of one of them (escalated rather than resolved). The probabilities of the benchmark run itself were not recorded (this run predates recording the triage categories, which later runs keep), and the hand-built replay may differ from the in-run request, so the in-run cause is not proven. No bug in our extraction, email handling or stop selection was found, so this is reported as a known limitation: the calibrated gate is conservative on a third-person email ETA and sends borderline cases to a Hero, which is the designed failure direction (no wrong value is accepted). Unclear news. bench:1:048, 067, 072 and 085 are harmless FYI messages from the broker or dispatcher (for example “Heads up, the shipper called asking about the load.”). The decision model returned probabilities in the gray band, the result was unclear, and each became a Hero task. This is the same conservative gate as above and the largest source of avoidable Hero tasks in this run. Noise. Six chatter messages (013, 028, 060, 064, 076, 080, for example “radio keeps cutting out on me”) were read as new information and recorded as TMS notes: harmless but they put chatter in the TMS. Three more (002, 053, 070) were read as questions with no known topic and went to the planner, which produced a Hero task each (in 002 the planner proposed no steps, which also becomes a Hero task): “did you see the game last night” twice and “anyone know a decent diner near the next exit?”. The decision model is calibrated on ETA, temperature and seal questions, so open chatter is out of its distribution. A wrong-topic answer. bench:1:063 (“can I leave the trailer in the lot overnight?”) was labelled an instructions question and answered with the load instructions, which do not address the question. The oracle expects a planner call and a Hero task. This is the one case where the agent texted a driver something unhelpful. Label corrections. bench:1:023 (“Appreciate the updates today.”) and bench:1:045 (“Thanks for keeping us posted on this one.”) were labelled news by the generator; the model’s reading (an acknowledgement, nothing to do) is right. The generator labels were corrected after this run; the events did not change, so the run is still valid. Predates fixes. This run predates the Unicode punctuation change in the acknowledgement filter (no message in this timeline contains such punctuation, so routing is unchanged) and the recording of triage categories per call.

Acknowledgements. The first run of this benchmark found that a bare thumbs-up emoji was not recognised by the deterministic acknowledgement filter, so it cost a triage call. The filter now accepts thumbs-up, ok-hand, check mark and folded-hands emoji (with skin tones, repeated, or next to “ok”). In this run 5 of 14 acknowledgement inputs reached a model. Before the fix (bench/results/2026-10-08-typesafe-deepseek-before-ack-fix.json, same seed): 48 triage calls and 52 model calls in total, 55.2% of inputs reaching a model. After: 45 triage calls, 49 model calls, 51.7%.

Planner calls on the strong model. 4 planner calls on deepseek-v4-pro, from inputs of category: 1 from driver_question, 3 from noise. The synthetic mix has 2 driver questions that no load fact can answer; a call from any other category means the triage model read chatter as a question on no known topic. The calls cost $0.0200 (18% of the load’s cost at the assumed TypeSafe price) and the slowest took 20.99 s (the slowest run of the day was 21.21 s). Decision paths of those runs: 1 escalated_unknown, 3 planner.

Proposed lever (not implemented). Route a question with no topic and low planner value straight to a non-urgent Hero task, with no planner call. The alternative is to run the planner on deepseek-flash for those cases. Either change must be validated by evals first: a question the planner would have answered well must still be answered. Trade-off: fewer planner calls and a lower worst-case latency, against more Hero tasks for questions the strong model could have handled.

  • DeepSeek list prices from https://api-docs.deepseek.com/quick_start/pricing, fetched 2026-10-08. Tier used: peak (peak hours 01:00-04:00 and 06:00-10:00, weekdays; peak is the conservative choice).
  • TypeSafe: $0.0020 per call, an assumption.
  • Cost is recomputed here from reported tokens, not taken from the clients’ own cost field. Input tokens include cached ones; cached tokens are charged at the cache-hit price.
uv run load-agent bench --dry-run # no network, predicted call counts
uv run load-agent bench # real providers, spends money
uv run load-agent bench --render bench/results/<file>.json

The real run needs LOAD_AGENT_LLM_PROVIDER=typesafe+deepseek, TYPESAFE_AI_API_KEY and DEEPSEEK_API_KEY, and stops before exceeding --max-llm-calls (default 150). It is never part of per-push CI.

Prepared for Freight Hero by Marcus Caum Source on GitHub