Benchmark: cost and latency of one load over a realistic day
Generated from bench/results/2026-10-08-typesafe-deepseek.json by uv run load-agent bench --render. Do not edit by hand.
This measures the assignment’s budget: about $1 of LLM inference per load over 50 to 100 inbound messages, and at most 5 minutes per run.
Verdict
Section titled “Verdict”- Cost per load: $0.1100 against a budget of $1.00: within budget ($0.020019 DeepSeek, measured tokens at list prices; $0.0900 TypeSafe, 45 calls at an assumed $0.0020 per call).
- TypeSafe break-even: the load reaches $1.00 at $0.0218 per TypeSafe call (45 calls).
- Slowest run: 21.21 s against a limit of 300 s: within the limit. This is
process_eventonly (no dispatch, queue wait or pump lag), and retries could push the worst case higher (see Method). - Inputs that reach a model: 51.7% of the 87 scripted inputs (49.5% of all 91 runs, timers included); the rest stop in deterministic code.
Method
Section titled “Method”Time inside the agent is simulated (FakeClock): the 30 minute follow-ups fire at the right moment of the synthetic day without waiting. Channels are fakes, so no SMS or email was sent. The timer pump runs once per simulated minute.
Latency was measured from this cloud container with the network to TypeSafe and DeepSeek included. It is wall-clock time of each process_event call. Production latency depends on where the workers run and on provider load at the time, so treat it as one sample, not a guarantee.
The latency figures cover process_event only: they exclude the outbox dispatch, the queue wait and the timer pump lag. The default per-attempt timeout is 30 s with 2 retries inside a 120 s run deadline, so a slow provider could push a run well beyond the slowest observed here.
The call cap counts calls at the client interface. HTTP retries inside the clients are invisible to it, so the real ceiling is about --max-llm-calls x (1 + LOAD_AGENT_LLM_MAX_RETRIES) requests (default retries 2: up to 450 for a cap of 150). Whether a retried request is billed is up to the provider. Tokens and cost here count the final successful attempt of each call only.
The TypeSafe per-call price is unknown. The TypeSafe cost below is calls multiplied by an assumed price, not a quote. Replace typesafe.assumed_usd_per_call in bench/prices.yaml and re-render when the real price is known.
Provider: typesafe-deepseek. Seed: 1. Date: 2026-10-08.
Mix of the synthetic day
Section titled “Mix of the synthetic day”87 scripted inputs plus 4 timer firings observed (timers are generated by the agent, not scripted).
| Category | Count | What it is |
|---|---|---|
| ping_regular | 26 | location pings; most repeat the current risk band, a few change it. Pings barely move cost: the filter and the band diff drop them in code, so the count only adds volume |
| ping_stale | 3 | pings whose recorded_at is older than the last applied ping (filtered in code, no model call) |
| ack | 14 | ok, thanks, thumbs up and similar replies from the driver |
| driver_eta | 8 | driver ETA replies in varied phrasing (4 before pickup, 4 for delivery) |
| driver_question | 8 | PO, address, appointment, instructions, and 2 questions no load fact answers |
| dispatcher_email | 7 | dispatcher emails without an ETA (one outside the main thread) |
| dispatcher_email_eta | 2 | dispatcher emails that carry an ETA |
| broker_slack | 5 | broker Slack news |
| noise | 10 | off-topic driver chatter |
| unknown_sender | 2 | SMS from a number that is not a contact of the load |
| stage_change | 2 | arrived at and departed from the pickup |
Documents (BOL, POD): 0 in this mix. A synthetic image would not measure the vision model, so document extraction cost is not measured here.
Results
Section titled “Results”| Measure | Value |
|---|---|
| Runs (scripted inputs plus timers fired) | 91 |
| Runs that called a model | 45 (49.5%) |
| Model calls per load | 49 |
| Failed model calls | 0 |
| Run latency p50 / p95 / max (all runs) | 0.02 s / 1.56 s / 21.21 s |
| Run latency p50 / p95 (runs that called a model) | 0.20 s / 7.01 s |
| Slowest run | bench:1:053 (noise), 21.21 s: 2 model call(s): triage on jev-1.13.0 (192 ms), plan on deepseek-v4-pro (20988 ms); decision path planner |
Calls by purpose and model
Section titled “Calls by purpose and model”| Purpose / model | Calls | Input tokens | Cached | Output tokens | Cost | p50 latency | Max latency |
|---|---|---|---|---|---|---|---|
| plan / deepseek-v4-pro | 4 | 4749 | 4608 | 4957 | $0.020019 | 6817 ms | 20988 ms |
| triage / jev-1.13.0 | 45 | 33026 | 0 | 5996 | $0.090000 | 186 ms | 1546 ms |
TypeSafe rows show the assumed per-call price in the cost column.
Decision paths
Section titled “Decision paths”| Path | Runs |
|---|---|
| escalated_unknown | 9 |
| filtered | 36 |
| planner | 3 |
| sop | 8 |
| triage | 35 |
Findings
Section titled “Findings”Triage agreement with the oracle
Section titled “Triage agreement with the oracle”Triage agreement with the oracle: 30/45. The oracle is the same pipeline answered by the generator’s own idea of what each message means. A run agrees when its action kinds (timer cancellations aside) equal the oracle’s for the same input. Disagreements after the first can be consequences of earlier ones (an ask left open changes what later messages do).
| Input | Category | Expected | Actual | Triage categories returned |
|---|---|---|---|---|
| bench:1:002 | noise | triage: noop | escalated_unknown: create_hero_task | not recorded |
| bench:1:013 | noise | triage: noop | triage: tms_note | not recorded |
| bench:1:028 | noise | triage: noop | triage: tms_note | not recorded |
| bench:1:048 | broker_slack | triage: tms_note | escalated_unknown: create_hero_task | not recorded |
| bench:1:053 | noise | triage: noop | planner: create_hero_task | not recorded |
| bench:1:060 | noise | triage: noop | triage: tms_note | not recorded |
| bench:1:063 | driver_question | escalated_unknown: create_hero_task | triage: send_sms | not recorded |
| bench:1:064 | noise | triage: noop | triage: tms_note | not recorded |
| bench:1:067 | dispatcher_email | triage: tms_note | escalated_unknown: create_hero_task | not recorded |
| bench:1:070 | noise | triage: noop | planner: create_hero_task | not recorded |
| bench:1:071 | dispatcher_email_eta | triage: cancel_timer, tms_note | escalated_unknown: create_hero_task | not recorded |
| bench:1:072 | broker_slack | triage: tms_note | escalated_unknown: create_hero_task | not recorded |
| bench:1:076 | noise | triage: noop | triage: tms_note | not recorded |
| bench:1:080 | noise | triage: noop | triage: tms_note | not recorded |
| bench:1:085 | broker_slack | triage: tms_note | escalated_unknown: create_hero_task | not recorded |
Cross-channel case, bench:1:071. A dispatcher email with an ETA (“Our driver expects to be at Valley DC at 8:30 am tomorrow.”) arrived while the delivery ETA ask was open and went to a Hero task (“could not be classified safely”) instead of resolving the ask. Investigation: the extractor finds the right candidate (“8:30 am tomorrow” read as 15:30 UTC, which is 8:30 PDT). The request was rebuilt by hand and replayed against the real TypeSafe decision model twice (2 calls). Both times the message was accepted as answering the ask: the candidate probability was 0.66 and 0.67 against the 0.60 acceptance threshold, the question probability 0.26 against a 0.30 gray band, kind new_info at 0.97. So the case sits close to two calibrated thresholds, and the benchmark run fell on the conservative side of one of them (escalated rather than resolved). The probabilities of the benchmark run itself were not recorded (this run predates recording the triage categories, which later runs keep), and the hand-built replay may differ from the in-run request, so the in-run cause is not proven. No bug in our extraction, email handling or stop selection was found, so this is reported as a known limitation: the calibrated gate is conservative on a third-person email ETA and sends borderline cases to a Hero, which is the designed failure direction (no wrong value is accepted).
Unclear news. bench:1:048, 067, 072 and 085 are harmless FYI messages from the broker or dispatcher (for example “Heads up, the shipper called asking about the load.”). The decision model returned probabilities in the gray band, the result was unclear, and each became a Hero task. This is the same conservative gate as above and the largest source of avoidable Hero tasks in this run.
Noise. Six chatter messages (013, 028, 060, 064, 076, 080, for example “radio keeps cutting out on me”) were read as new information and recorded as TMS notes: harmless but they put chatter in the TMS. Three more (002, 053, 070) were read as questions with no known topic and went to the planner, which produced a Hero task each (in 002 the planner proposed no steps, which also becomes a Hero task): “did you see the game last night” twice and “anyone know a decent diner near the next exit?”. The decision model is calibrated on ETA, temperature and seal questions, so open chatter is out of its distribution.
A wrong-topic answer. bench:1:063 (“can I leave the trailer in the lot overnight?”) was labelled an instructions question and answered with the load instructions, which do not address the question. The oracle expects a planner call and a Hero task. This is the one case where the agent texted a driver something unhelpful.
Label corrections. bench:1:023 (“Appreciate the updates today.”) and bench:1:045 (“Thanks for keeping us posted on this one.”) were labelled news by the generator; the model’s reading (an acknowledgement, nothing to do) is right. The generator labels were corrected after this run; the events did not change, so the run is still valid.
Predates fixes. This run predates the Unicode punctuation change in the acknowledgement filter (no message in this timeline contains such punctuation, so routing is unchanged) and the recording of triage categories per call.
Acknowledgements. The first run of this benchmark found that a bare thumbs-up emoji was not recognised by the deterministic acknowledgement filter, so it cost a triage call. The filter now accepts thumbs-up, ok-hand, check mark and folded-hands emoji (with skin tones, repeated, or next to “ok”). In this run 5 of 14 acknowledgement inputs reached a model.
Before the fix (bench/results/2026-10-08-typesafe-deepseek-before-ack-fix.json, same seed): 48 triage calls and 52 model calls in total, 55.2% of inputs reaching a model. After: 45 triage calls, 49 model calls, 51.7%.
Planner calls on the strong model. 4 planner calls on deepseek-v4-pro, from inputs of category: 1 from driver_question, 3 from noise. The synthetic mix has 2 driver questions that no load fact can answer; a call from any other category means the triage model read chatter as a question on no known topic. The calls cost $0.0200 (18% of the load’s cost at the assumed TypeSafe price) and the slowest took 20.99 s (the slowest run of the day was 21.21 s). Decision paths of those runs: 1 escalated_unknown, 3 planner.
Proposed lever (not implemented). Route a question with no topic and low planner value straight to a non-urgent Hero task, with no planner call. The alternative is to run the planner on deepseek-flash for those cases. Either change must be validated by evals first: a question the planner would have answered well must still be answered. Trade-off: fewer planner calls and a lower worst-case latency, against more Hero tasks for questions the strong model could have handled.
Price assumptions
Section titled “Price assumptions”- DeepSeek list prices from https://api-docs.deepseek.com/quick_start/pricing, fetched 2026-10-08. Tier used: peak (peak hours 01:00-04:00 and 06:00-10:00, weekdays; peak is the conservative choice).
- TypeSafe: $0.0020 per call, an assumption.
- Cost is recomputed here from reported tokens, not taken from the clients’ own cost field. Input tokens include cached ones; cached tokens are charged at the cache-hit price.
Reproduce
Section titled “Reproduce”uv run load-agent bench --dry-run # no network, predicted call countsuv run load-agent bench # real providers, spends moneyuv run load-agent bench --render bench/results/<file>.jsonThe real run needs LOAD_AGENT_LLM_PROVIDER=typesafe+deepseek, TYPESAFE_AI_API_KEY and DEEPSEEK_API_KEY, and stops before exceeding --max-llm-calls (default 150). It is never part of per-push CI.