Benchmark, interactive
One load over a synthetic day of 87 inputs, run against the real providers on 2026-10-08. Every figure on this page is read at build time from bench/results/2026-10-08-typesafe-deepseek.json, the same file the Benchmark page is generated from.
Headline
$0.020 DeepSeek measured, $0.090 TypeSafe at an assumed price
45 triage calls, 4 planner calls, over 91 runs (87 inputs and 4 timers)
The rest stop in deterministic code: filters, the risk band diff, the acknowledgement filter, compiled SOPs
bench:1:053: triage on jev-1.13.0 192 ms, plan on deepseek-v4-pro 21.0 s
What if the price or the volume changes?
LLM cost per load against the $1 budget
Move the TypeSafe price per call and the inbound volume. The load stays within budget until the price passes the break-even; DeepSeek, the measured part, is 18.2% of the measured total.
- DeepSeek, measured
- $0.020
- TypeSafe, 45 calls x price
- $0.090
- Break-even price per call
- $0.0218
- Break-even vs price
- 10.89x
Table view
| Measure | Value |
|---|---|
| TypeSafe price per call | $0.0020 |
| Inbound messages per load | 56 (measured) |
| TypeSafe calls | 45 |
| DeepSeek cost | $0.020 |
| TypeSafe cost | $0.090 |
| Cost per load | $0.110 |
| Budget | $1.00 |
| Break-even price per call | $0.0218 |
Where inputs go
Inputs per category: stopped in code or reached a model
45 of 87 inputs reached a model. Location pings, stale pings, unknown senders and stage changes never do; 9 of 14 acknowledgements stop in the deterministic filter. The 4 timers that fired are not inputs and all ran in code.
Table view
| Category | Inputs | Stopped in code | Reached a model | Share reaching a model |
|---|---|---|---|---|
Location ping ping_regular | 26 | 26 | 0 | 0.0% |
Acknowledgement ack | 14 | 9 | 5 | 35.7% |
Driver chatter noise | 10 | 0 | 10 | 100.0% |
Driver ETA driver_eta | 8 | 0 | 8 | 100.0% |
Driver question driver_question | 8 | 0 | 8 | 100.0% |
Dispatcher email dispatcher_email | 7 | 0 | 7 | 100.0% |
Broker Slack broker_slack | 5 | 0 | 5 | 100.0% |
Stale ping ping_stale | 3 | 3 | 0 | 0.0% |
Dispatcher email, ETA dispatcher_email_eta | 2 | 0 | 2 | 100.0% |
Unknown sender unknown_sender | 2 | 2 | 0 | 0.0% |
Stage change stage_change | 2 | 2 | 0 | 0.0% |
Run latency over the synthetic day
Wall time per run, log scale, by decision path
Runs that stay in code finish in milliseconds. Runs whose only model call is triage take at most 1.56 s (bench:1:071); every slower run called the planner on the strong model; even the slowest run uses 7.1% of the 5 min limit.
Table view (91 runs)
| # | Event | Category | Path | Wall time | Model calls |
|---|---|---|---|---|---|
| 1 | bench:1:000 | Location ping | Filtered | 1 ms | none |
| 2 | bench:1:001 | Acknowledgement | Filtered | 1 ms | none |
| 3 | bench:1:002 | Driver chatter | Escalated | 13.7 s | triage on jev-1.13.0, 577 ms; plan on deepseek-v4-pro, 13.1 s |
| 4 | bench:1:003 | Location ping | Filtered | 1 ms | none |
| 5 | bench:1:004 | Driver question | Triage | 566 ms | triage on jev-1.13.0, 563 ms |
| 6 | bench:1:005 | Location ping | Filtered | 1 ms | none |
| 7 | bench:1:006 | Broker Slack | Triage | 188 ms | triage on jev-1.13.0, 185 ms |
| 8 | bench:1:007 | Location ping | Filtered | 1 ms | none |
| 9 | bench:1:008 | Acknowledgement | Filtered | 1 ms | none |
| 10 | bench:1:009 | Location ping | Filtered | 1 ms | none |
| 11 | bench:1:010 | Unknown sender | Escalated | 2 ms | none |
| 12 | bench:1:011 | Driver ETA | Triage | 161 ms | triage on jev-1.13.0, 159 ms |
| 13 | bench:1:012 | Location ping | Filtered | 2 ms | none |
| 14 | bench:1:013 | Driver chatter | Triage | 193 ms | triage on jev-1.13.0, 189 ms |
| 15 | bench:1:014 | Dispatcher email | Triage | 212 ms | triage on jev-1.13.0, 209 ms |
| 16 | bench:1:015 | Location ping | Filtered | 2 ms | none |
| 17 | bench:1:016 | Acknowledgement | Triage | 189 ms | triage on jev-1.13.0, 185 ms |
| 18 | bench:1:017 | Driver ETA | Triage | 206 ms | triage on jev-1.13.0, 203 ms |
| 19 | bench:1:018 | Driver question | Triage | 156 ms | triage on jev-1.13.0, 151 ms |
| 20 | bench:1:019 | Location ping | SOP | 5 ms | none |
| 21 | bench:1:020 | Acknowledgement | Triage | 225 ms | triage on jev-1.13.0, 219 ms |
| 22 | bench:1:021 | Location ping | Filtered | 4 ms | none |
| 23 | bench:1:022 | Stale ping | Filtered | 4 ms | none |
| 24 | bench:1:023 | Broker Slack | Triage | 169 ms | triage on jev-1.13.0, 163 ms |
| 25 | bench:1:024 | Dispatcher email | Triage | 271 ms | triage on jev-1.13.0, 264 ms |
| 26 | bench:1:025 | Location ping | Filtered | 3 ms | none |
| 27 | evt:tmr:481207:ask:eta_request:bench:1:019:1 | Timer fired | SOP | 3 ms | none |
| 28 | bench:1:026 | Location ping | Filtered | 3 ms | none |
| 29 | bench:1:027 | Acknowledgement | Triage | 195 ms | triage on jev-1.13.0, 190 ms |
| 30 | bench:1:028 | Driver chatter | Triage | 197 ms | triage on jev-1.13.0, 191 ms |
| 31 | bench:1:029 | Location ping | Filtered | 4 ms | none |
| 32 | evt:tmr:481207:ask:eta_request:bench:1:019:2 | Timer fired | SOP | 5 ms | none |
| 33 | bench:1:030 | Driver question | Triage | 177 ms | triage on jev-1.13.0, 171 ms |
| 34 | bench:1:031 | Driver ETA | Triage | 189 ms | triage on jev-1.13.0, 179 ms |
| 35 | bench:1:032 | Dispatcher email, ETA | Triage | 180 ms | triage on jev-1.13.0, 173 ms |
| 36 | bench:1:033 | Location ping | Filtered | 6 ms | none |
| 37 | bench:1:034 | Driver ETA | Triage | 168 ms | triage on jev-1.13.0, 161 ms |
| 38 | bench:1:035 | Location ping | Filtered | 4 ms | none |
| 39 | bench:1:036 | Acknowledgement | Triage | 167 ms | triage on jev-1.13.0, 162 ms |
| 40 | bench:1:037 | Driver chatter | Triage | 201 ms | triage on jev-1.13.0, 195 ms |
| 41 | bench:1:038 | Stage change | SOP | 9 ms | none |
| 42 | bench:1:039 | Location ping | Filtered | 5 ms | none |
| 43 | bench:1:040 | Acknowledgement | Filtered | 5 ms | none |
| 44 | bench:1:041 | Location ping | Filtered | 5 ms | none |
| 45 | bench:1:042 | Driver question | Triage | 164 ms | triage on jev-1.13.0, 157 ms |
| 46 | bench:1:043 | Location ping | Filtered | 7 ms | none |
| 47 | bench:1:044 | Acknowledgement | Filtered | 7 ms | none |
| 48 | bench:1:045 | Dispatcher email | Triage | 205 ms | triage on jev-1.13.0, 198 ms |
| 49 | bench:1:046 | Stale ping | Filtered | 8 ms | none |
| 50 | bench:1:047 | Location ping | Filtered | 11 ms | none |
| 51 | bench:1:048 | Broker Slack | Escalated | 164 ms | triage on jev-1.13.0, 152 ms |
| 52 | bench:1:049 | Stage change | SOP | 6 ms | none |
| 53 | bench:1:050 | Acknowledgement | Triage | 187 ms | triage on jev-1.13.0, 177 ms |
| 54 | bench:1:051 | Driver question | Triage | 254 ms | triage on jev-1.13.0, 245 ms |
| 55 | bench:1:052 | Location ping | Filtered | 7 ms | none |
| 56 | bench:1:053 | Driver chatter | Planner | 21.2 s | triage on jev-1.13.0, 192 ms; plan on deepseek-v4-pro, 21.0 s |
| 57 | bench:1:054 | Dispatcher email | Triage | 557 ms | triage on jev-1.13.0, 546 ms |
| 58 | bench:1:055 | Driver question | Triage | 196 ms | triage on jev-1.13.0, 184 ms |
| 59 | bench:1:056 | Location ping | Filtered | 8 ms | none |
| 60 | bench:1:057 | Unknown sender | Escalated | 6 ms | none |
| 61 | bench:1:058 | Acknowledgement | Filtered | 6 ms | none |
| 62 | bench:1:059 | Acknowledgement | Filtered | 8 ms | none |
| 63 | bench:1:060 | Driver chatter | Triage | 210 ms | triage on jev-1.13.0, 202 ms |
| 64 | bench:1:061 | Dispatcher email | Escalated | 195 ms | triage on jev-1.13.0, 186 ms |
| 65 | bench:1:062 | Location ping | Filtered | 7 ms | none |
| 66 | bench:1:063 | Driver question | Triage | 195 ms | triage on jev-1.13.0, 186 ms |
| 67 | bench:1:064 | Driver chatter | Triage | 178 ms | triage on jev-1.13.0, 167 ms |
| 68 | bench:1:065 | Location ping | Filtered | 11 ms | none |
| 69 | bench:1:066 | Acknowledgement | Filtered | 12 ms | none |
| 70 | bench:1:067 | Dispatcher email | Escalated | 193 ms | triage on jev-1.13.0, 179 ms |
| 71 | bench:1:068 | Driver ETA | Triage | 190 ms | triage on jev-1.13.0, 181 ms |
| 72 | bench:1:069 | Location ping | SOP | 10 ms | none |
| 73 | bench:1:070 | Driver chatter | Planner | 5.93 s | triage on jev-1.13.0, 177 ms; plan on deepseek-v4-pro, 5.74 s |
| 74 | evt:tmr:481207:ask:eta_request:bench:1:069:1 | Timer fired | SOP | 9 ms | none |
| 75 | bench:1:071 | Dispatcher email, ETA | Escalated | 1.56 s | triage on jev-1.13.0, 1.55 s |
| 76 | bench:1:072 | Broker Slack | Escalated | 204 ms | triage on jev-1.13.0, 187 ms |
| 77 | bench:1:073 | Location ping | Filtered | 12 ms | none |
| 78 | evt:tmr:481207:ask:eta_request:bench:1:069:2 | Timer fired | SOP | 13 ms | none |
| 79 | bench:1:074 | Acknowledgement | Filtered | 13 ms | none |
| 80 | bench:1:075 | Driver ETA | Triage | 239 ms | triage on jev-1.13.0, 225 ms |
| 81 | bench:1:076 | Driver chatter | Triage | 217 ms | triage on jev-1.13.0, 200 ms |
| 82 | bench:1:077 | Location ping | Filtered | 14 ms | none |
| 83 | bench:1:078 | Driver ETA | Triage | 193 ms | triage on jev-1.13.0, 176 ms |
| 84 | bench:1:079 | Driver question | Planner | 7.01 s | triage on jev-1.13.0, 179 ms; plan on deepseek-v4-pro, 6.82 s |
| 85 | bench:1:080 | Driver chatter | Triage | 560 ms | triage on jev-1.13.0, 541 ms |
| 86 | bench:1:081 | Location ping | Filtered | 16 ms | none |
| 87 | bench:1:082 | Dispatcher email | Triage | 202 ms | triage on jev-1.13.0, 184 ms |
| 88 | bench:1:083 | Stale ping | Filtered | 17 ms | none |
| 89 | bench:1:084 | Acknowledgement | Filtered | 17 ms | none |
| 90 | bench:1:085 | Broker Slack | Escalated | 182 ms | triage on jev-1.13.0, 162 ms |
| 91 | bench:1:086 | Driver ETA | Triage | 221 ms | triage on jev-1.13.0, 201 ms |
Cost and calls by purpose and model
Cost, calls and latency per purpose and model
Triage on the small decision model makes most of the calls and, at the assumed price, most of the cost. The planner is called 4 times, but its median call takes 6.82 s against 186 ms for triage, which is where the latency tail comes from.
Cost per load
* calls x assumed price
Calls per load
Latency per call, p50 to max (log)
Table view
| Purpose / model | Calls | Input tokens | Cached | Output tokens | Cost | p50 | Max |
|---|---|---|---|---|---|---|---|
| triage / jev-1.13.0 | 45 | 33,026 | 0 (0.0%) | 5,996 | $0.0900 (assumed price) | 186 ms | 1.55 s |
| plan / deepseek-v4-pro | 4 | 4,749 | 4,608 (97.0%) | 4,957 | $0.0200 | 6.82 s | 21.0 s |
Triage agreement
Triage runs whose actions match the oracle
The oracle replays the same pipeline with the generator's own reading of each message. 14 of the 15 disagreements fail safe, toward a Hero task or a TMS note; one sent a contact an unhelpful reply.
The 15 disagreements
bench:1:002 | Driver chatter | Triage: nothing | Escalated: Hero task | Fails safe: Hero task |
bench:1:013 | Driver chatter | Triage: nothing | Triage: TMS note | Fails safe: TMS note |
bench:1:028 | Driver chatter | Triage: nothing | Triage: TMS note | Fails safe: TMS note |
bench:1:048 | Broker Slack | Triage: TMS note | Escalated: Hero task | Fails safe: Hero task |
bench:1:053 | Driver chatter | Triage: nothing | Planner: Hero task | Fails safe: Hero task |
bench:1:060 | Driver chatter | Triage: nothing | Triage: TMS note | Fails safe: TMS note |
bench:1:063 | Driver question | Escalated: Hero task | Triage: SMS | Unhelpful reply |
bench:1:064 | Driver chatter | Triage: nothing | Triage: TMS note | Fails safe: TMS note |
bench:1:067 | Dispatcher email | Triage: TMS note | Escalated: Hero task | Fails safe: Hero task |
bench:1:070 | Driver chatter | Triage: nothing | Planner: Hero task | Fails safe: Hero task |
bench:1:071 | Dispatcher email, ETA | Triage: Timer cancelled, TMS note | Escalated: Hero task | Fails safe: Hero task |
bench:1:072 | Broker Slack | Triage: TMS note | Escalated: Hero task | Fails safe: Hero task |
bench:1:076 | Driver chatter | Triage: nothing | Triage: TMS note | Fails safe: TMS note |
bench:1:080 | Driver chatter | Triage: nothing | Triage: TMS note | Fails safe: TMS note |
bench:1:085 | Broker Slack | Triage: TMS note | Escalated: Hero task | Fails safe: Hero task |
Table view of all 45 triage runs
| Input | Category | Outcome |
|---|---|---|
bench:1:002 | Driver chatter | Fails safe: Hero task |
bench:1:004 | Driver question | Agrees |
bench:1:006 | Broker Slack | Agrees |
bench:1:011 | Driver ETA | Agrees |
bench:1:013 | Driver chatter | Fails safe: TMS note |
bench:1:014 | Dispatcher email | Agrees |
bench:1:016 | Acknowledgement | Agrees |
bench:1:017 | Driver ETA | Agrees |
bench:1:018 | Driver question | Agrees |
bench:1:020 | Acknowledgement | Agrees |
bench:1:023 | Broker Slack | Agrees |
bench:1:024 | Dispatcher email | Agrees |
bench:1:027 | Acknowledgement | Agrees |
bench:1:028 | Driver chatter | Fails safe: TMS note |
bench:1:030 | Driver question | Agrees |
bench:1:031 | Driver ETA | Agrees |
bench:1:032 | Dispatcher email, ETA | Agrees |
bench:1:034 | Driver ETA | Agrees |
bench:1:036 | Acknowledgement | Agrees |
bench:1:037 | Driver chatter | Agrees |
bench:1:042 | Driver question | Agrees |
bench:1:045 | Dispatcher email | Agrees |
bench:1:048 | Broker Slack | Fails safe: Hero task |
bench:1:050 | Acknowledgement | Agrees |
bench:1:051 | Driver question | Agrees |
bench:1:053 | Driver chatter | Fails safe: Hero task |
bench:1:054 | Dispatcher email | Agrees |
bench:1:055 | Driver question | Agrees |
bench:1:060 | Driver chatter | Fails safe: TMS note |
bench:1:061 | Dispatcher email | Agrees |
bench:1:063 | Driver question | Unhelpful reply |
bench:1:064 | Driver chatter | Fails safe: TMS note |
bench:1:067 | Dispatcher email | Fails safe: Hero task |
bench:1:068 | Driver ETA | Agrees |
bench:1:070 | Driver chatter | Fails safe: Hero task |
bench:1:071 | Dispatcher email, ETA | Fails safe: Hero task |
bench:1:072 | Broker Slack | Fails safe: Hero task |
bench:1:075 | Driver ETA | Agrees |
bench:1:076 | Driver chatter | Fails safe: TMS note |
bench:1:078 | Driver ETA | Agrees |
bench:1:079 | Driver question | Agrees |
bench:1:080 | Driver chatter | Fails safe: TMS note |
bench:1:082 | Dispatcher email | Agrees |
bench:1:085 | Broker Slack | Fails safe: Hero task |
bench:1:086 | Driver ETA | Agrees |
What the run's own notes say
- Cross-channel case, bench:1:071. A dispatcher email with an ETA ("Our driver expects to be at Valley DC at 8:30 am tomorrow.") arrived while the delivery ETA ask was open and went to a Hero task ("could not be classified safely") instead of resolving the ask. Investigation: the extractor finds the right candidate ("8:30 am tomorrow" read as 15:30 UTC, which is 8:30 PDT). The request was rebuilt by hand and replayed against the real TypeSafe decision model twice (2 calls). Both times the message was accepted as answering the ask: the candidate probability was 0.66 and 0.67 against the 0.60 acceptance threshold, the question probability 0.26 against a 0.30 gray band, kind new_info at 0.97. So the case sits close to two calibrated thresholds, and the benchmark run fell on the conservative side of one of them (escalated rather than resolved). The probabilities of the benchmark run itself were not recorded (this run predates recording the triage categories, which later runs keep), and the hand-built replay may differ from the in-run request, so the in-run cause is not proven. No bug in our extraction, email handling or stop selection was found, so this is reported as a known limitation: the calibrated gate is conservative on a third-person email ETA and sends borderline cases to a Hero, which is the designed failure direction (no wrong value is accepted).
- Unclear news. bench:1:048, 067, 072 and 085 are harmless FYI messages from the broker or dispatcher (for example "Heads up, the shipper called asking about the load."). The decision model returned probabilities in the gray band, the result was unclear, and each became a Hero task. This is the same conservative gate as above and the largest source of avoidable Hero tasks in this run.
- Noise. Six chatter messages (013, 028, 060, 064, 076, 080, for example "radio keeps cutting out on me") were read as new information and recorded as TMS notes: harmless but they put chatter in the TMS. Three more (002, 053, 070) were read as questions with no known topic and went to the planner, which produced a Hero task each (in 002 the planner proposed no steps, which also becomes a Hero task): "did you see the game last night" twice and "anyone know a decent diner near the next exit?". The decision model is calibrated on ETA, temperature and seal questions, so open chatter is out of its distribution.
- A wrong-topic answer. bench:1:063 ("can I leave the trailer in the lot overnight?") was labelled an instructions question and answered with the load instructions, which do not address the question. The oracle expects a planner call and a Hero task. This is the one case where the agent texted a driver something unhelpful.
- Label corrections. bench:1:023 ("Appreciate the updates today.") and bench:1:045 ("Thanks for keeping us posted on this one.") were labelled news by the generator; the model's reading (an acknowledgement, nothing to do) is right. The generator labels were corrected after this run; the events did not change, so the run is still valid.
- Predates fixes. This run predates the Unicode punctuation change in the acknowledgement filter (no message in this timeline contains such punctuation, so routing is unchanged) and the recording of triage categories per call.
Before and after the acknowledgement fix
Calls, cost and model reach, before and after the fix
The first run sent a bare thumbs-up emoji to the triage model. Teaching the deterministic filter common acknowledgement emoji removed 3 triage calls on the same seed. Only $0.0060 of the $0.0089 cost difference comes from those calls; the other $0.0029 is DeepSeek variance between the two runs (the same 4 planner calls, different token and cache counts) with a different number of timers firing (3 before, 4 after). Each measure has its own scale from zero.
Model calls per load
Triage calls per load
LLM cost per load
Inputs reaching a model
Table view
| Measure | Before | After |
|---|---|---|
| Model calls per load | 52 | 49 |
| Triage calls per load | 48 | 45 |
| LLM cost per load | $0.1190 | $0.1100 |
| Inputs reaching a model | 55.2% | 51.7% |
Before: bench/results/2026-10-08-typesafe-deepseek-before-ack-fix.json. After: bench/results/2026-10-08-typesafe-deepseek.json.
Method, price sources and how to reproduce the run: Benchmark. Keyboard: Tab to a chart, then the arrow keys move between marks; Escape hides the tooltip.