Build log · Part 5 of 6
Cost
$997.48 across 85,613 calls, itemised by month, model, stage, token, and unit of output
The tenet catalogued in Part 4 has two halves. The subscription half is rp_authorship_log, one row
per commit. The metered half is eng_llm_traces, one row per paid model call, written by a logging
wrapper with cost, token counts, latency, status, and parent attribution.
No untracked AI work. Every AI-driven action leaves a structured trace: cost for metered runtimes, provenance for subscription runtimes.
Direct SDK instantiation outside the wrapper module was prohibited. The figures below are the full table, not a sample.
Totals
| metric | value |
|---|---|
| Model spend | $997.48 |
| Calls | 85,613 |
| Period | 2026-04-09 → 2026-07-24 |
| Prompt tokens | 221,700,000 |
| Output tokens | 28,700,000 |
Mean latency (status='ok') | 6,816 ms |
| Companies in graph | 1,787 |
| Signals collected | 55,406 |
| Cost per company | $0.558 |
| Cost per signal | $0.018 |
Two aggregates behind those totals are exact rather than rounded: sum(cost_usd) is
$997.480519 and the token columns sum to 221,700,073 prompt and 28,740,244 output. The trace
table itself occupied 470 MB of the 2,031 MB database at teardown — 23% of the graph’s storage
was the record of what it cost to build the graph.
Instrumentation coverage is not uniform, and the rest of this accounting depends on knowing where it is thin:
| field | rows populated | share of rows | share of spend covered |
|---|---|---|---|
cost_usd | 85,613 | 100% | 100% |
latency_ms, status | 85,613 | 100% | 100% |
prompt_tokens / output_tokens | 51,917 | 60.6% | 31.2% |
parent_type | 60,341 | 70.5% | 48.8% |
Cost was captured from the first call. Token counts and parent attribution were added to the wrapper later, so the token analysis below is restricted to the months where coverage is complete and the attribution analysis is stated with its coverage attached.
By month
| month | calls | spend | mean/call |
|---|---|---|---|
| April | 10,976 | $174.43 | $0.01589 |
| May | 30,339 | $528.50 | $0.01742 |
| June | 33,598 | $206.29 | $0.00614 |
| July (to 24th) | 10,700 | $88.26 | $0.00825 |
June ran 10.7% more calls than May at 39% of the cost. Mean cost per call fell 64.8%.
Two changes landed between those months: model tiering (bulk extraction moved to Haiku, judgment retained on Sonnet, Opus restricted to design and review) and the deterministic-first rule, which suppresses model calls for facts already produced by a deterministic path.
Neither change reduced corpus output. The reduction came from call composition and call avoidance, not from a cheaper provider or a reduced workload.
The same four months, measured on reliability and instrumentation rather than spend:
| month | error rate | rows with token counts | metered prompt:output |
|---|---|---|---|
| April | 2.00% | 29 / 10,976 | 2.98 : 1 |
| May | 8.55% | 8,091 / 30,339 | 3.15 : 1 |
| June | 0.77% | 33,127 / 33,598 | 3.55 : 1 |
| July | 0.28% | 10,670 / 10,700 | 6.72 : 1 |
May is the outlier in every column.
What the $528.50 May consisted of
May carried 53.0% of total model spend across 35.4% of the calls. It divides into four movements: one agent running continuously at Sonnet prices, a four-day credential outage in the middle of it, a final-week cost escalation in that same agent, and a large volume of cheap warfare enrichment arriving at the end of the month.
By calendar week, counting only rows dated in May:
| week | calls | spend | mean/call |
|---|---|---|---|
| May 1–3 | 720 | $15.75 | $0.02188 |
| May 4–10 | 5,496 | $128.65 | $0.02341 |
| May 11–17 | 5,650 | $124.90 | $0.02211 |
| May 18–24 | 4,995 | $63.09 | $0.01263 |
| May 25–31 | 13,478 | $196.10 | $0.01455 |
By agent, the top of the month:
| agent | calls | spend | share of May | mean/call |
|---|---|---|---|---|
business_events | 8,177 | $262.15 | 49.6% | $0.03206 |
triage | 4,377 | $51.00 | 9.6% | $0.01165 |
editorial_qa | 117 | $26.88 | 5.1% | $0.22972 |
attack_event_enrich | 7,597 | $25.02 | 4.7% | $0.00329 |
editorial_review | 1,127 | $21.26 | 4.0% | $0.01886 |
conflict_extractor | 366 | $19.27 | 3.6% | $0.05266 |
cluster_conflict | 122 | $14.88 | 2.8% | $0.12193 |
analyze | 184 | $11.59 | 2.2% | $0.06301 |
pipeline_qa_sonnet | 164 | $10.36 | 2.0% | $0.06316 |
signal_enricher | 2,091 | $10.07 | 1.9% | $0.00482 |
business_events accounts for 49.6% of May. Its entire lifetime spend — 8,177 calls, $262.15,
26.3% of the whole project — falls inside May, and every one of those calls ran on
anthropic/claude-sonnet-4-6. It ran daily from 4 May to 30 May and never ran again.
Its per-call cost was not flat across that run:
| week | calls | spend | mean/call | mean latency |
|---|---|---|---|---|
| May 4–10 | 2,556 | $66.00 | $0.02582 | 2,308 ms |
| May 11–17 | 2,563 | $62.65 | $0.02444 | 1,648 ms |
| May 18–24 | 1,535 | $31.92 | $0.02080 | 1,406 ms |
| May 25–31 | 1,523 | $101.57 | $0.06669 | 2,387 ms |
The final week ran the fewest calls of any full week and cost the most. Mean cost per call rose 3.2× against the preceding week while mean latency rose 1.7×, which is the signature of longer documents entering the same extractor rather than a change in call count. The last four days of the run — 27 to 30 May — cost $83.61 between them, against $9.27 on 13 May for a comparable call count.
The single most expensive day of May was 30 May at $37.90 across 1,357 calls. The busiest day was
31 May at 7,920 calls for $29.61: attack_event_enrich starting up, at $0.00329 per call.
Between those two patterns sits the outage:
| day | calls | errors | spend |
|---|---|---|---|
| 2026-05-19 | 795 | 0 | $16.55 |
| 2026-05-20 | 679 | 1 | $16.02 |
| 2026-05-21 | 695 | 328 | $8.00 |
| 2026-05-22 | 702 | 702 | $0.00 |
| 2026-05-23 | 651 | 651 | $0.00 |
| 2026-05-24 | 763 | 504 | $6.25 |
| 2026-05-25 | 1,077 | 184 | $19.43 |
| 2026-05-26 | 774 | 0 | $24.93 |
On 22 and 23 May every call the system made failed. Spend for those two days was $0.00 and output was nothing. The workers, timers and cron jobs ran on schedule throughout; the calls left the host, were rejected at the provider, and were logged as traces with a cost of zero. Nothing in the spend series marks the failure — it reads as two cheap days.
By model
| model id | calls | spend | mean/call |
|---|---|---|---|
anthropic/claude-sonnet-4-6 | 15,157 | $464.26 | $0.03063 |
anthropic/claude-haiku-4.5 | 64,437 | $377.49 | $0.00586 |
anthropic/claude-opus-4 | 182 | $44.11 | $0.24234 |
anthropic/claude-opus-4.6 | 638 | $39.18 | $0.06141 |
grok-4.3 | 1,166 | $31.41 | $0.02694 |
anthropic/claude-sonnet-4.5 | 300 | $28.59 | $0.09530 |
anthropic/claude-haiku-4-5 | 2,301 | $10.45 | $0.00454 |
openai/gpt-5.4 | 6 | $0.03 | $0.00500 |
text-embedding-3-small | 521 | $0.01 | — |
Haiku ran 4.4× Sonnet’s call volume at 81% of Sonnet’s total cost. Opus variants totalled 820 calls across the whole run, all of them schema design, architecture decisions, or review passes.
Fifteen further model ids appear in the table at under $0.05 each, $0.1282 combined across 1,429
calls. Thirteen of them are routing experiments that were not adopted — deepseek/deepseek-v3.2-exp
(81 calls, $0.0000), z-ai/glm-4.5-air (7, $0.0012), google/gemini-2.5-flash (6, $0.0083),
openai/gpt-5-nano (6, $0.0000), x-ai/grok-4-fast (6, $0.0000), anthropic/claude-3-haiku
(6, $0.0030), minimax/minimax-m2.5 (2, $0.0001) among them. The cost of trying an alternative
provider, in this build, was cents.
Restricting to rows that carry token counts, the effective blended rate per million tokens (prompt and output combined, at the mix each stage actually sent):
| model id | calls | prompt tok | output tok | spend | $/M tokens |
|---|---|---|---|---|---|
anthropic/claude-haiku-4.5 | 48,098 | 99,608,120 | 22,990,930 | $212.46 | $1.733 |
anthropic/claude-sonnet-4-6 | 1,136 | 5,561,439 | 2,325,555 | $51.13 | $6.483 |
grok-4.3 | 948 | 7,911,032 | 2,394,225 | $25.75 | $2.499 |
anthropic/claude-opus-4.6 | 241 | 1,570,393 | 329,695 | $15.93 | $8.386 |
anthropic/claude-sonnet-4.5 | 96 | 826,045 | 180,576 | $5.13 | $5.101 |
text-embedding-3-small | 1,260 | 1,362,173 | 0 | $0.03 | $0.020 |
The Haiku-to-Sonnet blended ratio is 3.7×, not the 10× or more that headline list prices imply, because the stages routed to Sonnet sent proportionally more output tokens than the stages routed to Haiku. Tiering moves work between price points; it does not move it between the token mixes each kind of work requires.
By stage
| agent | calls | spend | share | mean/call |
|---|---|---|---|---|
business_events | 8,177 | $262.15 | 26.3% | $0.03206 |
triage | 8,584 | $106.71 | 10.7% | $0.01243 |
attack_event_enrich | 29,805 | $95.99 | 9.6% | $0.00322 |
conflict_extractor | 1,264 | $74.56 | 7.5% | $0.05899 |
signal_enricher | 5,487 | $60.94 | 6.1% | $0.01111 |
editorial_qa | 175 | $43.04 | 4.3% | $0.24594 |
editorial_review | 1,804 | $37.67 | 3.8% | $0.02088 |
x_search_listener | 1,166 | $31.41 | 3.1% | $0.02694 |
cluster_conflict | 221 | $27.23 | 2.7% | $0.12321 |
analyze | 443 | $26.81 | 2.7% | $0.06052 |
content_remediation | 708 | $16.98 | 1.7% | $0.02398 |
business_events extracted funding rounds, contracts, partnerships and acquisitions from full
documents into linked structured output. Highest per-call token count in both directions, 26.3% of
total spend, and the source of deal records on 78% of companies.
attack_event_enrich had the lowest mean cost of any substantial stage at $0.00322, and the highest
volume at 34.8% of all calls in the project.
Sixty-five distinct agent_slug values appear in the table. The eleven above account for 78.5% of
spend; the remaining fifty-four account for $214.00. One of them is unknown — 1,216 calls and
$28.31 between 9 April and 7 May, before the wrapper required a slug. That is 2.8% of total spend
whose stage cannot be recovered, and the argument for making the attribution field non-nullable
before the first call rather than after the first month.
Token economics
The headline ratio is 221,700,073 prompt tokens against 28,740,244 output tokens, or 7.71 : 1. That number is not a spend ratio, and the difference matters more than the number.
Of the 221.7M prompt tokens, 104,654,955 — 47.2% — come from twenty rows. They are
codex_dispatch: build jobs handed to a subscription-runtime coding agent, which the wrapper logs
for provenance at cost_usd = 0. The largest single row carries 20,581,575 prompt tokens against
43,378 output tokens and cost nothing. Removing the subscription runtime leaves the metered
population:
| population | calls | prompt tok | output tok | ratio | spend |
|---|---|---|---|---|---|
| All rows with token counts | 51,917 | 221,700,073 | 28,740,244 | 7.71 : 1 | $311.43 |
Metered only (excl. codex_dispatch) | 51,900 | 117,045,118 | 28,255,229 | 4.14 : 1 | $311.43 |
codex_dispatch (subscription) | 20 | 104,654,955 | 485,015 | 215.78 : 1 | $0.00 |
4.14 : 1 is the metered pipeline’s ratio. 7.71 : 1 describes a population dominated by twenty free calls.
Per stage, restricted to June and July where token coverage is 98.6% and 99.7%:
| agent | calls | mean prompt | mean output | ratio | spend |
|---|---|---|---|---|---|
classify | 1,633 | 6,180 | 188 | 32.88 : 1 | $11.51 |
people_extract | 233 | 5,606 | 198 | 28.25 : 1 | $4.57 |
discovery | 510 | 13,249 | 495 | 26.74 : 1 | $7.98 |
triage | 2,234 | 11,689 | 595 | 19.66 : 1 | $32.43 |
chief-engineer | 1,700 | 1,628 | 98 | 16.70 : 1 | $3.56 |
product_extract | 260 | 6,181 | 533 | 11.59 : 1 | $2.28 |
product_enrich | 233 | 5,901 | 793 | 7.44 : 1 | $6.83 |
attack_event_sector | 7,139 | 744 | 101 | 7.34 : 1 | $8.84 |
backfill_cut_discriminators | 234 | 5,563 | 1,104 | 5.04 : 1 | $2.57 |
analyze | 233 | 6,454 | 1,347 | 4.79 : 1 | $15.21 |
listen_extract | 2,332 | 696 | 176 | 3.96 : 1 | $3.63 |
x_search_listener | 949 | 8,336 | 2,523 | 3.30 : 1 | $25.75 |
attack_event_enrich | 22,136 | 1,047 | 438 | 2.39 : 1 | $70.97 |
signal_scan | 261 | 5,669 | 2,610 | 2.17 : 1 | $4.84 |
signal_enricher | 1,751 | 4,424 | 3,112 | 1.42 : 1 | $34.65 |
conflict_extractor | 635 | 4,291 | 3,215 | 1.33 : 1 | $38.48 |
The ratio is a direct function of what the stage is asked to return, and the stages sort into three regimes.
Classification and routing — 16:1 to 33:1. classify, triage, people_extract and
discovery receive a document or a batch of documents and return a label, a boolean, or a short
list of names. triage sends 11,689 prompt tokens on average and returns 595. classify sends
6,180 and returns 188. Input size is set by the source material and cannot be reduced without
discarding evidence; output size is set by the schema and is already minimal. These stages are
priced almost entirely on input, so their total cost tracks the input price of the model they are
routed to and almost nothing else.
Structured extraction — 3:1 to 8:1. product_extract, product_enrich, attack_event_sector
and x_search_listener return populated records. Output scales with how much the document actually
contains.
Generation and synthesis — 1:1 to 2.5:1. conflict_extractor at 1.33:1 and signal_enricher
at 1.42:1 return nearly as many tokens as they receive. These stages cannot be made cheap by
shrinking prompts; their cost is dominated by output, which is priced higher per token on every
model in the table. conflict_extractor ran 1,264 calls for $74.56 — 7.5% of total spend on 1.5%
of calls.
The monthly metered ratio moved from 2.98:1 in April to 6.72:1 in July. The composition shifted towards classification and routing, and away from generation, as the corpus filled and the work changed from writing records to deciding whether new material was worth a record at all.
Cost per unit of output
Cost per call is a property of the model. Cost per unit of retained output is a property of the system, and the two diverge whenever a stage runs many cheap calls to produce one durable row.
| unit | count | attributed spend | cost per unit | basis |
|---|---|---|---|---|
| Company in the graph | 1,787 | $997.48 | $0.558 | total spend ÷ companies |
| Company, enrichment stages only | 1,787 | $91.65 | $0.05129 | 9 corpus agents, 8,389 calls |
| Signal collected | 55,406 | $997.48 | $0.018 | total spend ÷ signals |
| Signal, collection-to-enrichment stack | 55,406 | $210.93 | $0.003807 | 6 signal agents |
| Signal, triage decision only | 55,406 | $106.71 | $0.001926 | triage |
| Deal record | 5,287 | $262.15 | $0.04958 | business_events |
| Attack event | 7,218 | $207.07 | $0.02869 | warfare layer |
| Published article | 2,794 | $99.28 | $0.03553 | editorial layer |
The two figures given per company differ by a factor of eleven. $0.558 divides everything — including the warfare layer, the editorial layer, the credential outage and the design work — by the entity count. $0.05129 is what the nine agents that read a company and wrote structured fields onto it actually cost. Neither is wrong; they answer different questions, and the second omits the warfare layer, the editorial layer, the credential outage and the design work — an order of magnitude, itemised in the sections below.
The signal stack is the cheapest unit in the system. triage decided the fate of 55,406 signals for
$106.71, or $0.001926 each, because triage was batched: 8,584 calls covering 55,406 signals, 6.5
signals per call. 36,834 of those signals — 66.5% — ended up linked to at least one company through
39,912 rows in signal_companies. Cost per signal that survived triage and acquired a company link
is $0.005727 across the whole stack.
The editorial layer carried a second, independent cost record: artifacts.generation_cost, written
per artifact by the generator rather than derived from traces.
| artifact status | count | generation cost | per artifact | mean words |
|---|---|---|---|---|
PUBLISHED | 2,719 | $112.13 | $0.04124 | 991 |
REJECTED | 412 | $18.50 | $0.04490 | 1,356 |
ARCHIVED | 197 | $10.61 | $0.05386 | 1,184 |
| All | 3,328 | $141.25 | $0.04244 | — |
13.1% of article generation cost went to articles the pipeline itself rejected. Rejected drafts were
longer on average than published ones — 1,356 words against 991 — which is the quality gate
catching over-generation rather than under-generation. A further 526 artifacts carried a
remediation_count above zero: 310 revised once, 65 twice, 151 three times. The nominal cost of a
published article is $0.04124; the cost of the pipeline that produces publishable articles is
$0.04244 per attempt at a 81.7% acceptance rate.
Attribution
parent_type and parent_id were meant to make every dollar traceable to the thing it produced.
They do so for 48.8% of spend.
| parent_type | calls | distinct parents | spend | per parent | calls per parent |
|---|---|---|---|---|---|
| (null) | 25,272 | — | $510.85 | — | — |
artifact | 4,896 | 3,239 | $116.10 | $0.03584 | 1.51 |
attack_event | 36,944 | 121 | $104.83 | $0.86634 | 305.32 |
signal | 3,006 | 2,398 | $93.03 | $0.03880 | 1.25 |
job | 3,990 | 3,990 | $72.16 | $0.01809 | 1.00 |
listening | 3,230 | 3,058 | $47.42 | $0.01551 | 1.06 |
feed_source | 3,499 | 14 | $35.04 | $2.50308 | 249.93 |
resolution_target | 913 | 532 | $9.74 | $0.01831 | 1.72 |
research_report | 1,720 | 1,670 | $3.56 | $0.00213 | 1.03 |
company | 575 | 467 | $0.98 | $0.00209 | 1.23 |
backlog | 20 | 0 | $0.00 | — | — |
$510.85 — 51.2% of all spend — has no parent recorded. Two rows expose the cost of that: 36,944
calls are attributed to attack_event, but to only 121 distinct events, 305 calls apiece. The
attribution was written against the batch handle rather than the individual event, so the column
reports which run the call belonged to and not which record it produced. feed_source has the same
shape at 249.93 calls per parent.
Where the attribution is per-record it holds up. artifact runs 1.51 calls per artifact at $0.03584
each; artifacts.generation_cost, written by a different code path into a different table, gives
$0.04244. Two independently maintained columns agree on the cost of an article to within $0.0066.
The backlog join produces nothing. Twenty calls carry parent_type = 'backlog' and all twenty
have a null parent_id, so cost per work item cannot be computed from the trace table at all —
across 1,188 backlog rows, zero dollars are attributable to a specific task.
Burst and time of day
The system ran on timers and cron, not on human attention. Spend follows the schedule.
| hour (UTC) | calls | spend |
|---|---|---|
| 00 | 11,186 | $70.54 |
| 01 | 7,722 | $55.33 |
| 02 | 7,032 | $53.99 |
| 03 | 1,873 | $35.21 |
| 04 | 1,835 | $33.50 |
| 05 | 1,966 | $33.63 |
| 06 | 1,871 | $33.55 |
| 07 | 1,960 | $36.50 |
| 08 | 2,070 | $34.70 |
| 09 | 1,877 | $34.44 |
| 10 | 1,820 | $33.71 |
| 11 | 2,359 | $35.09 |
| 12 | 6,114 | $47.74 |
| 13 | 2,939 | $38.42 |
| 14 | 2,597 | $37.98 |
| 15 | 1,936 | $34.57 |
| 16 | 1,863 | $30.69 |
| 17 | 4,952 | $54.61 |
| 18 | 5,752 | $59.57 |
| 19 | 3,679 | $50.68 |
| 20 | 2,017 | $45.77 |
| 21 | 4,785 | $44.61 |
| 22 | 3,479 | $35.36 |
| 23 | 1,929 | $27.30 |
The 00:00 UTC hour ran 6.0× the volume of the 16:00 hour. Three peaks are visible: midnight UTC (the daily collection and publish cycle), 12:00, and a 17:00–21:00 band. The floor never reaches zero — the quietest hour of the day still carried $27.30 across the run, because the always-on workers did not stop.
Aggregate burst statistics across 2,514 active hours and 107 active days:
| measure | value |
|---|---|
| Mean calls per active hour | 34.1 |
| Median calls per active hour | 16 |
| Busiest single hour | 5,148 calls (2026-06-11 00:00 UTC, $16.74) |
| Mean spend per active day | $9.32 |
| Most expensive single day | $41.81 (2026-06-10) |
| Top decile of hours (252 h) | $433.10 — 43.4% of spend |
| Top 10 days | $302.09 — 30.3% of spend |
Median 16 calls per hour against a peak of 5,148 is a 322× spread. 43.4% of the entire bill was incurred in 252 hours — ten and a half days’ worth of clock time out of 107 active days. A budget alarm sampling hourly averages would have detected none of it.
By weekday:
| day | calls | spend |
|---|---|---|
| Sunday | 14,805 | $134.39 |
| Monday | 7,032 | $118.60 |
| Tuesday | 7,373 | $138.04 |
| Wednesday | 18,183 | $173.61 |
| Thursday | 22,088 | $171.67 |
| Friday | 7,587 | $127.83 |
| Saturday | 8,545 | $133.33 |
Spend varies by 1.5× across days of the week and call volume by 3.1×. Weekends are not cheaper. The schedule, not the operator, determined the shape.
Total model runtime, summing latency_ms across every row, was 164.6 hours — 155.8 of them on
calls that returned status='ok'. Against a 107-day window that is 6.4% duty cycle on a single
serial timeline, achieved by a machine with four vCPUs running four always-on workers, five systemd
timers and nine cron jobs concurrently.
Spend on deleted layers
| layer | calls | spend | share of total |
|---|---|---|---|
Warfare intelligence (attack_*, conflict, CARVER) | 38,441 | $207.07 | 20.8% |
| Editorial (2,794 articles) | 2,873 | $99.28 | 10.0% |
| Combined | 41,314 | $306.35 | 30.7% |
Both layers functioned as built. Both were removed at teardown — the warfare layer for export-control exposure and scope reduction, the editorial layer because a maintained news surface was not wanted.
30.7% of total model spend produced output that was subsequently deleted.
Per unit, the deleted layers were the cheapest output the system produced: $0.02869 per attack event across 7,218 events, and $0.03553 per published article. Neither was expensive to make. Both were deleted because the corpus they built was not worth maintaining, which is a judgment about value rather than about unit cost, and no cost metric available inside the trace table would have predicted it.
Error composition
3,105 calls returned status='error', a 3.63% error rate, costing $20.74.
| error | count | share of errors |
|---|---|---|
401 Unauthorized (OpenRouter) | 2,185 | 70.4% |
403 Forbidden | 265 | 8.5% |
403 Key limit exceeded (daily limit) | 246 | 7.9% |
AttributeError: 'NoneType' object has no attribute 'strip' | 15 | 0.5% |
json_parse_failed | 7 | 0.2% |
A 3.63% error rate reads as within tolerance. The same data grouped by error string identifies one fixable configuration defect responsible for the majority of failures. Rate and composition are different measurements and the second is the actionable one.
Grouped by cost rather than by count, the ordering inverts:
| error class | count | direct cost | mean latency | wall time consumed |
|---|---|---|---|---|
401 Unauthorized | 2,185 | $0.0000 | 263 ms | 0.16 h |
403 Forbidden | 265 | $0.0000 | 159 ms | 0.01 h |
403 Key limit exceeded | 246 | $0.0000 | 2,428 ms | 0.17 h |
AttributeError: NoneType | 15 | $0.0000 | 6,929 ms | 0.03 h |
| other (404, 503, timeouts) | 22 | $0.0000 | 261,464 ms | 1.60 h |
json_parse_failed | 372 | $19.6958 | 53,579 ms | 5.54 h |
The transport failures cost nothing. A rejected request is rejected before tokens are billed, so all
2,185 401s, all 511 403s and every timeout in the table have cost_usd = 0. 100% of the
$19.6958 booked against status='error' is json_parse_failed — calls that ran to completion,
generated and billed output tokens, and returned text the parser could not use. 372 such calls at a
mean latency of 53,579 ms consumed 5.54 hours of wall time and produced nothing.
json_parse_failed concentrates in three agents:
| agent | model | count | cost | mean latency |
|---|---|---|---|---|
signal_enricher | claude-haiku-4.5 | 150 | $8.7004 | 25,803 ms |
qa_remediation | claude-haiku-4.5 | 108 | $5.7845 | 70,306 ms |
content_remediation | claude-haiku-4.5 | 110 | $4.4023 | 71,928 ms |
conflict_extractor | claude-sonnet-4-6 | 2 | $0.4421 | 143,110 ms |
cluster_conflict | claude-sonnet-4.5 | 1 | $0.3664 | 265,293 ms |
All three of the large contributors are generation-regime stages routed to Haiku — the same stages that sit at 1.3:1 to 1.4:1 in the token table. The failure mode is a long free-text response where strict JSON was required, and it is the direct cost of tiering a schema-constrained task down to a cheaper model without a constrained-decoding path.
A separate retry status records the recovery attempts. 16 calls carry status='json_retry' at a
combined $1.0400 and a mean latency of 55,231 ms — a second full call, at full price, for one
malformed first response. $19.6958 of failed parses plus $1.0400 of retries reconciles to the
$20.74 total. A further 211 calls carry status='success', all of them x_search_listener on
grok-4.3 at $5.6539: a second status string for the same outcome, from a code path that never
adopted the shared vocabulary.
Latency distributions by status show what a failure costs in time rather than in dollars:
| status | count | p50 | p95 | max |
|---|---|---|---|---|
ok | 82,281 | 3,805 ms | 25,331 ms | 1,644,894 ms |
error | 3,105 | 133 ms | 64,337 ms | 1,800,034 ms |
json_retry | 16 | 65,626 ms | 96,848 ms | 97,968 ms |
success | 211 | 15,735 ms | 32,098 ms | 40,915 ms |
The error median of 133 ms is the credential rejections returning immediately. The error p95 of 64,337 ms and the 30-minute maximum are the parse failures and timeouts. Errors are bimodal: a failure is either 200× faster than a success or 17× slower, and averaging them produces a number that describes neither.
The 401s were a single credential misconfiguration confined to one window: first at
2026-05-21 13:05:20 UTC, last at 2026-05-24 18:24:33 UTC, 2,185 calls over four days, spanning
twelve agents — triage (924), business_events (462), editorial_review (231),
campaign_sender (154), signal_enricher (95) and seven others. Direct cost was $0.00. The
indirect cost was two entire days on which every pipeline stage executed, failed, and wrote nothing
to the graph. May’s 8.55% error rate is that window; strip 21–24 May out and the remaining 27 days
of May run at 1.49%.
Nothing detected it for four days. The traces recorded every failure faithfully in real time, with correct statuses and correct error strings. What did not exist was anything that read them.
Infrastructure
Model spend was instrumented per call. Infrastructure was not, and invoices were not read at teardown. What follows is list price against the documented specification, not measured spend — it establishes the order of magnitude and nothing more.
The host, read from the running machine:
| property | measured value |
|---|---|
| CPU | AMD EPYC-Milan, 4 cores |
| Memory | 16 GB |
| Root filesystem | 75 GB |
| Product | vServer (Hetzner Cloud) |
The Hetzner Cloud plan matching 4 AMD vCPU / 16 GB is CCX23 (4 dedicated AMD vCPU, 16 GB RAM, 160 GB NVMe). Hetzner adjusted CCX pricing on 15 June 2026, mid-build:
| plan | vCPU / RAM | list price to 15 Jun 2026 | list price from 15 Jun 2026 |
|---|---|---|---|
| CCX13 | 2 / 8 GB | €15.99 / month | €42.99 / month |
| CCX23 | 4 / 16 GB | €31.49 / month | €85.99 / month |
| CCX33 | 8 / 32 GB | €62.49 / month | €138.49 / month |
All prices excluding VAT. Hetzner’s notice states the adjustment “applies to new orders and cloud instance rescales starting from 15 June 2026; 8 AM CEST” — orders placed earlier retain the previous rate. At the pre-adjustment rate, four months of CCX23 is €125.96 at list; at the post-adjustment rate, €343.96. The same specification ordered on either side of one Monday morning differs by 2.73×.
The database was Supabase Postgres 17.6 with PostGIS 3.3.7 and pgvector 0.8.0, 2,031 MB at teardown. Supabase list pricing:
| plan | monthly | database disk included | overage | egress included | overage |
|---|---|---|---|---|---|
| Free | $0 | 500 MB | — | 5 GB | — |
| Pro | $25 | 8 GB per project | $0.125 / GB | 250 GB | $0.09 / GB |
| Team | $599 | 8 GB per project | $0.125 / GB | 250 GB | $0.09 / GB |
At 2.0 GB the database was 4.1× the Free tier’s 500 MB ceiling and 25% of the Pro tier’s 8 GB
inclusion, so it sat inside Pro at $25/month with no storage overage — $100 at list over four
months. 470 MB of that 2,031 MB was eng_llm_traces.
Object storage and hosting stayed inside free tiers. Cloudflare R2 lists a permanent free allowance of 10 GB-month of Standard storage, 1 million Class A operations and 10 million Class B operations per month, with egress free; beyond that, $0.015 per GB-month, $4.50 per million Class A and $0.36 per million Class B. Cloudflare Pages hosted the site on its free plan throughout.
| component | specification | list price |
|---|---|---|
| VPS | Hetzner CCX23 — 4 AMD vCPU / 16 GB / 160 GB NVMe | €31.49–€85.99 / month, ex VAT |
| Database | Supabase Pro — Postgres 17.6, PostGIS 3.3.7, pgvector 0.8.0, 2.0 GB | $25 / month |
| Object storage | Cloudflare R2, under the 10 GB-month free allowance | $0 |
| Hosting | Cloudflare Pages, free plan | $0 |
| Providers | ~15 API keys, several on paid tiers | not enumerated |
Four months of VPS and database at list price is €125.96 + $100 at the pre-June rate. Against $997.48 of model spend, infrastructure at list is roughly one quarter of the bill. It is stated at list because it was never measured, and it is stated separately because the two behave differently: model spend reached zero at the moment workers were stopped; infrastructure cost persists until each service is cancelled.
Post-teardown running cost: Cloudflare Pages free tier, R2 below the 10 GB free threshold, and domain registration.
Reconstruction estimate
For the same corpus scope, built with the tiering and deterministic-first rules applied from the start rather than from month three:
| component | estimate |
|---|---|
| Corpus construction (excluding deleted layers) | $600–700 |
| Design and review (Opus-tier) | $150–250 |
| Infrastructure, 4 months | a few hundred |
Total under $1,000 for a provenance-tracked graph of ~1,800 entities with products, people, deals and sourcing.
Four measured findings bound where a rebuild’s savings would and would not come from. April and May
ran 41,315 calls for $702.93; the same call count at June’s post-tiering mean of $0.00614 would have
cost $253.67, an excess of $449.26. The single largest line item, business_events at
$262.15, was one agent on one model for twenty-seven days, and is a routing decision. The
credential outage cost $0.00 in direct spend and two days of throughput, which no budget change
recovers. And 30.7% of the bill bought layers that were deleted — a scoping decision taken at
teardown that could have been taken at the start, and the only one of the four large enough to
change the total on its own.
Next: Part 6 — Results.