GCP project electro_x_evolve · invoice window 2026-03-10 → 2026-07-12
· reconciled against 112 Gemini runs in outputs/
electro_x_evolve, 4 models, 242 SKU-day rowsOrdered by how much they change the number. The first two are established by direct evidence; the rest bound what the traces cannot show.
The invoice bills gemini-2.5-flash
($14,489.78) and
gemini-2.5-flash-lite
($63.38). All 1,655 run configs
under outputs/ were scanned, plus every source file in the repository: no
gemini-2.5-* model appears anywhere except unused rate constants in three
pricing.py files, and no model id is ever built dynamically. That spend also runs
Mar 28 – Jul 12, almost entirely after this project's Gemini activity
ended (last configured run 2026-04-27). This is the one figure that does not depend on
date alignment — it is a model-identity argument, so no timezone or clock assumption
can move it.
Of our models' spend, $21,049.39 is fresh input and $12,402.50 is cached input, against only $5,133.48 of output (13.3%). Gemini bills reasoning as output, so even if every billed output token were a thinking token, the ceiling on that explanation is $5,133.48. The actual driver is volume: 15.25B fresh plus 68.29B cached input tokens, which is what repeatedly re-sending retrieved context costs at Gemini 3 Pro's $2–$4 per 1M. ReAct pipelines re-send a growing transcript every step and map_reduce re-sends every chunk, so input is paid many times per query.
The fields total_input_tokens / total_output_tokens equal
len(text.split()) exactly — verified on 102 of 102 map_reduce queries in a
gemini-3.1-pro run (ratio 1.000; a chars/4 estimate gave 0.585). Recording is also inconsistent:
the vrag shape records no input tokens at all, and map_reduce records no token fields whatsoever.
Any cost figure previously derived from these fields is an undercount.
Re-tokenizing every prompt and every piece of generated text we stored yields a floor of
3.52B input and 37.01M output tokens, against 83.55B
and 530.24M billed. No raw API usage object was persisted in any of the 112
runs, so the gap cannot be decomposed from traces alone. It is some mixture of cached context
re-sends, retried and failed calls, runs whose outputs were never saved, and thinking tokens.
We can bound but not attribute it.
Gemini prices requests above a per-request input threshold higher. Rates derived from this
invoice: Gemini 3 Pro input $2.00→$4.00 per 1M (2.00×) and output
$12.00→$18.00 (1.50×). dashboard/pricing.py encodes only the short-tier
rates, so it understates every long-context call. This is a context-length tier, not a
concurrency charge — Gemini has no concurrency-based pricing.
Our run timestamps are UTC, but the billing export does not record its own timezone, so an exact-day join can misplace spend by up to a day. Under exact-day matching, $3,111.97 of our-model spend falls on days with no saved run; allowing ±1 day that falls to $654.19, and ±2 days to $37.64. The honest statement is therefore a band: $654–$3,112 unattributed, not a point estimate. The total charged to our models ($38,585.37) is unaffected by this choice.
Confirmed by the team: a colleague ran this BrowseComp-Plus work on the same project API key.
What is stored in outputs/browsecomp_plus/ is only the ingest — all 12
directories are stamped within a 90-second window on 2026-05-06 and carry a source
block pointing at browsecomp_results/runs/, which is owned by another uid and
unreadable here. The generation billed on an earlier date we cannot recover: the traces carry
timestamps: null, and the invoice has no Gemini charges anywhere in May
(Apr 27 → May 23 is empty). Its dollars sit inside the our-models total;
they are only absent from the run-matched row. Workload floor: 5,736 queries,
896.28M input and 10.03M output tokens, about $544 at
short-tier rates.
Three runs on 2026-07-27 routed google/gemini-* through
openrouter.ai, which bills separately. The invoice ending 2026-07-12 with no later
Gemini charges independently confirms provider routing was read correctly.
Each browsecomp_plus run carries three llm_judge_report_*.json files,
which reads as three billed judge passes on gemini-3-pro-preview. They are not: the
model-generated fields are byte-identical across all three (same decisions, same explanations,
same hash), so the evals were copied in and re-stamped, and the extra files cost nothing. The
hazard is that nothing on disk distinguishes a copied eval from a real re-run, which would look
identical and would silently multiply cost.
Daily invoiced spend, split by whether the billed model is one this project ever configured. Spend is overwhelmingly concentrated in March 2026 ($48,660.55, 92% of the total); the single worst day is 2026-03-25 at $8,969.91.
The same data accumulated, on its own axis — the March ramp is where the incident actually happened.
Billed token volume for our two models, taken from invoice SKUs rather than our instrumentation — these are the only trustworthy token counts available. Each panel keeps its own axis because the three kinds differ by two orders of magnitude: cached input dominates volume while output dominates cost.
| Billed model | Cost | Tokens | Configured in this project? |
|---|---|---|---|
| gemini-3-pro | $35,843.49 | 47.18B | ours |
| gemini-2.5-flash | $14,489.78 | 7.33B | NOT ours |
| gemini-3.1-flash-lite | $2,741.88 | 36.89B | ours |
| gemini-2.5-flash-lite | $63.38 | 638.19M | NOT ours |
Invoice SKUs name Gemini 3 Pro as
“gemini 3 pro”; our configs call the same model
gemini-3.1-pro-preview, and our own run directories use
gemini_3_pro_preview, which corroborates the mapping.
Rates below are derived from this invoice (cost ÷ tokens per SKU), not from a maintained price table.
| Gemini 3 Pro | Short-context $/1M | Long-context $/1M | Premium |
|---|---|---|---|
| input | $2.00 | $4.00 | 2.00× |
| cached input | $0.20 | $0.40 | 2.00× |
| output | $12.00 | $18.00 | 1.50× |
Full reconciliation table — the table view backing every chart above.
| Date | Our models $ | Other models $ | Total $ | Cumulative $ | Runs | Fresh in | Cached in | Output |
|---|---|---|---|---|---|---|---|---|
| 2026-03-10 | $3.09 | $0.00 | $3.09 | $3.09 | 0 | 1.21M | 0 | 275.89K |
| 2026-03-11 | $3.10 | $0.00 | $3.10 | $6.19 | 0 | 551.68K | 0 | 469.98K |
| 2026-03-12 | $3.14 | $0.00 | $3.14 | $9.33 | 8 | 616.13K | 0 | 411.91K |
| 2026-03-13 | $27.34 | $0.00 | $27.34 | $36.67 | 0 | 15.37M | 57.20K | 970.50K |
| 2026-03-14 | $15.19 | $0.00 | $15.19 | $51.86 | 0 | 20.46M | 85.80K | 1.61M |
| 2026-03-15 | $0.00 | $0.00 | $0.00 | $51.86 | 4 | 0 | 0 | 0 |
| 2026-03-16 | $48.07 | $0.00 | $48.07 | $99.93 | 2 | 197.37K | 0 | 3.99M |
| 2026-03-17 | $111.09 | $0.00 | $111.09 | $211.02 | 0 | 28.46M | 0 | 6.72M |
| 2026-03-18 | $0.00 | $0.00 | $0.00 | $211.02 | 6 | 0 | 0 | 0 |
| 2026-03-19 | $234.71 | $0.00 | $234.71 | $445.73 | 4 | 44.36M | 0 | 15.46M |
| 2026-03-20 | $1,897.51 | $0.00 | $1,897.51 | $2,343.24 | 21 | 1.77B | 8.68B | 18.83M |
| 2026-03-21 | $5,877.29 | $326.53 | $6,203.82 | $8,547.06 | 5 | 3.13B | 17.17B | 43.05M |
| 2026-03-22 | $6,174.78 | $544.92 | $6,719.70 | $15,266.76 | 14 | 2.60B | 12.13B | 56.52M |
| 2026-03-23 | $7,645.88 | $149.64 | $7,795.52 | $23,062.28 | 21 | 3.40B | 10.62B | 101.74M |
| 2026-03-24 | $7,313.65 | $978.88 | $8,292.53 | $31,354.81 | 6 | 1.85B | 9.44B | 76.74M |
| 2026-03-25 | $6,144.56 | $2,825.35 | $8,969.91 | $40,324.72 | 3 | 1.19B | 7.21B | 50.27M |
| 2026-03-26 | $2,028.74 | $1,631.06 | $3,659.80 | $43,984.52 | 0 | 687.88M | 2.32B | 78.46M |
| 2026-03-27 | $613.46 | $501.42 | $1,114.88 | $45,099.40 | 0 | 411.07M | 620.12M | 53.06M |
| 2026-03-28 | $0.00 | $363.06 | $363.06 | $45,462.46 | 0 | 0 | 0 | 0 |
| 2026-03-29 | $0.00 | $1,307.07 | $1,307.07 | $46,769.53 | 0 | 0 | 0 | 0 |
| 2026-03-30 | $0.00 | $1,527.34 | $1,527.34 | $48,296.87 | 0 | 0 | 0 | 0 |
| 2026-03-31 | $0.00 | $363.68 | $363.68 | $48,660.55 | 0 | 0 | 0 | 0 |
| 2026-04-01 | $0.00 | $11.21 | $11.21 | $48,671.76 | 0 | 0 | 0 | 0 |
| 2026-04-11 | $6.97 | $148.65 | $155.62 | $48,827.38 | 0 | 2.90M | 0 | 936.27K |
| 2026-04-12 | $0.00 | $310.04 | $310.04 | $49,137.42 | 0 | 0 | 0 | 0 |
| 2026-04-13 | $9.51 | $261.39 | $270.90 | $49,408.32 | 0 | 894.02K | 16.47M | 64.63K |
| 2026-04-14 | $1.26 | $256.45 | $257.71 | $49,666.03 | 0 | 1.51M | 0 | 563.17K |
| 2026-04-15 | $19.90 | $22.26 | $42.16 | $49,708.19 | 0 | 8.03M | 5.86M | 1.30M |
| 2026-04-18 | $0.86 | $451.38 | $452.24 | $50,160.43 | 1 | 3.14M | 1.60M | 23.86K |
| 2026-04-19 | $12.94 | $108.96 | $121.90 | $50,282.33 | 2 | 703.19K | 216.08K | 1.09M |
| 2026-04-20 | $70.80 | $0.00 | $70.80 | $50,353.13 | 2 | 315.48K | 0 | 6.24M |
| 2026-04-21 | $49.21 | $0.00 | $49.21 | $50,402.34 | 3 | 8.11M | 4.34M | 2.69M |
| 2026-04-22 | $0.00 | $0.00 | $0.00 | $50,402.34 | 1 | 0 | 0 | 0 |
| 2026-04-23 | $0.00 | $71.25 | $71.25 | $50,473.59 | 0 | 0 | 0 | 0 |
| 2026-04-24 | $0.00 | $420.19 | $420.19 | $50,893.78 | 0 | 0 | 0 | 0 |
| 2026-04-25 | $0.00 | $473.75 | $473.75 | $51,367.53 | 0 | 0 | 0 | 0 |
| 2026-04-26 | $272.32 | $688.19 | $960.51 | $52,328.04 | 0 | 76.05M | 75.90M | 8.75M |
| 2026-04-27 | $0.00 | $434.39 | $434.39 | $52,762.43 | 1 | 0 | 0 | 0 |
| 2026-05-23 | $0.00 | $152.33 | $152.33 | $52,914.76 | 0 | 0 | 0 | 0 |
| 2026-06-11 | $0.00 | $22.80 | $22.80 | $52,937.56 | 0 | 0 | 0 | 0 |
| 2026-06-12 | $0.00 | $26.34 | $26.34 | $52,963.90 | 0 | 0 | 0 | 0 |
| 2026-06-13 | $0.00 | $14.24 | $14.24 | $52,978.14 | 0 | 0 | 0 | 0 |
| 2026-07-11 | $0.00 | $103.79 | $103.79 | $53,081.93 | 0 | 0 | 0 | 0 |
| 2026-07-12 | $0.00 | $56.60 | $56.60 | $53,138.53 | 0 | 0 | 0 | 0 |
The invoice is treated as ground truth. Traces are used to attribute it, and to establish a floor — never to replace it.
What this cannot establish. Thinking tokens, cached-input volume, and retried or failed calls were never recorded, so the bottom-up floor covers only 4.4% of billed output and cannot be used to price the incident. The $3,111.97 billed on our models on days with no saved run is genuinely ambiguous: it may be unsaved or failed runs of ours (the browsecomp_plus generation is a known candidate, since it definitely ran on this key on a date we cannot recover), or further third-party use. Resolving the remainder requires per-request API logs or Cloud Logging, which the billing export does not contain.
Generated 2026-08-12 18:13 UTC from
sky-gcp-electro_x_evolve-20260310-20260811.csv and
/data_large/electro_results/outputs.
Reproduce with python -m gemini_audit.audit then
python -m gemini_audit.reconcile.