published Aug 26, 2026, 10:04 AM · updated Aug 26, 2026, 10:04 AM
Overview
This report forecasts github/gh-aw agentic workflow spending using AI Inference
Cost (AIC) data sampled over a 30-day history window (as of 2026-08-26T09:56:23Z,
forecast period: month). Of 50 workflows discovered, 40 had at least one sampled
run with AIC data; those 40 are the basis for all totals below. 10 workflows had
zero sampled runs and are excluded from spend totals (see Data Quality section).
Executive summary
- Total observed AIC (sum of all sampled agentic-run costs, 30-day window):
$44,336.56 across 1,004 sampled runs, 40 active workflows. - Weekly forecast (sum across all active workflows' weekly Monte Carlo bands):
- P10 (10th percentile — optimistic scenario: 9 out of 10 weeks will cost at
least this much): $2,971.05 - P50 (50th percentile — median/expected scenario: equal probability of
spending more or less than this): $7,694.94 - P90 (90th percentile — conservative scenario: only 1 out of 10 weeks is
expected to exceed this amount): $14,265.48
- P10 (10th percentile — optimistic scenario: 9 out of 10 weeks will cost at
- Monthly forecast (sum across all active workflows' monthly Monte Carlo bands):
- P10 (optimistic): $21,538.53
- P50 (median): $34,451.77
- P90 (conservative): $50,084.93
- Top spend drivers: Go Logger Enhancement ($9,219.6 observed, 31 runs, 71.0%
success), Agentic Workflow Audit Agent ($7,647.5, 30 runs, 80.0% success), and
Semantic Function Refactoring ($6,114.8, 31 runs, 83.9% success) — together
~52% of total observed spend.
Charts
Key metrics — workflow forecast table
Sorted by total observed AIC, descending. All AIC values in USD-equivalent cost units.
"Weekly proj (P10–P90)" is the weekly Monte Carlo (10,000 iterations) confidence range.
| Workflow | Sampled runs | Observed AIC | P50/run | P95/run | Weekly proj (P10–P90) | Monthly proj | Success rate |
|---|---|---|---|---|---|---|---|
| Go Logger Enhancement | 31 | 9219.6 | 335.6 | 466.6 | 2151.3 (586–2608) | 9219.6 | 71.0% |
| Agentic Workflow Audit Agent | 30 | 7647.5 | 272.9 | 408.0 | 1784.4 (576–2381) | 7647.5 | 80.0% |
| Semantic Function Refactoring | 31 | 6114.8 | 212.3 | 315.5 | 1426.8 (500–1998) | 6114.8 | 83.9% |
| CLI Version Checker | 32 | 4165.4 | 128.1 | 300.6 | 971.9 (334–1372) | 4165.4 | 84.4% |
| Tidy | 32 | 2184.9 | 67.1 | 150.6 | 509.8 (147–648) | 2184.9 | 75.0% |
| Smoke Copilot | 48 | 2025.8 | 47.3 | 82.6 | 472.7 (83–389)* | 2025.7 | 47.9% |
| Copilot Agent PR Analysis | 29 | 1911.6 | 75.0 | 94.5 | 446.0 (150–625) | 1911.6 | 82.8% |
| Lockfile Statistics Analysis Agent | 30 | 1732.3 | 56.9 | 88.1 | 404.2 (172–618) | 1732.3 | 93.3% |
| Duplicate Code Detector | 30 | 1396.7 | 17.7 | 141.4 | 325.9 (55–460) | 1396.7 | 73.3% |
| Dev | 32 | 1373.8 | 22.3 | 101.8 | 320.6 (71–438) | 1373.8 | 75.0% |
| Smoke Claude | 14 | 996.9 | 75.9 | 110.7 | 232.6 (61–416) | 996.9 | 92.9% |
| Daily News | 24 | 969.3 | 36.6 | 83.5 | 226.2 (85–373) | 969.3 | 95.8% |
| Terminal Stylist | 33 | 876.3 | 26.4 | 47.4 | 204.5 (82–291) | 876.3 | 87.9% |
| GitHub MCP Remote Server Tools Report Generator† | 4 | 835.3 | 238.4 | 248.7 | 194.9 (0–497) | 835.3 | 100% |
| Scout | 16 | 693.0 | 0 | 131.4 | 161.7 (0–206) | 692.9 | 43.8% |
| Weekly Workflow Analysis† | 5 | 607.7 | 137.7 | 214.7 | 141.8 (0–352) | 607.7 | 80.0% |
| Documentation Unbloat | 32 | 459.3 | 14.5 | 23.7 | 107.2 (48–163) | 459.3 | 93.8% |
| Smoke Codex | 19 | 310.9 | 6.5 | 56.0 | 72.5 (0–97) | 310.9 | 52.6% |
| Daily Documentation Updater | 31 | 307.1 | 0 | 30.4 | 71.7 (21–126) | 307.1 | 93.5% |
| Weekly Issue Summary† | 6 | 288.9 | 39.7 | 113.5 | 67.4 (0–113) | 288.9 | 50.0% |
| Artifacts Usage Report† | 6 | 120.3 | 0 | 46.5 | 28.1 (0–47) | 120.3 | 50.0% |
| Repository Tree Map Generator† | 7 | 79.6 | 17.4 | 26.1 | 18.6 (0–26) | 79.6 | 42.9% |
| Plan Command† | 2 | 19.6 | 0 | 19.6 | 4.6 (0–20) | 19.6 | 50.0% |
| Copilot Setup Steps, Commit Changes Analyzer†, Copilot cloud agent, MCP Inspector Agent†, Mergefest†, Notion Issue Summary†, Poem Bot†, Rebuild the documentation†, Resource Summarizer Agent†, Doc Build - Deploy, Go Pattern Detector, CodeQL, Dependabot Updates, CI, Basic Research Agent†, Video Analysis Agent†, Dev Hawk† | 1–100 | 0 | 0 | 0 | 0 | 0 | 0–100% |
* Smoke Copilot's weekly P10–P90 band (83–389) does not bracket its point weekly
projection (472.7) — a Monte Carlo sampling artifact from its high run-frequency
(48 sampled runs) combined with high per-run AIC variance; treated as low-confidence.
† Workflow has fewer than 8 sampled runs and/or an unreliable Monte Carlo flag (see
Data Quality below) — treat its weekly/monthly projections as directional only.
Data quality and accuracy
Zero-AIC workflows (28 of 40 active workflows show $0 observed spend)
Many workflows (CodeQL, CI, Doc Build - Deploy, Copilot cloud agent,
Dependabot Updates, Go Pattern Detector, Copilot Setup Steps, etc.) show
sampled_runs > 0 but avg_aic = 0 and all run_samples[].aic = 0. These are
standard (non-agentic) CI/CD or dependency-management workflows that do not invoke
an AI engine, so $0 AIC is expected and correct, not a data gap. A smaller
group (Commit Changes Analyzer, MCP Inspector Agent, Mergefest, Notion Issue Summary, Poem Bot, Rebuild the documentation..., Resource Summarizer Agent,
Basic Research Agent, Video Analysis Agent, Dev Hawk) show 0% success rate
alongside $0 AIC — this looks like workflow runs that failed before invoking the
AI engine (e.g., setup/config errors) rather than genuinely free runs. This
distinction could not be verified further without pulling individual run logs,
which was out of scope for this forecast; flagging for follow-up rather than
inventing a corrected value.
Zero-sampled-run workflows (10 workflows excluded from totals)
Q, .github/workflows/test-proxy, Sentry Issue Analyzer, CI Failure Doctor,
Smoke OpenCode, Test, Test Claude, Test Copilot CLI Engine, Test Copilot GitHub Integration, and Format, Lint, Build and Commit returned sampled_runs: 0,
avg_aic: 0, and monte_carlo: null — no runs occurred in the 30-day window (likely
disabled, manually-triggered-only, or newly added workflows). They are excluded from
all totals and forecasts above since there is no evidence to project from. No
follow-up evidence collection was attempted for these, as a genuine absence of runs
in a fixed 30-day window is the most likely and unremarkable explanation.
Low-sample and unreliable Monte Carlo workflows (16 flagged)
16 of 40 active workflows have is_reliable: false on their (weekly and/or monthly)
Monte Carlo simulation, generally correlating with fewer than 8 sampled runs in the
30-day window: GitHub MCP Remote Server Tools Report Generator (4), Weekly Workflow Analysis (5), Weekly Issue Summary (6), Artifacts Usage Report (6),
Repository Tree Map Generator (7), Commit Changes Analyzer (2), MCP Inspector Agent (7), Mergefest (1), Notion Issue Summary (1), Plan Command (2), Poem Bot (2), Rebuild the documentation... (1), Resource Summarizer Agent (2), Basic Research Agent (2), Video Analysis Agent (1), Dev Hawk (2). Their weekly/monthly
projections are included in the totals above for completeness but should be treated
as low-confidence directional estimates only — sample sizes of 1–7 runs are too
sparse to support a reliable budget decision. No rerun of gh aw forecast was
performed for these, since the underlying constraint (few runs in 30 days) cannot be
resolved by re-sampling the same window; a longer history window would be needed to
improve confidence.
History window and consistency checks
All 40 active workflows report a consistent history_days: 30 window, and period: month throughout — no inconsistent date windows detected. Run counts per workflow
(2–100 sampled runs) are plausible given differing trigger cadences (scheduled daily
jobs vs. manually-triggered or PR-triggered workflows); no implausible run
frequencies were found. No AIC value in the sampled data was negative or obviously
corrupted. forecast.json parsed cleanly with 50 top-level workflow entries and no
schema errors.
Assumptions
- Forecast date: 2026-08-26 (as_of:
2026-08-26T09:56:23Z). - History window: 30 days, consistent across all workflows.
- Totals and charts include only the 40 workflows with at least one sampled run;
the 10 zero-sampled workflows are excluded from spend aggregates (see above). - Weekly/monthly bands are the sum of each workflow's independent Monte Carlo
(10,000-iteration) percentile projections; summing percentiles across workflows
is an approximation (it assumes the projections are additive at each percentile
rank) and is not a joint simulation across the full portfolio — treat portfolio-
level P10/P90 as indicative, not statistically exact. - No original
forecast.jsonvalues were modified, imputed, or fabricated;
all flagged issues are reported as-is.
Next actions
- Investigate the 0%-success, $0-AIC workflows (
Commit Changes Analyzer,MCP Inspector Agent,Mergefest,Notion Issue Summary,Poem Bot,Rebuild the documentation...,Resource Summarizer Agent,Basic Research Agent,Video Analysis Agent,Dev Hawk) to confirm whether they are failing before AI
invocation or are intentionally cost-free. - Extend the history window (e.g., 60–90 days) for the 16 low-sample/unreliable
workflows to improve Monte Carlo confidence before using their projections in
budget decisions. - Monitor
Smoke Copilot's Monte Carlo band inconsistency (P10–P90 not bracketing
the point weekly projection) on the next forecast run to see if it resolves with
more samples. - Track
Go Logger Enhancement,Agentic Workflow Audit Agent, andSemantic Function Refactoringclosely, as they represent the largest and most reliable
cost drivers (52% of observed spend, all with ≥30 samples and reliable Monte
Carlo bands).
Reference: §32955167587
Generated by 📈 Daily Spending Forecast · copilot · auto · 42.7 AIC · ⌖ 5.45 AIC · ⊞ 11.3K · ◷
- expires on Sep 2, 2026, 2:04 AM UTC-08:00

