[spending-forecast] Daily spending forecast - 2026-08-25

Daily Spending Forecast · issue · closed

Filter2mode:review mode:live
All recorded Export JSON
github-actions[bot]

published Aug 25, 2026, 9:59 AM · updated Aug 26, 2026, 10:04 AM

Overview

This report forecasts spending for github/gh-aw agentic workflows based on a
30-day history window (2026-07-26 to 2026-08-25) as of 2026-08-25T09:52:46Z.
Terminology used throughout: P10 (10th percentile — optimistic scenario: 9
out of 10 months will cost at least this much), P50 (50th percentile —
median/expected scenario), P90 (90th percentile — conservative scenario:
only 1 out of 10 months is expected to exceed this amount).

Of 50 tracked workflows, 25 had sampled runs in the window and 25 had zero
sampled runs
(inactive, disabled, or not triggered — see Data Quality section).

Executive summary (active workflows only):

  • Total observed AIC over 30 days (429 sampled runs, 25 workflows): 43,909.0
  • Weekly forecast: P10 = 3,299 · P50 = 8,571 · P90 = 15,799
  • Monthly forecast: P10 = 23,710 · P50 = 38,089 · P90 = 56,064

Charts

Spending Trend — per-run AIC over the last 30 days for the top 5 workflows by total AIC, with 7-day rolling average overlay

Weekly Forecast Distribution — P10 (optimistic), P50 (median), P90 (conservative) projected weekly AIC for the top 10 workflows by median forecast

Key metrics

Top spenders by 30-day observed AIC: Go Logger Enhancement (9,048.0), Agentic
Workflow Audit Agent
(7,402.9), Semantic Function Refactoring (5,966.3), CLI
Version Checker
(4,036.9), Tidy (2,140.3). These five account for ~65% of total
observed spend.

Notable data-quality flags (see next section for detail): 7 workflows have
Monte Carlo simulations marked unreliable (typically due to very small sample
counts of 1–4 runs), and 25 workflows show zero sampled runs.

Workflow table (active workflows, sorted by observed AIC)

Workflow Samples Observed AIC P50/run P90/run Weekly P50 Monthly P50 Success Rate Weekly Range (P10–P90)
Go Logger Enhancement 28 9048.0 355.2 466.6 1535.8 6756.9 75% 620–2713
Agentic Workflow Audit Agent 26 7402.9 281.8 408.0 1501.6 6547.5 88% 635–2578
Semantic Function Refactoring 28 5966.3 212.3 315.5 1093.9 4917.1 82% 461–1923
CLI Version Checker 27 4036.9 132.6 300.6 873.3 3886.9 96% 385–1527
Tidy 29 2140.3 70.7 150.6 394.5 1773.7 83% 169–704
Smoke Copilot 37 2025.8 52.1 86.4 286.6 1255.9 62% 129–485
Copilot Agent PR Analysis 24 1853.2 77.9 94.5 403.0 1781.1 96% 174–702
Lockfile Statistics Analysis Agent 28 1668.0 56.9 88.1 366.6 1619.0 96% 170–615
Duplicate Code Detector 24 1414.3 55.3 141.4 298.5 1350.1 96% 100–568
Dev 25 1394.6 57.4 101.8 306.9 1405.4 100% 118–565
Smoke Claude 15 1197.9 77.3 110.7 266.8 1202.0 100% 79–516
Daily News 22 909.9 36.6 83.5 204.7 908.8 100% 83–364
Terminal Stylist 31 896.6 27.4 47.4 197.9 867.8 97% 100–324
GitHub MCP Remote Server Tools Report Generator 4 835.3 238.4 248.7 238.4 830.9 100% 0–497 ⚠️
Scout 7 693.0 109.3 131.4 131.4 693.7 100% 0–372 ⚠️
Weekly Workflow Analysis 4 607.7 137.7 214.7 137.7 607.7 100% 0–393 ⚠️
Documentation Unbloat 30 451.4 14.5 23.7 98.8 436.5 97% 48–164
Smoke Codex 12 346.9 21.1 56.0 64.7 318.8 92% 6–156
Daily Documentation Updater 13 327.0 24.9 33.8 73.7 327.8 100% 21–147
Weekly Issue Summary 4 288.9 67.2 113.5 39.7 213.9 75% 0–174 ⚠️
Artifacts Usage Report 4 159.8 39.5 46.5 39.5 161.7 100% 0–103 ⚠️
Q 1 122.2 122.2 122.2 0.0 122.2 100% 0–122 ⚠️
Repository Tree Map Generator 4 79.6 17.4 26.1 17.4 60.8 75% 0–43 ⚠️
Smoke OpenCode 1 22.9 22.9 22.9 0.0 22.9 100% 0–23 ⚠️
Plan Command 1 19.6 19.6 19.6 0.0 19.6 100% 0–20 ⚠️

⚠️ = weekly and/or monthly Monte Carlo simulation flagged is_reliable: false by the
forecast engine (sample count too small, typically 1–4 runs) — treat these ranges as
low-confidence.

Data quality and accuracy

Unreliable Monte Carlo simulations (7 workflows)

The forecast engine itself flags is_reliable: false for both weekly and monthly
Monte Carlo projections on: GitHub MCP Remote Server Tools Report Generator,
Scout, Weekly Workflow Analysis, Weekly Issue Summary, Artifacts Usage
Report
, Q, Repository Tree Map Generator, Smoke OpenCode, and Plan
Command
. Root cause: each has only 1–7 sampled runs in the 30-day window, too few
for a statistically meaningful weekly/monthly distribution — several show a P10 of
0.0, meaning the lower bound collapses to "may not run at all this week," which is
plausible for low-frequency or manually-triggered workflows but should not be treated
as a hard floor for budgeting. Forecast impact: their contribution to the P10 total
is likely understated; their contribution to P90 carries wide uncertainty. No
follow-up evidence was pulled since sample counts are inherently small (these are
low-frequency workflows, not a collection failure) and the totals above already
reflect the raw Monte Carlo output as-is.

Zero sampled runs (25 workflows)

The following workflows returned sampled_runs: 0 in the 30-day window: Copilot
Setup Steps, .github/workflows/test-proxy, Notion Issue Summary, MCP Inspector
Agent, Poem Bot - A Creative Agentic Workflow, Copilot cloud agent, Rebuild the
documentation after making changes, Go Pattern Detector, Resource Summarizer Agent,
Doc Build - Deploy, Commit Changes Analyzer, Sentry Issue Analyzer, CodeQL,
Mergefest, CI Failure Doctor, Dependabot Updates, CI, Test, Test Claude, Test
Copilot CLI Engine, Test Copilot GitHub Integration, Basic Research Agent, Video
Analysis Agent, Format Lint Build and Commit, Dev Hawk.

Many of these are conventional CI/CD or setup workflows (CodeQL, CI, Test, Format
Lint Build and Commit, Dependabot Updates) that are expected to have zero or
near-zero AIC since they are not LLM-driven agentic runs, so their exclusion from the
forecast is expected and does not indicate a collection failure. A smaller subset
(Notion Issue Summary, MCP Inspector Agent, Poem Bot, Copilot cloud agent, Go
Pattern Detector, Resource Summarizer Agent, Sentry Issue Analyzer, Basic Research
Agent, Video Analysis Agent) are agentic-style workflows with zero recent runs,
suggesting they are currently disabled, unscheduled, or simply idle in this window —
this is consistent with intermittent/manual triggers rather than a data-collection
gap, since 25 other workflows across the same window did produce non-zero samples
from the same gh aw forecast collection pass. No rerun was performed because
sampled_runs was not zero for all workflows (only a subset), which is the
documented threshold for escalating to direct run/artifact inspection.

Low success-rate observations

Smoke Copilot shows the lowest success rate among active workflows (62%,
23/37 runs), consistent with it being a smoke-test workflow that is expected to
surface engine failures. No other active workflow falls below 75% success.
Success rate does not directly bias the AIC forecast (both successful and failed
runs incur cost), but a low success rate for a production (non-smoke) workflow
would warrant investigation; here it is limited to explicitly-named smoke/test
workflows.

Window and consistency checks

All 25 active workflows report history_days: 30, consistent with the forecast
period. sampled_runs matches observed_runs_per_period for every active
workflow, and p50_aic_per_run values fall within the observed min/max of each
workflow's run_samples[].aic, confirming internal consistency of the reported
percentiles. No AIC values were zero, negative, or missing for any run sample
among the 25 active workflows.

Assumptions

  • Forecast date: 2026-08-25 (as_of 2026-08-25T09:52:46Z)
  • History window: 30 days (2026-07-26 to 2026-08-25)
  • Only the 25 workflows with sampled_runs > 0 are included in aggregate
    totals and forecasts; the 25 workflows with zero samples are excluded from
    spend projections (see Data Quality section).
  • Monte Carlo projections use each workflow's own iteration count (10,000) and
    reliability flag as reported by gh aw forecast; flagged workflows are noted
    above but not excluded from totals.
  • AIC = Actions Infrastructure Cost as computed by the gh aw forecast command
    from actual workflow run billing/usage data, not an external estimate.

Reference

Full run: §32833886632

Generated by 📈 Daily Spending Forecast · copilot · auto · 38.8 AIC · ⌖ 12.3 AIC · ⊞ 11.3K ·

  • expires on Sep 1, 2026, 1:59 AM UTC-08:00