Daily A/B Testing Advisor

Durable reports produced by .github/workflows/ab-testing-advisor.md in github/gh-aw.

standalone

.github/workflows/ab-testing-advisor.md

View authored workflow

Reports

0 Open 2 Resolved

[ab-advisor] Experiment campaign for ci-coach: A/B test prompt_style

🧪 Experiment Campaign: ci-coach Workflow file: .github/workflows/ci-coach.md Selected dimension: promptstyle Triggered by: ab-testing-advisor on 2026-05-15 Background The ci-coach workflow is a daily CI optimization coach that analyzes workflow runs and proposes efficiency improvements via pull request. Its current prompt is highly prescriptive — it defines six numbered phases with detailed time budgets, extensive safety checklists, and repeated prohibitions. This verbosity may be consuming more tokens than necessary without improving output quality, making promptstyle the highest-impact dimension to experiment on. Hypothesis H0 (null): Switching from the current detailed, phase-structur...

closed review issue

[ab-advisor] Experiment campaign for agent-performance-analyzer: A/B test caveman_mode

🧪 Experiment Campaign: agent-performance-analyzer Workflow file: .github/workflows/agent-performance-analyzer.md Selected dimension: cavemanmode Triggered by: ab-testing-advisor on 2026-05-19 Background The Agent Performance Analyzer is a meta-orchestrator that analyzes AI agent performance across the repository. It currently uses a comprehensive 648-line prompt with detailed instructions across 5 phases. This experiment tests whether extreme prompt compression (the "caveman" principle: "why use many token when few do trick") preserves output quality, allowing us to identify prompt verbosity waste and reduce token consumption without sacrificing effectiveness. Hypothesis H0 (Null): Promp...

closed review issue