On this page · 11 sections
- What actually expires, and what does not
- The credit-pool arithmetic
- Four defaults that turn a shortfall into an invoice
- The four-step forecast to run this week
- Where the credits actually go
- Levers worth pulling before 31 August
- India-specific considerations
- What we would not do
- FAQ
- How eCorpIT can help
- References
Summary. Two separate GitHub Copilot promotions expire at the end of August 2026, and both land on the same invoice. The larger one: existing Copilot Business and Copilot Enterprise customers have been receiving 3,000 and 7,000 included AI credits per user per month since 1 June 2026, and on 1 September those allowances return to the standard 1,900 and 3,900. That is a 36.7% cut for Business and a 44.3% cut for Enterprise, with no change in seat price. GitHub prices 1 AI credit at $0.01, so a 100-seat Business org loses 110,000 pooled credits, worth $1,100 a month. The smaller one: Claude Sonnet 5 has been billed at a promotional $2.00 per 1M input tokens and $10.00 per 1M output tokens "through August 31, 2026", and GitHub has not published what it costs on 1 September. Nothing about your usage has to change for your September bill to rise.
The awkward part is that the overage is opt-out, not opt-in. GitHub's documentation states that additional usage is enabled by default for organizations and enterprises, so a team that quietly exceeds the smaller pool keeps working and keeps billing. There is also no cheap-model safety net: GitHub removed the old fallback behaviour when it moved off premium requests. Below is the arithmetic, the four defaults that convert a credit shortfall into a real invoice, and a forecast you can run before the end of the month.
What actually expires, and what does not
It helps to separate the three things people conflate.
Seat prices are not changing. Mario Rodriguez, Chief Product Officer at GitHub, was explicit about this when the model was announced on 27 April 2026: "Base plan pricing is not changing. Copilot Pro remains $10/month, Pro+ remains $39/month, Business remains $19/user/month, and Enterprise remains $39/user/month." That still holds. What changes is how much model usage each of those seats buys you.
The included allowance is changing. GitHub's billing documentation describes a promotional band for "the first three months of usage-based billing (June 1 – September 1, 2026)", available to "existing Copilot Business and Copilot Enterprise customers", and states plainly: "After the promotional period, included usage returns to the standard amounts above."
One model rate is changing, or at least may be. The single footnote on GitHub's pricing reference reads: "Claude Sonnet 5 is available at the promotional pricing of $2.00 per 1M input tokens, $0.20 per 1M cached input tokens, $2.50 per 1M cache write tokens, and $10.00 per 1M output tokens through August 31, 2026." The page does not say what replaces it. Every other generally available Claude Sonnet on the same page, including Sonnet 4, 4.5 and 4.6, sits at $3.00 input and $15.00 output. Treat that as the planning worst case rather than a published fact, because GitHub has not confirmed it.
The credit-pool arithmetic
Credits pool at the billing entity level, not per seat. GitHub's own worked example: "an enterprise with 100 Copilot Business users gets a shared pool of 190,000 AI credits rather than 100 individual buckets." That pooling is helpful when usage is uneven, and it is exactly why the September change is easy to miss. Individual heavy users will not see a wall; the pool absorbs them until it does not.
| Line | June to August 2026 (promotional) | From 1 September 2026 (standard) | Change |
|---|---|---|---|
| Copilot Business, per user per month | 3,000 credits ($30) | 1,900 credits ($19) | -1,100 credits, -36.7% |
| Copilot Enterprise, per user per month | 7,000 credits ($70) | 3,900 credits ($39) | -3,100 credits, -44.3% |
| 100-seat Business pool | 300,000 credits ($3,000) | 190,000 credits ($1,900) | -110,000 credits, -$1,100/month |
| 250-seat Enterprise pool | 1,750,000 credits ($17,500) | 975,000 credits ($9,750) | -775,000 credits, -$7,750/month |
| 1,000-seat Business pool | 3,000,000 credits ($30,000) | 1,900,000 credits ($19,000) | -1,100,000 credits, -$11,000/month |
| Seat price | $19 Business / $39 Enterprise | $19 Business / $39 Enterprise | Unchanged |
The dollar figures follow directly from GitHub's fixed conversion, stated twice in the documentation: "1 AI credit = $0.01 USD, so a $10 USD budget covers 1,000 AI credits." Annualise the third row and a 100-seat Business org is looking at $13,200 of new metered exposure it did not have in August, assuming consumption stays flat.
Consumption rarely stays flat in September. August is a holiday month across Europe and North America, and a team that ran at 85% of a 3,000-credit allowance during a quiet month is running at 134% of a 1,900-credit allowance at the same absolute token volume. The percentage you should care about is not your August utilisation. It is your August absolute credit burn divided by the September pool.
Four defaults that turn a shortfall into an invoice
None of these are bugs. They are documented behaviours that happen to compound.
Additional usage is on by default. GitHub's documentation carries an explicit note: "Additional usage is enabled by default for organizations and enterprises. If you want to prevent any spending beyond your included AI credits, an administrator must explicitly disable the AI credits paid usage policy in your enterprise's or organization's AI Controls settings." If nobody in your organisation has touched that setting, September overage bills silently.
There is no fallback to a cheaper model. Under the old premium-request scheme, exhausting your allowance dropped you to a lower-cost model and you kept working. That is gone. GitHub said so at announcement: "Fallback experiences will no longer be available." The documentation repeats it: "There is no automatic fallback to lower-cost models when a budget is exhausted." So the choice at the end of the pool is billing or blocking, and the default is billing.
Credits do not roll over. "Included AI credits do not carry over between months. Unused credits are forfeited, and the pool resets to the full monthly amount at 00:00:00 UTC on the first day of each calendar month." A quiet August buys you nothing in September. The reset is a fixed calendar-month boundary and does not track your billing date, which matters if your finance team reconciles on a non-calendar cycle.
Output tokens dominate, and reasoning depth drives output. Across the frontier models on GitHub's own table, output is priced five to six times input: GPT-5.6 Terra at $2.00 in and $12.00 out, GPT-5.6 Sol at $5.00 and $30.00, Claude Opus 5 at $5.00 and $25.00. GitHub shipped a per-task reasoning selector for Copilot cloud agent on 3 August 2026 and stated the consequence directly: a higher level "can improve answers to complex problems, but it consumes more tokens, and therefore more credits." GitHub has published no multiplier for it. Neither the billing reference nor the changelog gives a number, so the only way to know your delta is to measure it.
The four-step forecast to run this week
You have until 31 August. This is a two-hour exercise, not a project.
Step 1: pull August credit burn in absolute terms. Since 20 July 2026, Copilot Business and Enterprise users can see credits consumed per billing cycle on their Copilot usage page even without an individual budget; before that, the page only showed a percentage of a budget most people did not have. At the admin level, export the AI credit data from the AI usage page in billing settings, where GitHub says you can "group, filter, and export" it. Do not read the percentage. Read the raw credit count.
Step 2: divide by the September pool, not the August one.
seats = 100 # Copilot Business
august_credits_used = 240000 # absolute, from the AI usage export
september_pool = seats * 1900 # = 190,000 credits
projected_overage = august_credits_used - september_pool
= 240000 - 190000
= 50,000 credits
projected_overage_usd = 50000 * 0.01 # 1 credit = $0.01
= $500 per month
The team in that example was comfortably inside its 300,000-credit promotional pool in August, at 80% utilisation, and is $500 a month over from 1 September without changing a thing.
Step 3: add a September growth assumption and a model-mix assumption. If a meaningful share of your output tokens ran through Claude Sonnet 5 at the promotional $10.00 per 1M, model the same volume at $15.00 as a ceiling case. That is a 50% increase on the output side of that slice, and on a mix where Sonnet 5 carried, say, 40% of output tokens it moves the whole bill roughly 20% on that component alone. Label it clearly as a scenario in whatever you hand finance, because GitHub has not published the post-promotional rate.
Step 4: decide the policy before the pool empties, not after. Three settings, in the AI Controls and budgets screens: whether AI credits paid usage stays enabled, what the enterprise spending limit is, and whether user-level budgets exist. GitHub notes that "a $0 USD user-level budget blocks the user immediately", and that "a user can also be blocked by an enterprise spending limit before they reach their individual user-level budget, if the spending limit runs out first." Blocking a senior engineer mid-sprint on 22 September is a worse outcome than a $500 overage, so pick deliberately.
One tooling change to note while you do this: GitHub retired the Copilot Billing Preview app on 4 August 2026. If your September forecast was going to come from that app, it will not. The replacement path is the AI usage page, usage reports and the billing API.
Where the credits actually go
The instinct is to blame chat. The billing reference disagrees. "Code completions and next edit suggestions are not billed in AI credits. They remain unlimited for all paid Copilot plans." The billed surface is agentic: "Copilot Chat, Copilot CLI, Copilot cloud agent, Copilot Spaces, Spark, and third-party coding agents."
| Surface | Consumes AI credits | Also consumes something else | Cost control available |
|---|---|---|---|
| Code completions | No | No | Not applicable, unlimited on paid plans |
| Next edit suggestions | No | No | Not applicable, unlimited on paid plans |
| Copilot Chat | Yes | No | Model choice, prompt discipline |
| Copilot cloud agent | Yes | No | Model choice plus per-task reasoning level |
| Copilot code review | Yes | GitHub Actions minutes | Lite or Balanced effort level, org default |
| Copilot CLI, Spaces, Spark | Yes | No | Model choice |
| Third-party coding agents | Yes | No | Agent-level tracking via usage metrics API |
Copilot code review is the line item most often missed, for two reasons. It bills twice: token consumption in AI credits plus the Actions minutes that run the review, billed at the same per-minute rates as any other Actions workflow. And GitHub states that for code review "the model is selected automatically and is not disclosed, so per-token costs may vary between reviews." You cannot pin its unit cost from the rate card.
What you can do, since 7 August 2026, is control its depth. Lite and Balanced effort levels went generally available that day, replacing the preview names Low and Medium, with Balanced routing to "a higher-reasoning model". Organization admins can set a default that repositories inherit, and each review is labelled with the effort level that ran. Setting the org default to Lite and reserving Balanced for security-sensitive or cross-service pull requests is the cheapest structural saving available before September, because it does not require anyone to change how they work.
The other visibility gain landed the same day. The Copilot usage metrics API added a totals_by_3rd_party_agent array, one entry per recognised agent app, in the enterprise, organization, enterprise-user and organization-user 1-day and 28-day reports. If your teams run Claude or Codex agents alongside Copilot cloud agent, that spend was previously one undifferentiated bucket. GitHub's own framing is that it lets you "ground rollout and licensing decisions in real usage rather than assumption." Two cautions from the same changelog: group on agent_id rather than agent_name, because display names change, and do not sum the nested user_initiated_interaction_count with the top-level field of the same name.
Levers worth pulling before 31 August
Model routing is the biggest one, and the spread on GitHub's own rate card is wide enough that routing beats rationing.
| Model | Input per 1M tokens | Output per 1M tokens | Category |
|---|---|---|---|
| GPT-5.6 Luna | $0.20 | $1.20 | Lightweight |
| GPT-5 mini | $0.25 | $2.00 | Lightweight |
| Raptor mini | $0.25 | $2.00 | Versatile |
| MAI-Code-1-Flash | $0.75 | $4.50 | Lightweight |
| Kimi K2.7 Code | $0.95 | $4.00 | Versatile |
| GPT-5.6 Terra | $2.00 | $12.00 | Versatile |
| Claude Sonnet 5 | $2.00 (promotional) | $10.00 (promotional) | Versatile |
| Kimi K3 | $3.00 | $15.00 | Powerful |
| GPT-5.6 Sol | $5.00 | $30.00 | Powerful |
| Claude Opus 5 | $5.00 | $25.00 | Powerful |
| Claude Fable 5 | $10.00 | $50.00 | Powerful |
Read the output column. GPT-5.6 Luna produces output at $1.20 per 1M against Claude Fable 5 at $50.00, a factor of 41. Almost no organisation needs its most expensive model for dependency bumps, changelog drafts, test scaffolding or docstring passes. GitHub added Kimi K3 to Copilot on 6 August 2026 as an open-weight option billed at $3.00 and $15.00, which sits below Sol and Fable 5 for agentic work. Deciding which classes of task get which tier, then enforcing it through model policy rather than a wiki page, is where the money is. We have written separately about Copilot AI credit pools and cost control and about reasoning level configuration for the Copilot cloud agent; both mechanisms matter more from 1 September than they did in June.
Three more levers, in rough order of return:
Cache deliberately. Cached input is priced at 10% of input on GitHub's OpenAI and Anthropic rows: GPT-5.6 Terra at $0.20 against $2.00, Claude Sonnet 5 at $0.20 against $2.00. Grok 4.5 is the exception at 25% of input, $0.50 against $2.00. Anthropic and GPT-5.6 models also carry a cache write cost, so a cache that is written constantly and read rarely costs money rather than saving it.
Watch the long-context thresholds. GPT-5.4, GPT-5.5, GPT-5.6 Sol and GPT-5.6 Terra switch to a long-context tier above 272K input tokens; GPT-5.6 Luna, Gemini 3.1 Pro and Grok 4.5 switch above 200K. Input doubles in every case. A repository-wide context dump that crosses the threshold silently doubles the input side of that request.
Instrument before you legislate. The Copilot impact dashboard added a "Potential return on investment" section on 7 August 2026 showing cost per developer per month derived from actual credit consumption, that cost as a share of payroll, and pull requests per developer per month, split between chat-and-completion users and agent-first developers. GitHub is careful about the caveat: "Cost figures are estimates based on AI credit consumption, and the salary selector is a modeling input rather than actual payroll data. Treat these metrics as directional." Directional is still better than a spreadsheet of guesses, and the same release widened cohort counts to include every user active across the full 28-day window rather than only the final day, so expect your phase counts to look higher than they did in July.
India-specific considerations
For Indian engineering organisations and global capability centres, the September step has a different shape than it does for a 40-person startup in Berlin.
Seat counts are larger, so the pooled loss scales linearly and lands as a single number. A 1,200-seat Copilot Business estate loses 1,320,000 pooled credits a month, or $13,200, without a single new user or a single behaviour change. That is a budget line, not a rounding error, and it arrives in a quarter where most Indian finance teams have already locked FY27 tooling budgets.
Second, the exposure is dollar-denominated while the budget is usually rupee-denominated, so the metered overage carries currency risk that the flat seat fee did not. A fixed $19 per seat is easy to hedge; a variable overage that depends on how many agent runs your teams start in a given sprint is not. Setting an enterprise spending limit converts that open-ended exposure back into a capped, forecastable number, which is usually worth more to an Indian CFO than the marginal productivity of an unbounded pool.
Third, offshore delivery teams tend to be agent-heavy, because agentic runs map well onto well-specified tickets. That is precisely the usage pattern the credit model prices highest. Teams whose work is mostly completions and inline chat are barely affected, since completions are not billed at all. Teams running Copilot cloud agent across large repositories are the ones that will feel 1 September.
Personal data does not enter the credit calculation, but it does enter the decision about which models and third-party agents you enable, and Indian teams are making that decision under the Digital Personal Data Protection Act 2023. GitHub shipped the enforcement point on 6 August 2026: allowedMcpServers and deniedMcpServers keys in copilot/managed-settings.json let enterprise owners control which Model Context Protocol servers Copilot clients may run, matched by remote URL, local command or name. GitHub notes that these policies "fail closed, meaning a malformed or unverifiable configuration is blocked rather than allowed", that a server must pass every policy layer, and that serverName matching "is only supplied as a convenience, not a security control, since users can rename servers". Enforcement currently covers the Copilot app, Copilot CLI and VS Code. That makes the same file a cost surface and a data-governance surface; we have covered enterprise managed settings and Copilot model policy separately.
What we would not do
We would not disable paid additional usage across the board on 31 August as a reflex. It converts a cost problem into an availability problem, and the first time a staff engineer is blocked mid-incident, the policy gets reversed loudly and permanently. Set an enterprise spending limit with headroom instead, and put user-level budgets only on cohorts whose consumption you have actually measured.
We would not rewrite the model policy in the same week the pool shrinks. Change one variable at a time or you will not know which change moved the number. Effort levels for code review are the cleanest first move, because the org default is one setting and the effect is visible in the next review's label.
We would not build the September forecast on percentage-of-budget figures. Most users have no individual budget, which is why GitHub changed the usage page in July to show absolute credits. Percentages of a denominator that is about to change by 37% will mislead you in exactly the direction you cannot afford.
The real cost here is not the model rate. It is running an agentic development workflow on a pool you have never measured in absolute terms.
FAQ
How eCorpIT can help
eCorpIT builds and runs engineering cost-control practice for teams that adopted AI coding tools faster than they adopted the reporting around them. We model your actual credit burn against the September pool, set model routing and effort-level policy through enterprise managed settings rather than guidance documents, and wire the usage metrics API into whatever your finance team already reads. eCorpIT is CMMI Level 5, MSME certified and ISO 27001:2022 certified, and we work with Indian and global engineering organisations from Gurugram. If your September Copilot forecast is currently a percentage rather than a number, talk to our engineering team and we will help you build the number. For broader context, see our guide to cloud FinOps for Indian teams and our analysis of hybrid LLM routing for API spend.
References
- Models and pricing for GitHub Copilot - GitHub Docs, per-model token rates, the Claude Sonnet 5 promotional footnote, long-context thresholds and the AI credit conversion.
- Usage-based billing for organizations and enterprises - GitHub Docs, standard and promotional credit allowances, pooling, expiry, budgets and default overage behaviour.
- GitHub Copilot is moving to usage-based billing - Mario Rodriguez, GitHub Blog, 27 April 2026, seat prices, promotional amounts and the removal of fallback experiences.
- Copilot users can now see AI credits used per billing cycle - GitHub Changelog, 20 July 2026.
- Retiring the Copilot Billing Preview app - GitHub Changelog, 4 August 2026.
- Copilot code review effort levels are generally available - GitHub Changelog, 7 August 2026, Lite and Balanced levels and organization defaults.
- Copilot usage metrics API adds agent app activity - GitHub Changelog, 7 August 2026, the
totals_by_3rd_party_agentarray and its field caveats.
- Copilot impact dashboard adds a return on investment section - GitHub Changelog, 7 August 2026, cost per developer and cohort counting changes.
- Customize the reasoning level for Copilot cloud agent - GitHub Changelog, 3 August 2026, reasoning level and its credit consequence.
- Kimi K3 is now available in GitHub Copilot - GitHub Changelog, 6 August 2026.
- Budgets for usage-based billing - GitHub Docs, budget levels for enterprises, cost centers and users.
- Copilot usage metrics REST API - GitHub Docs, the reporting endpoints behind the figures above.
- MCP allowlists in enterprise managed settings - GitHub Changelog, 6 August 2026, the
allowedMcpServersanddeniedMcpServerskeys and their fail-closed behaviour.
Last updated: 9 August 2026.