Two things changed for Microsoft Copilot over the summer, and between them they change what a business running on Microsoft 365 should be commissioning. Neither is about what the technology can do. Both are about who pays for what.

The questions that follow from those changes are the ones we get asked in the room and rarely see answered in public: what does an agent actually cost per interaction, how far will the platform carry a real process, and where does it stop? This piece answers all three, with the numbers.

Copilot is now part of the subscription

Since 1 July 2026, Business Standard with Copilot and Business Premium with Copilot have been permanent plans for businesses of up to 300 seats. Copilot was previously a separate line item that had to justify itself on its own merits; it now arrives with the subscription.

The practical effect is a change of question. It is no longer whether to buy Copilot, but what you should be doing with something you already pay for — and in most of the businesses we speak to, nobody has yet been given the time to work that out.

Building agents is now metered

Copilot Studio operates three harnesses, and they are billed differently. The distinction is the most expensive thing on this page to misunderstand.

The Standard and Copilot Chat harnesses suit well-defined, rule-shaped work. A Microsoft 365 Copilot licence covers your licensed users on both.

The GitHub Copilot harness handles reasoning-heavy, multi-step work — the genuinely agentic end of the platform. Since 1 September 2026 it consumes Copilot Credits for every part of that work: authoring the agent, previewing it, testing it, generating evaluations, and running it in production. This applies whether or not you hold a Microsoft 365 Copilot licence. Agents and workflows built before 3 August moved onto the new arrangement on 1 September.

Read that once more, because it inverts a normal procurement assumption. On the capable harness, the meter runs while you are still establishing whether the idea works at all.

What a credit actually buys

Billing on the Standard harness moved from a flat count of messages to feature-based rates, and the rates stack. One user turn can draw from several meters at once, which is why an identical question costs different amounts depending on how the answer was built.

What happens Credits
Classic, scripted answer 1 per response
Generative answer 2 per response
Agent action — a trigger, a reasoning step, a topic transition, computer use 5 per action
Tenant-graph grounding — searching your Microsoft 365 content 10 per message
Agent flow actions 13 per 100 actions
Premium AI tools — what a reasoning model draws 100 per 10 responses

A credit costs $0.01 on a pay-as-you-go meter, or about $0.008 prepaid in packs of 25,000 for $200 a month.

Read the last row carefully, because it is the one most often quoted wrongly. Premium AI tools bill at 100 credits per ten responses, which is ten credits per response, not a hundred. Microsoft’s own worked example makes it explicit: a generative answer using a reasoning model costs two credits for the answer plus ten for the premium tool, prorated from that 100-per-10 rate.

So the stacking works out like this. An agent that searches your tenant and writes an answer costs 12 credits per response — 10 for the grounding, 2 for the generation. Switch it to a reasoning model and the same response costs 22. That is a real increase, and it is the kind of thing a design decision settles early without anyone pricing it, but it is not the tenfold jump that gets repeated in licensing write-ups.

The bigger saving is easy to miss. For employee-facing agents, classic answers, generative answers and tenant-graph grounding are included at zero cost when the user holds a Microsoft 365 Copilot licence. Point the same agent at customers or partners and all of it bills normally. That distinction moves more money than any per-feature rate, and it is a decision about audience rather than architecture.

Two more worth knowing before anyone builds. Autonomous work bills per action rather than per wake: Microsoft’s example of an order-processing agent making four action calls per trigger comes to 20 credits each time it fires — so an agent that wakes every fifteen minutes bills through the night whether or not it finds anything. And when consumption reaches 125% of prepaid capacity, custom agents are disabled rather than throttled; conversations already running finish, and the tenant admin gets an email.

None of this is the harness that changed on 1 September. The rates above are the Standard harness rate card. The GitHub Copilot harness bills on tokens and complexity — model choice, runtime, context size, external tool use — published as light, medium and heavy ranges rather than per-feature prices. There is no equivalent table to hand you, which is precisely the problem: on the harness that does the agentic work, the only honest forecast comes from a metered pilot.

Where the platform is strong

Three capabilities do most of the work in making straightforward cases genuinely straightforward.

Dataverse grounding. Agents can be grounded in Dataverse, which acts as the operational database beneath them: your own tables, carrying the access controls they already have, queried in place rather than copied somewhere else first.

Model Context Protocol servers. Agents pick up tools from external systems through MCP servers, which hand over each tool with its inputs and outputs already described and keep them in step as the underlying system changes.

Workflows as tools. An agent can call a workflow rather than reasoning its way through a subprocess — running a tested, deterministic flow and using the result to continue. Microsoft’s framing is that this lets the high-frequency and high-stakes parts of a process run with the consistency the business requires. That is the right instinct, and it answers the most common objection to agents in operational settings: the part that has to be exact does not have to live inside the model.

Taken together, a single well-scoped agent that knows your data, calls a handful of reliable flows, and reaches two or three external systems is quick to stand up and genuinely useful. For a business already inside Microsoft 365, that covers a substantial amount of day-to-day work for very little effort. It is the strongest case for the platform, and it is a strong one.

Where the platform stops

The limits are documented, specific, and worth knowing before rather than after.

Tools per agent. Generative orchestration supports up to 128 tools, but the guidance is to stay under roughly 25 to 30. Beyond that the agent begins selecting the wrong tool, and the failure mode is quiet: a plausible answer drawn from the wrong source.

No nesting. An orchestrator can connect to subagents, but those subagents cannot have connected agents of their own. Attempting it produces an “agent chaining detected” error, and the documented resolution is to flatten the hierarchy.

Subagents cannot address each other. The shape is a star, not a tree. Everything moves through the orchestrator on a declared contract: it passes Inputs, the subagent works within its own context, and it returns Outputs. The subagent’s raw data reads, reasoning and intermediate tool calls never reach the parent — only the declared result does. That keeps the orchestrator fast and insulates it from a subagent’s mistakes, which is sound design rather than an oversight.

The cost of that isolation is what to plan around. If one subagent needs something another has discovered, the orchestrator must carry it, and only if you declared it as an Output when you designed the schema. Every piece of cross-agent knowledge has to be anticipated in advance. What the platform offers is delegation — a manager handing out briefs and collecting reports from staff who have never met — rather than collaboration. That is excellent for routing a request to the right specialist and returning an answer. It does not support work whose value comes from agents building on each other’s findings.

Handoffs at scale. Two agents passing work between them behave as the demonstrations suggest. Reported experience at five or six is different: context does not carry cleanly, and a supervising agent selects the wrong subordinate often enough to matter. Citations may also fail to survive the return trip to the calling agent.

Knowledge limits. Without a Microsoft 365 Copilot licence in the same tenant, SharePoint files must be under 7 MB to be used for answers; with one, and with tenant graph grounding enabled, that rises to 200 MB. Files uploaded directly can be considerably larger. More consequentially, a SharePoint search feeds only the top three results into the answer — so an agent pointed at a large, undifferentiated document library will answer confidently from the wrong three files. Indexing a large library can also take days.

Autonomous agents run as their maker. Event triggers authenticate with the credentials of whoever built the agent, which means anyone using it operates with that person’s permissions. It is the most common route to someone retrieving data they could not have opened themselves. If you take a single governance point from this article, take that one.

What this means for build versus buy

Our June piece argued that the decision is between two costs rather than between cost and no cost. That still holds — with a third option now worth pricing alongside the other two.

If the work is rule-shaped, lives inside Microsoft, and your team is licensed, use what you already own. Commissioning a build for that is difficult to justify.

If the work needs the agentic harness, you are making an arithmetic decision rather than a platform one — and the arithmetic has to come from a metered pilot, because that harness publishes ranges rather than rates. Run it small, read the actual consumption for a month, then set the annualised figure against a fixed-price piece of software that does the same job and never sends a variable bill. Which one wins depends on volume and on design. Only one of those is knowable before you start, which is itself worth planning around.

Some work sits outside the agent altogether. Anything that must be exactly right every time — dangerous-goods documentation, invoice calculations, anything an auditor will read — wants deterministic code. That does not rule out an agent. It means the exact part belongs in a workflow the agent calls, not in its reasoning.

How we’re approaching it

For clients already standardised on Microsoft, we now implement inside the Microsoft stack rather than defaulting to something custom beside it. On several recent engagements, the honest answer to “what should we build here?” was “less than you expect, and inside what you are already paying for.”

The rest of that answer is knowing where the metered harness stops being the cheaper option, and which of these limits your particular process will run into first. Both are arithmetic, and both are worth doing before the pilot rather than after it.