AI Spending Has Entered a New Phase
For many organizations, AI spending started with licenses.
A Microsoft 365 Copilot subscription had a clear monthly cost. Azure AI services were often confined to pilot projects. Budgeting felt manageable because leaders could associate AI adoption with a predictable number of users and licenses.
That has changed.
As Microsoft expands consumption-based pricing across Copilot agents, Azure AI Foundry, GitHub Copilot, and other AI workloads, organizations are finding that AI costs don’t stop at the license fee.
The conversation is shifting from:
- How many AI licenses do we own?
- How many employees are using AI?
- How quickly can we increase adoption?
To:
- What is driving our AI consumption?
- Which models and agents account for the spend?
- Who owns the budget?
- What business value are we receiving?
During Quisitive’s recent webinar, Tokenomics 101: What Every Executive Needs to Know Before the Next AI Invoice, Bryan Blaine, Vice President, Technology & Innovation at Quisitive, and Steven Balusek, Executive Vice President, Solutions Development, Advisory, and Presales at Quisitive, examined how this shift is affecting enterprise AI budgets.
As Bryan explained, “It’s really going from, ‘We bought a seat that gives us access to a product,’ to now it’s all consumption-based for a lot of this stuff.”
Steven captured the budgeting problem clearly: “You can buy the license. You can feel like you’ve budgeted for AI, but yet you’re still going to be exposed.”
For IT and finance leaders, that exposure can come from model usage, agent activity, business data retrieval, supporting Azure infrastructure, monitoring, and AI capabilities embedded in third-party software.
What Is AI Tokenomics?
AI tokenomics is the practice of understanding, governing, and optimizing how AI consumption translates into cost and business value.
Tokens are units that AI models use to process and generate information. Prompts, responses, document analysis, agent workflows, and automated processes all consume tokens. More complex tasks may require larger prompts, longer outputs, additional data, or repeated interactions with models and tools.
Token usage matters because it contributes to cost. It should not, however, become the primary measure of AI success.
What is token maxing?
Token maxing refers to treating AI consumption as an indicator of productivity or progress. During the first wave of enterprise AI adoption, organizations faced pressure to prove that employees were using AI. Some responded by tracking raw consumption. High usage appeared to signal successful adoption.
That encouraged the wrong behavior. When employees and teams are rewarded for consuming more AI resources, they naturally find more ways to use them. The organization may generate more tokens, prompts, and agent activity without improving the underlying business result.
As Bryan explained during the webinar, “What happens if you have the wrong incentive? People are going to learn how to game it.”
It’s similar to evaluating software developers based on the number of lines of code they write. More activity can create more cost and complexity without creating a better product. The better objective is value maxing.
“It’s really the idea of how do we go from, ‘How many tokens do we use?’ to ‘What do those tokens actually produce?’” Bryan said. “How do we put the governance in place to turn AI consumption into something that the business ultimately cares about?”
Why Microsoft AI Costs Are Becoming Harder to Predict
Many IT leaders assume they have budgeted for AI once they purchase the required licenses. Consumption-based services make the full cost harder to predict.
Microsoft licensing is only part of the cost
A Microsoft AI workload may involve several cost variables:
- The model selected for the task
- The number of input and output tokens
- The context or organizational data retrieved
- The tools and connectors called
- The length and frequency of agent activity
- The Azure infrastructure supporting the workload
- Logging, monitoring, and security requirements
- Capacity commitments or pay-as-you-go pricing
A user may have a Microsoft 365 Copilot license while also triggering consumption associated with agents, Azure services, or other metered capabilities. That creates a gap between license budgeting and total AI budgeting.
Why does this matter for enterprise IT leaders?
Finance teams need to forecast costs for systems whose usage can change quickly as employees create new agents, connect more data, or expand successful pilots.
The challenge resembles the earlier shift from on-premises infrastructure to cloud services. Cloud adoption required organizations to adjust from purchasing fixed infrastructure to managing variable operating expenses. AI is creating a similar financial management challenge in a shorter period.
“You can buy the license. You can feel like you’ve budgeted for AI, but yet you’re still going to be exposed,” Steven said during the webinar.
Without shared ownership and regular cost reviews, IT may see the technical usage while finance sees the invoice and business units see only the outcome they want to achieve. No single group has the full picture.
The Biggest Drivers of Microsoft AI Costs
An unexpectedly high AI bill rarely has one cause. Several cost drivers tend to work together.
1. Model selection
Model selection is often one of the first places to look for savings.
Organizations may default to the newest or most capable model because they assume it will produce the best result. That reasoning makes sense for tasks that require advanced analysis. It can create unnecessary expense for routine work.
Common tasks that may not require a frontier reasoning model include:
- Summarizing documents
- Classifying support requests
- Extracting fields from forms
- Drafting standard responses
- Categorizing records
- Routing workflow items
“Most of the work being done doesn’t require the latest frontier reasoning models,” Bryan said. “When you’re looking at classifications, summarizations, and other things, you don’t need to overpay for that work.”
Steven summarized the practical implication: “Knowing which model to use for which task is in itself a cost control.”
The lowest-cost model isn’t automatically the right model. Organizations should test whether a smaller model can meet defined quality, security, latency, and accuracy requirements before making a change.
2. Agent compute and execution
AI agents can complete multistep work across models, data sources, tools, and applications. Each additional action may add cost.
An agent may need to:
- Retrieve data from SharePoint or another enterprise system
- Send information to an AI model
- Call an external tool or connector
- Evaluate the result
- Perform another action based on that result
- Log the activity for review
The organization may pay for more than the model interaction. Supporting compute, containers, storage, retrieval, and monitoring can contribute to the total cost.
The same agent may also behave differently depending on the task. A short, controlled workflow can have a different cost profile from an agent that runs for an extended period or repeatedly calls tools.
3. Enterprise data and context
AI becomes more useful when it can work with the organization’s data. That context may include documents, emails, records, policies, product information, or data stored across Microsoft Fabric, SharePoint, and other platforms.
Retrieving and processing that information can increase consumption.
The goal isn’t to remove business context to save money. An agent without the right context may deliver a less useful or inaccurate answer. The goal is to provide the data required for the task without repeatedly retrieving irrelevant information.
4. Production monitoring and controls
A proof of concept can run with limited operational support. A production AI solution needs more.
Enterprise requirements may include:
- Logging and traceability
- Identity and access controls
- Security monitoring
- Data protection
- Quality testing
- Compliance reviews
- Performance monitoring
- Incident response
- Cost reporting
Organizations often budget for an AI tool or model while overlooking the services required to operate it safely in production.
5. Capacity planning
Organizations must also decide how they will purchase AI capacity.
Pay-as-you-go pricing can provide flexibility when workloads are new or unpredictable. Reserved capacity or pre-purchased commitments may make sense once usage is stable and understood. Committing too early can lock the organization into an inefficient usage pattern.
Bryan’s recommendation during the webinar was straightforward: understand and optimize the workload before making a larger commitment.
Steven summarized it as, “Optimize first and then commit.”
6. Shadow AI
Shadow AI refers to AI tools, agents, workflows, or services adopted outside established IT and procurement processes.
Business teams may purchase an AI tool or create an agent because they need to solve an immediate problem. Their intent may be reasonable, but the organization can still face:
- Untracked subscription and consumption costs
- Duplicated tools
- Unapproved data access
- Inconsistent identity controls
- Security and compliance exposure
- Limited visibility into agents and workflows
- Difficulty connecting spending to outcomes
Policy alone rarely resolves the issue.
“If leaders in the organization can’t see what’s being used, where things are at, and what licenses have been procured outside of procurement and IT, you don’t have that visibility,” Bryan explained.
Organizations need a clear intake process, approved tools, practical education, and a way to discover AI activity across the business.
Can Organizations Reduce AI Costs Without Reducing Quality?
Yes. Many AI workloads can be made less expensive through better model selection, tighter context, improved prompts, caching, workflow changes, and the removal of unused resources.
The sequence matters.
Organizations should measure current performance, define acceptable quality, test a proposed change, and then compare the results.
How should organizations approach AI cost reduction?
Start with changes that remove obvious waste before accepting reductions in capability.
A practical order is:
- Identify unused licenses, agents, endpoints, and services.
- Find workloads using expensive models for routine tasks.
- Reduce unnecessary context and repeated data retrieval.
- Review agent workflows for excessive calls or execution time.
- Test smaller models against defined quality standards.
- Confirm that actual usage supports any capacity commitment.
- Retire workloads that consume resources without producing a clear outcome.
As Bryan noted, “Every time you make a change, there can be an impact on model behavior.”
That means quality testing must come before broad deployment. A cheaper model creates no business benefit if it generates poor answers, increases manual review, or causes a downstream process to fail.
The objective is cost per acceptable outcome, not simply the lowest possible token price.
Six Questions Every IT Leader Should Be Able to Answer
During the webinar, Bryan and Steven challenged leaders to assess whether they could explain their organization’s AI spending and value to the CFO.
These six questions provide a useful starting point.
1. Who owns AI costs?
AI cost ownership should be shared across IT, finance, procurement, and the business.
IT can manage the technical environment and consumption data. Finance can establish budget discipline and forecasting. Procurement can review contracts, commitments, and vendors. Business owners must be accountable for whether a use case produces the expected result.
“If the answer is nobody, that’s pretty much a diagnostic in itself,” Steven said.
Ownership also fails when responsibility sits with IT alone. IT can report consumption, but it may not be able to determine whether the resulting business outcome justifies the expense.
2. Where is AI consumption occurring?
Leaders should be able to identify:
- Which business units generate AI spend
- Which models support each workload
- Which agents are active
- Which Microsoft and third-party platforms are involved
- Which use cases account for the largest costs
- Which owners are accountable for those use cases
A total invoice doesn’t provide enough information to make a sound decision. Costs need to be attributed to a workload, team, application, or business result.
3. How are we paying for AI?
Organizations should know whether each workload uses:
- Pay-as-you-go pricing
- Pre-purchased credits
- Reserved capacity
- License-based access
- Usage-based add-ons
- Third-party subscriptions
They should also compare actual consumption with committed capacity. Overcommitting can waste budget, while underestimating demand can create capacity or cost issues.
4. Do we have budgets, limits, and active monitoring?
A dashboard is useful, but it won’t control spending on its own.
Organizations should define:
- Budget owners
- Spending thresholds
- Alerts
- Escalation paths
- Actions required when usage exceeds expectations
- A regular review schedule
As Bryan explained, the difference is whether information is merely displayed or actively monitored with follow-up action.
5. Are we measuring business value?
Adoption metrics can indicate whether employees have started using AI. They don’t prove that the organization is receiving value.
Depending on the use case, relevant measures may include:
- Cost per completed transaction
- Time saved
- Reduction in manual effort
- Revenue influenced
- Operating cost avoided
- Error reduction
- Time required to reach the expected outcome
The measure should derive from the reason the organization approved the use case.
6. Do we know what AI is running in our environment?
This includes sanctioned platforms and tools adopted outside standard processes.
An AI inventory should capture:
- Microsoft 365 Copilot licenses
- Copilot agents
- Azure AI Foundry workloads
- GitHub Copilot usage
- Supporting Azure resources
- AI features within existing software
- Third-party AI products
- Business-created agents and automations
As Bryan put it, “You can’t govern what you can’t see, and you can’t optimize it either.”
How Should Organizations Build an AI Cost Governance Program?
A practical program can be organized into five phases.
Phase 1: Catalog the AI environment
Create an inventory of AI licenses, agents, models, applications, infrastructure, vendors, owners, and costs.
The inventory should connect technical usage to a business unit and use case whenever possible.
Phase 2: Define shared ownership
Establish responsibility across:
- IT
- Finance
- Procurement
- Security
- Business and product leaders
An AI investment council can provide a cross-functional forum for reviewing new use cases, financial commitments, security requirements, and expected business outcomes.
Phase 3: Build financial visibility
Create reporting that connects usage to owners and budgets.
Useful controls may include:
- Tags for workloads and resources
- Showback reporting that displays costs by business unit
- Chargeback models that assign costs to the consuming unit
- Budget alerts
- Consumption thresholds
- Recurring reviews of actual versus forecasted spending
Microsoft services may expose cost information in different administrative portals, so organizations may need to bring data together to create a full view.
Phase 4: Establish technical and financial guardrails
Guardrails should define how teams select, build, purchase, and operate AI services.
Policies may address:
- Approved models and services
- Model selection by workload type
- Security and data requirements
- Required testing
- Purchasing and procurement
- Pay-as-you-go versus reserved capacity
- Ownership and tagging
- Monitoring and reporting
Standards should give teams a practical path to build approved solutions rather than serving only as restrictions.
Phase 5: Connect spending to outcomes
Each AI use case should have a defined business objective and a method for measuring it.
Three useful measures are:
- Cost per inference, which connects total inference cost to the number of requests.
- Cost per token, which helps compare model economics for similar work.
- Time to business value, which measures how long it takes for the benefit of a use case to justify its cost.
These measures should support a broader question: Is the organization receiving enough value from this workload to continue, expand, redesign, or retire it?
Conclusion
Microsoft AI spending is becoming harder to predict because licenses represent only one part of the total cost. Model selection, agent execution, enterprise data retrieval, Azure infrastructure, monitoring, security, and governance can all affect the final amount.
Organizations should begin by identifying the AI services and agents running across the business. From there, they can assign ownership, attribute costs, set limits, test less expensive models, and connect spending to measurable outcomes.
Bryan Blaine and Steven Balusek returned to the same principle throughout Quisitive’s Tokenomics 101 webinar: activity alone is a poor measure of AI success.
“Token buying without outcome measurement always backfires,” Steven said.
The goal is to know what the organization is spending, why it is spending it, and whether each workload produces enough value to justify continued investment.
Continue the Conversation in Part 2
Understanding AI costs is only the first step.
This webinar focused on the economics of AI consumption, the factors driving Microsoft AI costs, and the governance questions every organization should be asking. The next challenge is turning that visibility into a repeatable operating model that finance, IT, procurement, and business leaders can align around.
In Part 2 of the series, From Spend to Strategy: Building an AI Cost Governance Framework That Holds Up in the Boardroom, Bryan Blaine and Steven Balusek will take the discussion further by exploring:
- How to establish ownership for AI spending across the organization
- Practical approaches for AI budgeting and forecasting
- Chargeback and showback models for AI workloads
- Governance structures that balance innovation with financial accountability
- KPIs and reporting frameworks executives can use to evaluate AI investments
- How organizations can align AI spending with measurable business outcomes
If your organization is already experimenting with Microsoft Copilot, Azure AI Foundry, agents, or other AI-powered solutions, Part 2 will provide a practical framework for governing AI investments as adoption scales.