Token Economics Will Determine Which AI Strategies Actually Scale
- Manan Sharma

- 4 days ago
- 10 min read

AI is getting cheaper. Your AI bill is not necessarily getting smaller, and that gap is the whole game.
The strategies that scale won't be the ones running the cheapest model. They'll be the ones that treat every token like a unit of labour: allocated to the work that earns it back, and cut from the work that doesn't. That discipline, not model access, is what will separate firms that profit from AI from firms that just spend on it.
Key takeaway
Falling per-token prices do not lower total AI spend. Usage grows faster than price falls.
Cost per token is the wrong metric. Cost per accepted deliverable is the right one.
Consulting work is context-heavy, so poorly designed retrieval and prompting is what actually drives cost, not the model itself.
Firms that redesign both the workflow and the pricing model capture the margin. Firms that just automate existing tasks under time-and-materials pricing hand the savings to the client.
In founder-led firms, this is also a succession issue: the operating discipline you build around AI now determines whether the business can run, and eventually transfer, without the founder holding every process in their head.
For many consulting firms, the cost of AI adoption is becoming harder, not easier, to understand. The reason is token economics, and getting it right is quickly becoming a core token economics AI strategy question for firm leaders, not just an IT line item.
What tokens are, and why they add up
Tokens are the units large language models use to process and generate information. A token may represent a word, part of a word, a number, a symbol, or another piece of data.
When an AI system reads a client document, interprets a prompt, retrieves information, reasons through a problem, calls another tool, or generates a response, it consumes computational resources that are commonly measured and priced through tokens.
At the level of one interaction, that cost looks trivial. At the level of an entire firm, it becomes an operating, margin, and service-design decision.
The central question is not how much a token costs. It's whether the firm is producing enough business value for every token it consumes.
The real cost equation
Most commercial AI providers price input tokens, output tokens, and cached tokens differently, and some add charges for tools, storage, retrieval, or specialized processing. Output tokens are frequently priced higher than input tokens, while caching and batch processing can lower the cost of repeatable workloads (Anthropic, 2026; Google, 2026; OpenAI, 2026).
So the cost of an AI workflow isn't just a function of headcount. It's shaped by:
The model selected for the task
The volume of information fed to the model
The length and complexity of the expected response
How many times the system retries or revises its work
The number of tools and data sources involved
The amount of human verification and rework required
A simple summarization prompt uses relatively few tokens. An AI agent that reviews hundreds of documents, searches internal systems, compares evidence, drafts recommendations, checks its own work, and formats a final report can consume substantially more.
The real model is:
Total AI delivery cost = model consumption + workflow infrastructure + information retrieval + evaluation + human review + rework.
Token price is one line item in that equation, not the whole thing.
Falling prices don't shrink the problem
Inference costs have fallen fast. Stanford's 2025 AI Index reported that the cost of querying a model performing at roughly GPT-3.5 level on the MMLU benchmark dropped from $20 per million tokens in November 2022 to $0.07 by October 2024, a decline of more than 280 times (Stanford Institute for Human-Centered Artificial Intelligence, 2025).
That decline is real and it's driving adoption. It makes previously uneconomical use cases viable and lets firms experiment with larger volumes of information.
But cheaper tokens don't mean smaller total spend. As models get cheaper, firms use them more often, on bigger documents, in more workflows, through more autonomous agents.
One employee asking AI to tighten up an email costs almost nothing. Hundreds of consultants running AI through research, analysis, delivery, business development, and knowledge management is a different cost profile entirely.
It's the same pattern as cloud computing. Low entry cost invites experimentation. Ungoverned usage turns into an operating expense nobody can fully explain a year later.
Why consulting firms are especially exposed
Consulting sells knowledge, judgment, analysis, communication, and access to expertise. That's a strong fit for generative AI, and it's also context-heavy work.
A credible consulting output may require the AI system to understand:
The client's industry and competitive environment
Internal reports, interviews, financial information, and operating data
Previous recommendations and decisions
The firm's own methodologies and intellectual property
Regulatory, market, and technical evidence
The tone and expectations of senior stakeholders
Every added source is more context to process. A poorly designed system resends the same documents, instructions, examples, and client history to the model on every call. A better-designed system retrieves only what the current task needs.
That distinction is the whole ballgame. Two firms can use the identical model for the identical use case and land on very different economics.
One firm uploads the entire document library every time someone asks a question. Another uses structured retrieval to pull the few relevant passages. One runs a premium reasoning model for routine classification. Another routes routine work to a smaller model and saves the expensive model for decisions that justify it.
The competitive edge won't come from which model a firm can access. Most firms can buy the same technology. It will come from the quality of the operating architecture around it.
It's the same gap we've pointed to before: most firms adopt AI at the task level and never touch the operating model underneath it, which is exactly why the underlying bottlenecks don't move (see AI-Driven Operating Model: Why AI Isn't Improving How Your Firm Operates).
Stop measuring cost per token
A narrow focus on cost per million tokens leads to bad decisions.
The cheapest model isn't necessarily the most economical one. A low-cost model that produces an incomplete analysis, invents evidence, or needs several rounds of revision can end up costing more, total, than a premium model that gets it right the first time.
The better metric for a consulting firm is:
Cost per accepted deliverable.
An accepted deliverable meets the firm's bar for accuracy, relevance, evidence, structure, and client readiness. The cost of the first draft isn't the whole story. You have to count consultant review, fact-checking, revisions, formatting, and the cost of any errors that get through.
Research on more than 750 Boston Consulting Group consultants shows why this distinction matters. On tasks inside the model's capability, consultants using GPT-4 finished work more than 25% faster and produced results rated more than 40% higher in quality, and they completed more tasks. On a task outside the model's effective range, consultants using AI were more likely to land on a wrong answer than consultants working without it (Dell'Acqua et al., 2023).
AI isn't a uniform productivity bump. It's a huge win in some parts of a consulting workflow and added risk in others. The right question isn't "how much AI can we use." It's "where does AI reliably lower cost or raise quality, and where does it just add exposure."
Token economics will hit consulting margins directly
How this plays out on the P&L depends entirely on the firm's commercial model.
Under a fixed-fee engagement, cutting the hours it takes to deliver the work improves margin. AI-supported research, document analysis, proposal development, and report production can let a team deliver the same or more value in fewer delivery hours.
Under time-and-materials, the same productivity gain cuts billable hours. Unless pricing changes with it, the firm gets more efficient and poorer at the same time.
That's the tension. A firm can't build its AI strategy around making existing tasks faster while still pricing purely on human time consumed.
Token economics will push the industry toward pricing based on outcomes, IP, access, subscriptions, managed services, and measurable value. Firms that redesign delivery and pricing together will keep more of the upside. Firms that automate the work but keep selling hours will hand most of that value straight to the client.
Canadian context: this is also a succession issue
This isn't only a margin question for Canadian founder-led firms. CFIB estimates that more than $2 trillion in business assets will change hands across Canada over the next decade, with 76% of small business owners planning to exit, yet only 9% have a formal succession plan in place. A firm that has built disciplined, well-documented AI workflows, rather than a pile of ad hoc prompts living in one founder's head, is easier to value, easier to transfer, and easier to run without that founder in the room. Token economics and succession readiness turn out to be the same exercise: building a system that works independent of any one person.
How to actually navigate this
1. Establish the economics of each use case
Every priority use case needs a real baseline: current human time, cost, quality, cycle time, and error rate, measured before AI touches it. The comparison isn't AI output versus zero cost. It's full cost of the current process versus full cost of the redesigned one.
2. Match model capability to task value
Not everything needs a frontier model. Classification, extraction, formatting, routing, and basic summarization can run on smaller, cheaper models. Complex synthesis, ambiguous problem-solving, and high-stakes client recommendations justify the expensive one.
Model routing should be an operating decision, not an individual employee's habit.
3. Engineer the context, not just the prompt
Prompt engineering matters. Context engineering matters more at scale. Firms need to define what information the model actually needs, when it needs it, and how it gets retrieved.
Reusable instructions, structured templates, prompt caching, retrieval systems, and curated knowledge libraries cut unnecessary processing and improve consistency at the same time.
4. Put limits on agentic workflows
Agents can plan and execute multi-step work, but every extra step adds token consumption, latency, and a new chance to go wrong. Agents need defined stopping conditions, tool permissions, spending thresholds, and escalation rules.
Building these agents has never been easier. Tools like OpenAI's AgentKit now let non-technical teams stand one up in hours instead of months (see Build Custom AI Assistants Without Developers). That ease of creation is exactly why the guardrails matter more, not less. An agent that's simple to build is just as simple to leave unsupervised.
Without those guardrails, an agent will keep searching, reasoning, and revising without the output getting proportionately better.
5. Measure quality-adjusted economics
Track more than adoption rates and total spend. Useful measures:
Cost per completed workflow
Cost per accepted deliverable
Percentage of outputs requiring material revision
Consultant time saved after verification
Cycle-time reduction
Margin improvement
Client outcome or revenue contribution
These connect AI consumption to actual operating performance.
6. Redesign the workflow and the pricing model together
McKinsey's research points to workflow redesign as one of the strongest drivers of measurable financial impact from generative AI, yet most organizations still bolt AI onto existing processes instead of rebuilding how the work gets done (McKinsey & Company, 2025).
Resist using AI to just speed up individual tasks. The bigger opportunity is redesigning the engagement itself: how evidence gets gathered, how insight gets developed, how clients participate, how quality gets verified, and how the work gets priced.
This is the same trap we've flagged in a fragmented growth system: AI bolted onto broken handoffs between sales, delivery, and marketing doesn't fix the system, it just processes the same chaos faster (see AI Growth Strategy: AI Will Not Save a Fragmented Growth System).
Token economics is the same lesson applied to cost instead of growth.
From cost control to value architecture
Token economics shouldn't become an excuse to restrict experimentation or force everyone onto the cheapest model. That approach can lower the visible technology line while quietly increasing rework, inconsistency, and risk.
The goal isn't minimizing token consumption. It's building a system where token consumption is purposeful, measurable, and tied to value.
For a consulting firm, that means treating tokens like professional labour, data, or capital: an input to be allocated toward the activities where it produces the greatest return.
The firms that master token economics won't be the ones with the smallest AI bill. They'll be the ones that can show every dollar of AI spend drives faster delivery, stronger analysis, reusable IP, better margins, or better client outcomes.
That's the difference between experimenting with AI and building an AI-enabled consulting business.
Frequently asked questions
Will AI replace founders?
No. AI replaces specific bottlenecks that founders currently hold onto by default, not the founder's judgment, relationships, or decision-making. The tasks most exposed are the repeatable ones a founder still does personally: first-draft proposals, research synthesis, document review, and status reporting. ALTA's view is that AI's real job in a founder-led firm is to remove the founder as the single point of failure in day-to-day delivery, freeing them for the judgment calls only they can make.
How does AI help founder-led businesses scale?
AI helps founder-led firms scale by taking over the repeatable, document-heavy work that currently has to run through the founder because no one else has the context to do it. That includes research, proposal drafting, client reporting, and internal knowledge retrieval. Done well, this lets the firm take on more client work without a proportional increase in founder hours, which is usually the real ceiling on growth in a founder-led business, not demand.
What are founder bottlenecks?
Founder bottlenecks are the points in a business where growth is capped because critical decisions, client relationships, or institutional knowledge exist only in the founder's head. Common examples include the founder personally reviewing every proposal, being the only person who fully understands the firm's methodology, or being the default point of contact for every key client. AI reduces these bottlenecks when it's used to document, structure, and partially automate that knowledge, but it doesn't reduce them automatically just by being adopted.
How should founder-led firms manage AI costs as they scale?
Founder-led firms should manage AI costs by measuring cost per accepted deliverable, not cost per token, and by matching model capability to task value rather than defaulting every task to the most expensive model. In practice, that means routing routine work like classification and formatting to smaller models, reserving premium reasoning models for high-stakes client recommendations, and building retrieval systems that only pull in the context a task actually needs. Firms that skip this discipline tend to see AI costs grow faster than the value it produces, which erodes the margin gains AI was supposed to create.
If your firm is scaling AI use faster than you can explain the bill, that's an operating architecture problem, not a model problem. Talk to ALTA about Generative AI Services to build the routing, context, and governance that make AI spend pay for itself.




Comments