Most AI pricing conversations start with tokens, but tokens are only part of the cost structure. In real AI systems, costs also come from knowledge retrieval, workflow orchestration, infrastructure, and tool execution.
If companies design pricing based only on token costs, they risk mispricing their AI features and compressing margins.
In this article, I break down:
- the four layers of AI costs
- the lifecycle of an AI agent interaction
- how different AI products translate value into pricing while aligning with underlying costs
The Four Layers of AI Costs
AI systems typically incur costs across four layers.

Model Layer
This layer represents the cost of calling the AI model itself. Vendors typically charge based on usage metrics such as tokens, images, or audio and video seconds.
Common providers include:
- OpenAI
- Anthropic
- Google Gemini
- AWS Bedrock
For most AI applications, model inference is the most visible cost, but it is rarely the only cost.
Knowledge Layer
Before generating a response, AI systems often retrieve additional context.
This layer includes the cost of:
- creating embeddings for documents
- storing embeddings in vector databases
- retrieving relevant context during a query
In simple terms, this layer helps the AI access the right knowledge before generating an answer.
Workflow Layer
Modern AI systems increasingly operate as agents, meaning they plan and execute multi-step workflows.
Costs at this layer include:
- Agent reasoning model calls
- API/tool execution
- Orchestration compute
This layer becomes especially important for AI agents that perform actions rather than simply generate text.
Infrastructure Layer
Finally, there are the infrastructure costs required to operate the system.
These include:
- compute resources such as CPUs and GPUs
- storage for logs and datasets
- latency optimization
- monitoring and observability tools
- safety and evaluation
These costs are often less visible but become significant at scale.
Example: The Lifecycle of a Customer Support AI Agent

To understand how these layers work together, consider a simplified example.
A customer contacts an AI support agent and asks:
"Why was I charged twice last month?"
1. User Input
The customer submits a question through chat.
System actions
- Capture the message
- Retrieve session and user ID
- Load conversation history
Cost components
- Compute for processing the request
- Storage for retrieving session data
- Monitoring for system activity
2. Context Assembly
The system gathers relevant information.
Example context:
- Customer account: Premium
- Last Active: two days ago
Cost components
- Knowledge retrieval through vector search
- Data processing to construct the prompt
- Storage access for historical data
- RAG retrieval
3. Intent Understanding
The AI analyzes the user question.
Example prompt:
- User question: Why was I charged twice?
- Task: classify the intent and determine the next action.
Possible output:
- Intent: billing issue
- Next action: check billing system
Cost components
- Model inference (token input and output)
- Monitoring of model latency
4. Agent Planning and Workflow Logic
The agent determines the steps required to resolve the request.
Example reasoning chain:
- Check billing records
- Verify duplicate payment
- Confirm refund status
Cost components
- Workflow orchestration
- Compute for running the orchestration framework
- Agent reasoning (model inference)
5. Tool and API Calls
The agent retrieves information from external systems.
Examples:
- Billing system API to retrieve payment history
- CRM system to verify customer status
- Refund service to confirm refund issuance
Example retrieved data:
- Payments: Jan 2: $250, Jan 3: $250
- Refund issued Jan 4
Cost components
- API usage fees
- Backend compute resources
6. Response Generation
The AI composes the final response.
Example output:
"It appears your card was charged twice on January 2 and 3. A refund for the second charge was issued on January 4 and should appear on your statement within three to five business days."
Cost components
- Model inference for response generation
7. Response Delivery
The system sends the response back to the user.
Cost components
- Compute for response delivery
- Latency optimization such as streaming output
8. Logging and Evaluation
The system records the interaction.
System actions
- Store the conversation
- Log agent decisions
- Evaluate response quality
Cost components
- Storage for conversation logs
- Monitoring and observability tools
- Analytics for performance evaluation
Key Insight: AI Costs Are More Than Tokens
This example highlights an important point.
AI costs are not limited to LLM tokens.
As AI systems become more sophisticated, especially with autonomous agents, the cost structure expands. Workflow orchestration, tool execution, and infrastructure can represent meaningful portions of the total cost.
In contrast, simpler AI capabilities such as AI writing assistants primarily incur token costs because their main function is generating or refining text.
AI Capabilities, Pricing, and Cost Drivers

AI Infrastructure
Value: enabling developers to build AI-powered applications
Typical pricing models: tokens, images, seconds
Primary cost driver: model inference
Example vendors: OpenAI, Anthropic, Google Gemini, AWS Bedrock
AI-Enhanced SaaS
Value: improving individual productivity
Examples include AI writing, summarization, and copilots.
Typical pricing models: per seat, AI credits per seat
Cost drivers: model inference, knowledge retrieval, light workflow orchestration
Individual productivity usage is generally predictable, which makes seat-based pricing viable.
Examples:
- Notion AI included in higher-tier plans
- Airtable AI credits per seat
Some vendors choose credit-based models because they create a natural expansion path toward AI agents.
AI Agents
Value: automating task execution at scale
Examples include customer support agents, AI sales assistants, and workflow automation agents.
The value of AI agents is no longer tied to the number of internal users. Instead, it is tied to:
- the number of tasks executed
- the number of actions performed
- the quality of outcomes
For example, a company with five support agents serving 10,000 customers may consume more AI agent capacity than a company with twenty agents serving only 1,000 customers. In this case, AI usage scales with customer interactions rather than the number of internal seats.
Typical pricing models: AI credits, agent actions, outcomes or resolutions
Examples:
- ClickUp AI credits and Super Agents
- Salesforce Agentforce charging by conversations
- Intercom Fin charging per resolution
Cost drivers: workflow orchestration, tool execution, model inference
This category requires careful pricing design because usage can scale rapidly and is not constrained by seats.
AI Builders
Value: enabling users to create AI applications or workflows
The value of AI builder platforms comes from the complexity and scale of the applications users create. Usage can vary drastically between users, which is why credit-based pricing is common.
Typical pricing models: credits, credits combined with seats
Examples: Lovable, v0, Replit
Cost drivers: model inference, compute infrastructure, orchestration
Final Thoughts
When designing AI pricing, companies must clearly understand both:
- the value driver of the feature
- the underlying cost structure
AI agents and AI builder platforms tend to exhibit much more variable usage patterns than individual productivity tools. Their costs extend far beyond tokens and can grow rapidly as automation expands.
While pricing is often simplified into credits, those credits should internally reflect the full cost stack, including:
- model inference
- knowledge retrieval
- workflow orchestration
- infrastructure
Understanding these layers helps companies design AI pricing that is both simple for customers and sustainable for the business.
AI pricing must reflect the full cost stack, not just tokens.
⚡Want a Fast AI Pricing Diagnostic?
I share pricing diagnostics in a 15 to 20 minute intro call.