
As AI agents become part of everyday development workflows, one topic is becoming increasingly important: how efficiently those agents use AI resources. Recent changes across AI developer platforms have made consumption, efficiency, and cost management a bigger consideration for teams building AI-powered solutions.
This made me look more closely at something I was already experiencing while building AI agents for Business Central scenarios. The question was no longer only “Can the agent solve the problem?” but also “Can the agent solve the problem efficiently?”
When I started building AI agents, my first focus was capability. I wanted the agent to understand business context, analyze problems, and provide useful recommendations. The early results were impressive. The agent could process information, identify patterns, and generate responses that were genuinely useful.
However, as I tested different scenarios, I noticed another challenge. Token consumption was growing faster than I expected. A simple request could consume thousands of tokens because the agent was not only processing the user’s question. It was also processing instructions, conversation history, available tools, documents, examples, and additional context provided during the workflow.
This became an interesting engineering challenge because building an AI solution for real-world usage is very different from creating a simple demonstration. A production-ready AI agent needs to be accurate, consistent, fast, and cost-effective.
So I started experimenting with different approaches to reduce unnecessary token usage without reducing the quality of the responses.
The First Lesson: More Context Does Not Always Mean Better Results
My initial assumption was simple: if I provide the agent with more information, it will produce better answers. I started giving the agent more documentation, more examples, more instructions, and more historical information,but during testing, I realized something important. More context does not always create more intelligence. Sometimes it creates more noise.
The agent spends time processing information that is not relevant to the actual problem. The biggest improvement came when I changed my approach: instead of providing everything, provide what actually matters.
Rather than sending large amounts of information and asking the AI to find the answer, I started preparing focused context before the AI analysis. This improved the quality of responses and reduced unnecessary processing.
Creating Better Instructions
One of the biggest improvements I applied was separating instructions from conversations. When I started building AI agents, I noticed that I was repeatedly sending the same explanations in every prompt: what role the agent should play, how it should analyze information, what type of response it should provide, and what limitations it should follow. Although these instructions were important, repeating them increased token usage unnecessarily.
To solve this, I started creating dedicated instruction files such as .md files that define the agent’s behavior, analysis approach, response format, and boundaries. This allows the agent to understand how it should operate without consuming tokens repeatedly explaining the same rules.
The conversation can then focus on the actual business problem instead of carrying repeated instructions. This small change improved consistency, reduced unnecessary token usage, and made the agent workflow easier to maintain as the solution continued to evolve.
Giving Agents Only the Tools They Need
Another important learning came from how tools are provided to AI agents. When building AI workflows, it is easy to expose every available capability because it feels like giving the agent more power. However, I realized that more tools do not always mean better results. Each additional tool adds more context, more possible paths, and more complexity for the agent to evaluate.
While experimenting with different Business Central scenarios, I noticed that each type of agent has a different purpose. A Business Central error analysis agent does not need the same tools as an AL code review agent. Similarly, a performance analysis agent requires different capabilities compared to a documentation assistant.
I started following a simple approach: provide the agent only with the tools required for the specific task it needs to perform. This helped reduce unnecessary processing, made the responses more focused, and created workflows that were easier to control and maintain.
The lesson I learned was that a powerful AI agent is not the one with access to everything. It is the one that has the right capabilities available at the right time.
Preparing Data Before AI Analysis
Another major improvement I applied was changing where the analysis begins. Initially, it was natural to provide the AI agent with as much information as possible and expect it to identify the important details. However, I realized that this approach made the agent spend valuable processing time searching through information that was not always relevant to the actual problem.
Instead of saying, “Here is all the information. Find the problem,” I changed the approach to, “Here is the relevant information. Analyze this specific situation.”
This meant preparing structured inputs before sending them to the AI model. The agent receives the actual issue description, relevant objects, error details, business context, and any other information required for the analysis instead of large amounts of unnecessary data.
This small change made a big difference. The AI no longer needed to spend tokens filtering through information before starting the actual reasoning process. It could focus directly on understanding the problem, identifying possible causes, and providing meaningful recommendations.
The lesson I learned was that AI works best when we do the preparation work around the problem. The quality of the output depends not only on the model, but also on the quality and relevance of the information we provide.
Using Focused Agents
Another important lesson I learned was that one large AI agent is not always the best design approach. When building an agent that tries to handle every possible scenario, it becomes harder to optimize because it requires more instructions, more context, and more information to make the right decision.
Initially, it may seem convenient to create one powerful agent that can handle everything. However, different tasks require different knowledge and different ways of thinking. A code review agent needs to understand development patterns and architecture. An error analysis agent needs to focus on troubleshooting and root cause identification. A performance analysis agent requires a different set of data and investigation methods.
Based on this experience, I started moving towards smaller, focused agents with clear responsibilities. Each agent receives only the information and tools required for its purpose. This makes the workflow more efficient, easier to maintain, and easier to improve over time.
The biggest takeaway for me was that a smarter AI solution does not always come from making one agent bigger. Sometimes better results come from creating multiple focused agents that work together.
Improving Response Quality Through Structure
Another area where I noticed unnecessary token consumption was the output itself. AI models are naturally capable of generating very detailed responses, but in business scenarios, users usually need clarity more than volume. A long explanation does not always mean a better answer, especially when consultants or support teams need to quickly understand an issue and decide the next action.
I started guiding the agent to provide structured responses instead of asking for a general explanation. Rather than generating a large response, the agent focuses on specific areas such as the problem description, possible root cause, business impact, recommendation, and next action.
This small change improved both efficiency and usability. The responses became shorter, more consistent, and easier for users to consume. It also helped ensure that the agent focuses on delivering actionable information instead of generating unnecessary details.
The lesson I learned was that controlling the output is just as important as controlling the input. A well-designed AI workflow should manage both what the agent receives and what it produces.
Final Thoughts
Building AI agents is not only about selecting a powerful model. The model is important, but the overall design of the workflow around it plays an equally important role. The way we provide context, define instructions, select tools, and structure the interaction can have a major impact on the efficiency and quality of the final solution.
Just like we optimize database queries, APIs, and application performance, AI workflows also require optimization. A well-designed AI solution should not only produce accurate results but should also be practical, scalable, and efficient when used in real-world scenarios.
The future of AI development will not only be about creating smarter agents. It will also be about creating agents that can deliver consistent value while using resources effectively. A good AI agent is not the one that consumes the most tokens or processes the largest amount of information.
It is the one that understands the problem, uses the right context, applies the right tools, and delivers the right answer efficiently. Building successful AI solutions is becoming less about adding more capabilities and more about designing intelligent workflows that create real business value.
Leave a Reply