"Your agent is stupid."
That is not the feedback you want after months of work.
Especially when, the day before, every test said the opposite.
Our agent passed almost 100% of its evaluation tests. We shipped it convinced it worked very well, and soon after, users started telling us a completely different story.
The problem was not the model we had chosen, and it was not a bug. We had got something much subtler wrong.
We had assumed how people would use the agent.
We will come back to that shortly.
This article comes from episodes like that one.
Over the past few years at VLK Studio, we have built agents and LLM-based systems for our own products and for our clients, and put them in the hands of thousands of people.
We tried different models, changed architectures, and built tools, evals, retrieval systems, and workflows that, often only a few months later, we questioned all over again.
Every project left us with something.
And, curiously, the lessons that stayed with us most clearly were rarely of the kind "we needed a more powerful model" or "we should have written a better prompt".
Much more often they were small adjustments.
Knowing when to guide an agent and when to leave it room. When it should ask a question and when it can decide on its own. Which tools to give it. How much to trust an evaluation test. How to let it make mistakes without causing damage. And, above all, how to give it enough information to notice on its own when something is not working.
Some of these things seem almost obvious to us today. They were not obvious while we were learning them.
Models and frameworks keep changing, and many of the practices we use today may look naive to us a few years from now. Some lessons, though, seem to hold up.
This article collects the ones that most changed how we build agents. It is not a framework, and it is not a definitive list of best practices. These are small lessons learned after making a lot of mistakes.
| Lesson | In practice |
|---|---|
| Don't design every path | Define the agent's scope of action, and give it tools and building blocks it can choose and combine on its own. |
| Test the mess | Evals should include ambiguous requests, negative cases, out-of-scope input, and multi-turn conversations, not only happy paths. |
| Asking is an action | Ask for clarification when the missing information can change the result, the risk, or what the agent is allowed to do. |
| Mistakes aren't the problem | Give the agent observable signals so it can notice when something is off, understand why, and correct its course. |
| Feedback is a sensor, not a verdict | The result of a search or a tool describes what the agent observed. It does not necessarily describe everything that exists, or the whole truth. |
| Knowing what you don't know | The agent should know which tools it has, what it has observed, what it has not observed yet, and the limits of its environment. |
Don't outsmart your users
In traditional software development, it is normal to think through the interactions a user might have with the product, and to design a UX that tries to anticipate and steer those actions.
Once you give users a free-text field to talk to your app, things change.
You cannot predict what they will write. It sounds like a small detail, but it completely changes how you build the agent underneath.
Is the agent stupid?
We built an agent for one of our clients. To keep it brief, it had to find documents from users' questions and generate answers from the content of those documents.
At the start we built a very rigid workflow.
Every user query ran a semantic search over the documents.
A second pass interpreted the documents that came back, and then the agent generated an answer.
It was a reasonable approach. The request entered the system and triggered a search. The system interpreted the documents, and the agent wrote an answer.
To check that the workflow worked, we collected about 50 scenarios the client had reported and built a suite of evaluation tests.
The tests used another LLM as a judge, scoring each answer against an expected answer for that question.
The result looked excellent. The success rate was close to 100%.
We had good reason to think the system was ready to release.
Once we released the bot, we read the first user feedback. "The agent is stupid!" Not what we expected.
So we went and looked at what users were actually asking the bot. The result surprised us.
Questions like:
AD-12aa45, a seemingly meaningless string that was in fact the identifier of the document they were looking for, dropped into the chat with no context;- "Can you give me information on CD?", an internal acronym of the client's that meant nothing to the model, and risked being confused with known acronyms such as Compact Disc;
- "Find documents similar to the second one you suggested", a search based on the previous search, which our workflow had not provided for.
The problem was clear. We had imagined a way of using the agent that did not match real use.
The problem was the journey

In a traditional interface, where the user can fill in a form, follow a wizard, or click dedicated buttons, you can design a user journey and try to lead them toward the best path. When you give users full freedom to express themselves in natural language, you cannot think in exact flows. You have to think about the scope of action.
So the problem was not that the workflow was too simple. We had tried to design the user's path in advance, instead of clearly defining what the agent could do.
The opposable thumb
We then did something very simple.
Instead of adding more branches to the workflow, we gave the agent its opposable thumb. Tools it could choose and combine on its own.
One for searching the text, one for applying filters, such as a document identifier, one for recognizing internal acronyms, and so on.
The agent could then choose how to handle the request without losing sight of the task it had been built for. We did not have to predict every user journey. We had to define the goals well, and the building blocks that would let it reach them.
If you give users a text field with no limits, expect them to ask you for everything.
Define the goals clearly, figure out which building blocks can let the agent reach them, and then give it room to maneuver.
Does this apply only to agents, or to us ordinary mortals too?
An eval can be perfectly wrong
We had made the agent more flexible. When we tried to measure that improvement, something unexpected happened.
The better agent scored worse
To be thorough, we ran the evaluation tests again on the new version. Our sense was that the agent had clearly improved. It was more responsive, better at recovering from errors, more interactive. More intelligent, in short.
The tests disagreed. The success rate had dropped. We were barely scraping 70%.
We decided to trust our own sense of it and release the new agent behind feature flags, then wait for the first feedback. A sigh of relief. The chatbot could finally handle real cases.
Users were happy with how dynamic it was, and they also appreciated that it turned some questions down, asking for more information before producing a correct answer.
Our evals had the same bias we did
Back to the evals. What was going on?
The tests had the same bias as the workflow. They measured a very narrow slice of the problem very well and, for that reason, gave us a false sense of correctness.
The early cases were almost all happy paths, and they assumed users who knew what they were doing and could phrase clear, complete requests.
Questions like:
"I am looking for information on ACME Inc. and, based on what you find, I would like an analysis of which of our services I can offer their sales department."
are perfectly legitimate.
But they assume a discipline and a clarity that real users simply do not always have.
Test the mess

From a tool that claims to be intelligent, we expect it to handle much less structured questions too.
We also expect it to recognize a request that falls outside its purpose.
An agent tasked with searching and analyzing documents can certainly write a poem about the beautiful sea of Sardinia. That does not mean it should.
So we rewrote the test suite completely, adding less structured queries and negative cases.
In some cases the agent was not supposed to answer, because it did not have enough information to go on.
It should prefer to stop and ask for clarification instead of guessing, and then act on what it learns.
Conversations, not prompts
Another piece was missing. Testing the conversation.
We cannot test only interactions made of a single message. Users iterate on answers, ask for more depth, come back to earlier messages after long discussions, and sometimes contradict themselves.
Testing this kind of interaction is much harder, but it makes the difference between being called "stupid" and being regarded as
"a tool that completely changed the way I work, something I can't do without".
That last one was the most rewarding feedback on this project.
An eval can measure one part of the problem very precisely and, at the same time, tell you little about how useful the agent is in practice.
To know when the agent can continue and when it must stop, better tests are not enough.
It also has to know when something is missing.
Knowing when to stop
Asking is an action
Letting the agent stop and ask for confirmation or clarification was certainly the most important adjustment of all, and also one of the hardest to handle correctly.
This was one of the most interesting surprises of that release. One of the behaviors users perceived as most intelligent was, technically, not acting.
The agent had learned to stop and ask.
I often like this exercise.
For a moment, pretend the agent is a person. You ask:
"Give me the best documents you can find on artificial intelligence."
Best... but based on what?
The most read? The ones with the most accurate information? The ones published by the most authoritative sources?
Picking one of those interpretations at random means rolling a die and hoping it lands on the face you had in mind.
A person would probably ask:
"What do you mean by best?"
And, once they had the answer, they would continue the work.
Why, then, do we presume that an agent can guess the face of the die?
The cost of questions

Let's stay with the same exercise.
"How do I turn on airplane mode on my phone?"
How would you answer? Think about it for a moment. Here is one possible continuation of the conversation:
"Of course, I am here to help. Before I answer, I need to know what phone you have."
"I have an iPhone."
"Okay, which model?"
"Hmm, I don't know, I think it's the 17. The plane leaves soon, please just tell me how."
"I understand your frustration. It is very important to turn on airplane mode before the plane takes off. Which city are you leaving from?"
"..."
The conversation is deliberately exaggerated, of course, but it carries one point.
Clarifying makes sense when the missing information can change the answer or the agent's decision in a way that matters.
Too many questions cause frustration, and they raise the chance that the agent gets confused and loses track of its primary goal.
A good way to avoid that behavior is to reason backward. Start from the goal and ask which information you actually need in order to reach it.
"You are asking me how to turn on airplane mode. What do I need so I can answer?"
"I should know which phone, sure, but the steps are practically the same on every phone now."
"The only real exception is a very old phone, but I will assume a modern one."
Following that practice, the agent would have answered immediately, with no questions at all.
"Go to Settings and you will find the button to turn airplane mode on and off right away. Tap it to turn it on. Have a good trip!"
No interruption, no frustration. Hopefully the user is not carrying a Nokia 3310!
Defaults are decisions
And what if they really are carrying a Nokia 3310?
Proceeding with a default is not neutral. Proceeding without asking still means choosing for the user.
In this case it is simple. It would have been enough to make the assumption explicit. "On most modern phones..." Then leave the user a way out if that does not hold for their device.
The agent must ask when the missing information can change the result, the risk, or what it is authorized to do. It can proceed with a default when the assumption is explicit and any error that follows is cheap, reversible, and easy to correct.
Whether the agent proceeds on an assumption, or does so after asking the user to confirm, there is always a risk of taking the wrong path. Especially when the task is complex and takes many steps.
Mistakes are inevitable.
To err is human, to forgive divine.
Mistakes aren't the problem
The problem is not that an agent can make a mistake.
The problem is that it can keep making mistakes without noticing.
The more autonomy we give it, the more important it is to also give it something that can contradict it. A failing test, an error returned by a system, a figure that does not add up, an action whose result differs from the one expected.
This is one reason software is a particularly good environment for agents.
Almost every action leaves observable signals behind. The code compiles or it does not, a test passes or fails, an application leaves logs and stack traces, and a change can be compared through a diff.
And often the environment does not only say no.
It also says why.
The agent can try a solution, observe what happened, and use that information to decide what to do next.
Try. Observe. Correct. Try again.
As long as something keeps telling it whether it is moving closer to the result or farther away.
That is a large advantage.
But not every environment can do it in the same way.
Imagine, for example, asking an agent to write an article.
Any recent model can produce text that is grammatically correct, well structured, and apparently complete.
But how does it know whether the text is actually interesting?
A compiler can tell it that line 13 has a syntax error. There is no equally precise equivalent that says:
"You lost the reader here."
or:
"The sentence is formally correct, but it says nothing worth remembering."
That does not mean good writing cannot be evaluated. There are rules, editors, metrics and, above all, readers who can produce feedback.
That feedback, though, tends to be less immediate, less precise, and harder to turn into the next action.
And that is the difference we care about.
The more an environment can return signals that are fast, specific, and tied to what the agent just did, the more we can let it observe the result and correct its own course.
When those signals exist naturally, we already have an important part of the loop.
When they do not, we have to find a way to build them. With one caution. An environment that responds is not necessarily an environment that tells the truth.
Feedback is a sensor, not a verdict

Back to our stupid agent. Another report from the early tests was about inconsistent results.
The same request, from different users, or even repeated by the same user, could return different documents.
Why?
Take a request like:
"Find articles related to VLK Studio in application development."
Before running the search, the agent rewrote the request into a form better suited to semantic search.
One possible query was:
VLK Studio software engineering, web and mobile application development
and it might produce, for example, 4 results.
The next time, though, the query rewriter might read the same intent in a slightly different way:
VLK Studio digital product development and engineering case studies
An equally sensible query.
And yet it might retrieve three different documents, perhaps only partly overlapping with the previous set.
From the user's point of view this was hard to explain. They had asked the same question. Why was the agent giving them a different answer?
The problem starts when we treat a search result as if it were a verdict.
Those four documents do not mean:
"These are the VLK Studio articles related to application development."
They mean something much more limited:
"These are the documents this search, worded this way, brought to the surface."
A different wording, a different filter, or another tool might surface others.
A search result is therefore an observation of the corpus, not a complete picture of what the corpus contains.
It is a sensor.
And like any sensor, it can give us enough information to decide what to do next without necessarily telling us the whole truth.
To know how much to trust those four results, the agent also has to know something about its own process. Where it searched, with which query, which filters it applied, which sources were available, and which were not.
Not knowing is information

Back to the previous example.
If two different queries are both plausible and can return different results, how do we tell the agent:
"Stop, you found what you were looking for."
At first we tried to get around the problem by running two or three query rewrites in parallel.
In most cases they retrieved almost the same documents. When they did not, the problem was still unsolved. We had no guarantee that a fourth query would not find more.
If the user had asked:
"Find me all the articles in which VLK Studio is associated with software development."
we would have had no reliable way to know when to stop.
The problem was not finding a still better query.
We were missing information about the environment.
The agent knew which documents it had found.
It did not know which documents it could have found.
Sometimes, adding extremely simple sensors is enough to change the situation completely.
To keep the implementation details out of the way, let's simplify the problem a great deal and imagine giving the agent two small sensors.
The first answers a very plain question:
"How many documents in the index mention VLK Studio?"
Suppose the answer is:
10
Our semantic search found four.
That does not mean the other six are necessarily relevant. Some might be about design, branding, or any other activity that has nothing to do with software development.
But the agent now knows something it did not know before.
Six other documents exist that it has not evaluated yet.
So we add a second sensor:
"Given a document ID, tell me briefly what it is about."
The agent can now quickly check the six documents that did not come up in the initial search, and decide whether any of them are relevant.
The process then looks something like this:
- Semantic search for
VLK Studio software engineering, web and mobile application developmentreturns 4 results. - Sensor question,
How many documents mention VLK Studio?, returns 10. - The agent knows that 6 more documents still have to be checked.
- It asks for a short summary of the six missing documents.
- Five turn out to be irrelevant. One is actually about software development.
- It adds that document to the results and answers the user.
If the initial search had been worded differently and had returned only three documents, the process would have been the same.
We could have started from different points and still arrived at a very similar set of results.
Not because the search had become deterministic, but because the agent finally had enough information to understand how complete its own observation was.
And that is what I mean by self-awareness in an agent.
Not consciousness, introspection, or anything anthropomorphic.
Much more simply:
Knowing which tools it has, what it has observed, what it has not observed yet, and what limits its environment has.
It wasn't stupid. It was limited.
At the start of this article we told you about the rather cruel feedback we received on one of our first agents:
"Your agent is stupid."
At the time the field was much less mature, and we were learning almost everything while we built these systems.
We had to learn by doing. Make mistakes, understand why, and try again.
Since then we have changed, and so has the field. Many of the lessons told here are now part of a much broader conversation.
With hindsight, though, we can finally answer that first piece of feedback.
The agent was not stupid.
It was limited.
Limited by the paths we had decided for it in advance.
Limited by an environment that let it act, but told it too little about the consequences of its actions.
Limited by our difficulty in knowing when it should proceed and when it should stop, ask, or change direction.
And, above all, limited by what we had assumed about how users would use it.
In 2026, models have become extraordinarily capable. They can reason, use tools, write code, search for information, and take on tasks we would have considered out of reach only a few years ago.
But maybe the most important lesson we carry from these years is this one.
A good model is not enough to build a good agent.
The tools, the information, the feedback, the boundaries, and the freedom we build around the model determine, in large part, what the agent will be able to do.
In the end, an agent is as capable as the environment we build around it allows it to be.
Just like us.
