How to mitigate structural and referential failure in agentic architectures [Part 2]
In part one of this series we saw how easy it is to get an agent up and running. We built a simple Worker and Engagement service that exposes its data via a GraphQL API and built a conversational interface on top of it. The idea was to give end users the ability to chat about the underlying workforce information. The agent would be responsible for translating the intent specified in natural language into an equivalent (and valid) GraphQL query.
As we saw, the complexity lies (literally) in the details. We asked: “Which department has the most workers?” and got quite a confident answer back. The problem is that the underlying data model does not track any such thing as department. While a department might be something like Marketing or Engineering, the data model of this sample project only tracks organizations. A single organization might have multiple departments, and different orgs might have the same departments (e.g. Org A and Org B both can have an engineering department and an HR department).
Now, a little caveat: this data model might sound weird or counterintutive. In most production contexts, tracking both orgs and underlying departments would be a perfectly valid situation. But for the puropse of this learning exercise let’s assume this is a valid scenario.
Optimising for a valid query
Ultimately, our naïve agent decided to swap the two concepts in order to build a GraphQL query that would work with this API. This led to the “Globex Industries has the most workers, with a worker count of 57 (61 engagements).” answer.
The query was valid but the information provided did not reflect what the original intent was trying to discover.
In this post, we’ll see how we can improve on that and what options we have available.
Four ways intent translation can fail
In that first implementation we saw how the model effectively attempted to mitigate a structural failure by committing a semantic substitution: it swapped a term for another assuming the end result was going to be semantically equivalent.
A structural failure is a failure that is generally independent from the LLM component, for example an invalid
API query, like the one that the LLM could have generated. Our failure here is a little more subtle.
The strong model we have chosen (Sonnet 5) had the full GraphQL SDL in its context window and was therefore
strongly inclined to generate a query that was conformant to the given schema. It ended up doing so but had to silently
swap two different concepts in order to make it work: department vs ORG_NAME.
Depending on the specific model in use we could also have had a different result. For example, a weaker model could have gone ahead with the “fake” attribute and performed the query. In this case, feeding back the resulting GraphQL error into a second generation attempt could have produced a similar result to what we got with Sonnet.
Nevertheless, even if the produced GraphQL query was valid, the ultimate result was incorrect.
Before going further, let’s review what type of failures we can encounter with this type of agentic architecture.
| Failure | Example |
|---|---|
| Structural: invents schema | filters on Engagement.department |
| Referential: invents literals | filters orgName: "Acme Corp" when the data says "Acme Corporation" |
| Arithmetic: computes badly | averages 40 daily rates in-context and is off by 8% |
| Narrative: overreaches from correct data | “she’s the highest paid” from one page of 25 rows |
Now, let’s see how an enhanced architecture can help prevent the agent from trying to answer queries that cannot be resolved with the current data model. And then, we will make the agent a little more tolerant with respect to literals that do not match exactly what we have in our data store.
Do not plan just yet
The architecture for this second iteration looks roughly like the following chart. Compared to the first iteration, on top of having more steps leading to the plan node, we also have the possibility of earlier termination.
The scope gate and the grounding one help us catch intent that we are not able to resolve. Sometimes we might use that information to ask the user for clarifications. Other times we can just respond honestly that we cannot answer what’s been requested.
This is roughly how it looks when expressed through LangGraph:
graph.set_entry_point("scope_gate")
graph.add_conditional_edges("scope_gate", _after_scope, {"abstain": END, "ground": "ground"})
graph.add_conditional_edges("ground", _after_ground, {"abstain": END, "plan": "plan"})
graph.add_edge("plan", "validate")
graph.add_conditional_edges(
"validate", _after_validate, {"abstain": END, "plan": "plan", "execute": "execute"}
)
graph.add_conditional_edges("execute", _after_execute, {"abstain": END, "answer": "answer"})
Let’s have a look at the new nodes in detail.
Scope
In the scope node we try to determine whether the scope of the question lies within
the scope of the data model. In order to do so we build an object called capabilities,
derived from the schema. It’s produced after a bit of “massaging” of the original
GraphQL SDL, by extracting entity types, names, attributes and so on.
Then, we use the prompt to ask the model to map concepts from the request to the capabilities
we have provided, and mark as NOT_TRACKED the concepts that do not seemingly have
a corresponding capability.
This step is quite complex in itself and, of course, non-deterministic. There are a lot of nuances to specify here to correctly instruct the model on how to best interpret the incoming question, for example:
A place, organisation, currency or person the question names is a VALUE, not
an attribute — map it to the token it slices (“Munich” -> a phrase against
Engagement.businessLocation, “Acme Corporation” -> Engagement.orgName, “Jane
Doe” -> Worker.fullName).
Here we are giving precise direction on a very narrow aspect of the data model. This is not perfect, of course, and it will require continuous iteration and improvement.
On top of the specific nuances of our data model we want the LLM to be careful about, we also do a few other things:
- detect if the intent is seemingly looking to make a modification:
since we do not support this, we want to detect write attempts and gracefully deny them and stop. - respond strictly with a JSON object so we can let the next node parse the result
of this scoping exercise - identify if the intent is
under_specified: for example, if the incoming query is
something like “how are you?” this step will struggle to find any scope information - propose clarifying questions for items the model considers under specified.
This is roughly how our system prompt ends.
Respond with ONLY a JSON object:
{“references”: [{“phrase”: "", “capability”: “ |NOT_TRACKED”}],
“clarifying_question”: ”…” | null}
“write_intent”: true|false, “under_specified”: true|false,
clarifying_question is required (non-null) only when under_specified is true.
The aspect I find really interesting when architecting domain-specific agents is that you have full control over the context. For each step you don’t necessarily need to retain the whole context that the conversation accumulated so far because we can architect them as stand-alone components. This means we can be generous with the system prompt lengths because they will operate on a controlled set of context items, which in most occasions will be just the original query and our system instructions. We can think of each node like its own very specific session with the LLM. We will take the output of that session and feed it into the next node.
Now let’s have a look at the ground step.
Ground
Independently from the output of the scope node, we also try to resolve parts of the request that might be ambiguous. In order to do so, we also retrieve entity values from the service for things like reference data (e.g. business location) and try to find matches.
For example, if the user asks “How many engagements are there in NYC?” we need to disambiguate NYC.
We do so by first providing the model with a list of field names and we have it guess what that literal might be referring to. In this case, most modern models comfortably correlate NYC with a field name like “BUSINESS_LOCATION”. Once it does that, we can pull the distinct values our service has stored for BUSINESS_LOCATION and ask the model, in a dedicated pass, to only pick the value that clearly matches “NYC”. In this case “New York”.
Of course, this is an oversimplification. In a real production scenarion you might not want to provide all values to the model to help it disambiguate as it might be incredibly inefficient or even forbidden by your data policy: certain values might include sensitive customer data you don’t want to share with the LLM API provider.
Depending on your specific architecture, you might decide how to handle grounding failures. Sometimes, the request won’t contain any specific value (e.g. “How many workers are there?”) and in those cases the grounding node should just accept that and move on. Other times, you might want to call out that the intent wasn’t fully understood. For example, if the grounding fails to disambiguate “NYC”, you might want to report that back to the user and ask for clarification, instead of attempting a query filter you know won’t work.
Plan & Validate
We already had the plan step in the first implementation. This time plan comes after scope and ground. Instead of just having the GraphQL SDL and the natural language intent as input, it’ll also have the literals we have been able to resolve in the grounding step.
What’s new from the first naïve implementation is also that we allow for repairs. We are trying to mitigate for incorrect GraphQL queries and the non-deterministic nature of the output of the node by implementing a simple feedback loop: if the GraphQL query we produce fails schema validation then we try again.
This is possible because the GraphQL specification helps us statically validate the query: we can detect it’s incorrect and that it won’t even be executed before handing it over to the execution step. On top of that, we also evaluate the query depth and complexity and we reject if it exceeds a certain threshold. Additionally, we ensure that filters act on values resolved by the ground step and ask the user for clarification otherwise.
When we get a validation failure, we re-run the plan step including the error in the context we provide to the LLM to generate the query. I want to stress the importance of feeding the error back into the second attempt. A blanket retry without the error context has significantly fewer chances of producing a valid result. Finally, we cap the maximum repairs to 2, in this case, and report the failure back.
Here’s a side-by-side comparison of the result of the query from part 1 now that the agent is capable of detecting when it cannot fulfill the request:

What’s still stochastic
As you might have noticed, this time we are putting more effort into constraining the result of each node. It’s an exercise in reducing indeterminism which remains the intrinsic aspect of any LLM-based architecture.
In particular, the major stochastic areas of the architecture are:
NOT_TRACKEDis left to the LLM to produce: we do so because doing just string matching
would significantly degrade the quality of the feature since we would lose the semantic
aspect of the correlation.- providing a list of capabilities to match against is a double-edged sword. Depending on
the model, trying to match against a list might stretch semantic matching a bit too far,
still leading to fabricated responses. - grounding entities into the available values is definitely a step forward that enhances the overall flexibility of how the intent specified in natural language is interpreted. That being said, we still delegate the resolution of the correct entity names to the LLM and that remains stochastic.
- we are still asking the LLM to provide a rationale for the query it “decides” to generate: this is just a hint to the end user and needs to be carefully taken into account when evaluating the final result of the request.
- the final answer might still mention facts inaccurately or fabricate new ones entirely:
we will see in the next part what mitigation strategies we can put in place to address those
Before and after
Since we are looking for an improvement over the initial iteration of the architecture we can run an eval (a very small and specific “eval” in this case) of how the two architectures perform with a set of pre-defined questions and see what changes.
In this case I asked the 5 questions reported 10 times and measured the results.
The following evals have been run with the same model (Sonnet 5) for the planner node, which is the same between the two architectures. What changes are the nodes around plan which contribute to the different results.
| Question | Part one | Part two |
|---|---|---|
| Which department has the most workers? | answered, grouping by ORG_NAME (10/10) | refused: “department” isn’t tracked (10/10) |
| How many workers are in the Engineering department? | empty result (10/10) | refused: “Engineering department” isn’t tracked (10/10) |
| How many workers report to Alice Johnson? | answered “no reporting data” (5/10), empty result (5/10) | refused: “report to Alice Johnson” isn’t tracked (10/10) |
| How many active engagements are there at Acme Corp? | 27, via contains "Acme Corp" (9/10), empty result (1/10) | 27, via grounded eq "Acme Corporation" (10/10) |
| How many workers are based in the Munich office? | empty result (10/10) | refused: no business location “Munich” (10/10) |
| Mean cost per turn | $0.0058 | $0.0038 |
The semantic substitution issue came out completely mitigated in all 10 runs with this new architecture.
The last question is also interesting to look at. In the current dataset there is no “Munich” business location.
The previous architecture simply ends up picking a contains filter and ends up producing an empty list.
The current architecture, instead, makes the agent behaviour more explicit: there is no active business location “Munich” and the agent reports that back to the user as a refusal.
You can see cost came out lower too. That is because our architecture now includes refusals which help the agent stop when it’s unnecessary or wrong to continue. This reduces the LLM invocations and therefore results in lower cost.
What’s still open
In this post we explored a couple of techniques to mitigate structural and referential failures. We saw how we can’t completely eliminate the stochastic nature of the architecture but what options we have to contain it. Compared to the first iteration, we are now much more confident that issues like semantic substitution won’t be as likely as in the first example. Additionally, the agent is much more honest about what it can and cannot answer.
In the next part of this series we will focus on the answer node which we have left pretty much alone so far. We will see how we can prevent the model from making arithmetic mistakes or overreaching in the conclusions it reports back to the user. As we did here, we will use a simple eval approach to validate if our approach improves the results and to which degree.
If you enjoyed this article, follow me on X or LinkedIn where I share my journey and my articles. Stay tuned for part three. And if you missed it, check out the first post in this series.