alediaferia.com
Experiments

How to mitigate structural and referential failure in agentic architectures [Part 2]

In part one of this series we saw how easy it is to get an agent up and running. We built a simple Worker and Engagement service that exposes its data via a GraphQL API and built a conversational interface on top of it. The idea was to give end users the ability to chat about the underlying workforce information. The agent would be responsible for translating the intent specified in natural language into an equivalent (and valid) GraphQL query.

As we saw, the complexity lies (literally) in the details. We asked: “Which department has the most workers?” and got quite a confident answer back. The problem is that the underlying data model does not track any such thing as department. While a department might be something like Marketing or Engineering, the data model of this sample project only tracks organizations. A single organization might have multiple departments, and different orgs might have the same departments (e.g. Org A and Org B both can have an engineering department and an HR department).

Now, a little caveat: this data model might sound weird or counterintutive. In most production contexts, tracking both orgs and underlying departments would be a perfectly valid situation. But for the puropse of this learning exercise let’s assume this is a valid scenario.

Optimising for a valid query

Ultimately, our naïve agent decided to swap the two concepts in order to build a GraphQL query that would work with this API. This led to the “Globex Industries has the most workers, with a worker count of 57 (61 engagements).” answer.

The query was valid but the information provided did not reflect what the original intent was trying to discover.

In this post, we’ll see how we can improve on that and what options we have available.

Four ways intent translation can fail

In that first implementation we saw how the model effectively attempted to mitigate a structural failure by committing a semantic substitution: it swapped a term for another assuming the end result was going to be semantically equivalent.

A structural failure is a failure that is generally independent from the LLM component, for example an invalid API query, like the one that the LLM could have generated. Our failure here is a little more subtle. The strong model we have chosen (Sonnet 5) had the full GraphQL SDL in its context window and was therefore strongly inclined to generate a query that was conformant to the given schema. It ended up doing so but had to silently swap two different concepts in order to make it work: department vs ORG_NAME.

Depending on the specific model in use we could also have had a different result. For example, a weaker model could have gone ahead with the “fake” attribute and performed the query. In this case, feeding back the resulting GraphQL error into a second generation attempt could have produced a similar result to what we got with Sonnet.

Nevertheless, even if the produced GraphQL query was valid, the ultimate result was incorrect.

Before going further, let’s review what type of failures we can encounter with this type of agentic architecture.

FailureExample
Structural: invents schemafilters on Engagement.department
Referential: invents literalsfilters orgName: "Acme Corp" when the data says "Acme Corporation"
Arithmetic: computes badlyaverages 40 daily rates in-context and is off by 8%
Narrative: overreaches from correct data“she’s the highest paid” from one page of 25 rows

Now, let’s see how an enhanced architecture can help prevent the agent from trying to answer queries that cannot be resolved with the current data model. And then, we will make the agent a little more tolerant with respect to literals that do not match exactly what we have in our data store.

Do not plan just yet

The architecture for this second iteration looks roughly like the following chart. Compared to the first iteration, on top of having more steps leading to the plan node, we also have the possibility of earlier termination.

invalid, max 2 repairs

out of scope / unclear

unknown / ambiguous value

still invalid

empty result

Question

Scope gate

Entity grounding

Plan GraphQL

Validate

Execute

Answer

Response

Templated refusal or question

The scope gate and the grounding one help us catch intent that we are not able to resolve. Sometimes we might use that information to ask the user for clarifications. Other times we can just respond honestly that we cannot answer what’s been requested.

This is roughly how it looks when expressed through LangGraph:

graph.set_entry_point("scope_gate")
graph.add_conditional_edges("scope_gate", _after_scope, {"abstain": END, "ground": "ground"})
graph.add_conditional_edges("ground", _after_ground, {"abstain": END, "plan": "plan"})
graph.add_edge("plan", "validate")
graph.add_conditional_edges(
    "validate", _after_validate, {"abstain": END, "plan": "plan", "execute": "execute"}
)
graph.add_conditional_edges("execute", _after_execute, {"abstain": END, "answer": "answer"})

Let’s have a look at the new nodes in detail.

Scope

In the scope node we try to determine whether the scope of the question lies within the scope of the data model. In order to do so we build an object called capabilities, derived from the schema. It’s produced after a bit of “massaging” of the original GraphQL SDL, by extracting entity types, names, attributes and so on.
Then, we use the prompt to ask the model to map concepts from the request to the capabilities we have provided, and mark as NOT_TRACKED the concepts that do not seemingly have a corresponding capability.

derived at import

closed list of tokens

department → NOT_TRACKED

re-check every pick

SDL

Capability manifest

Question

LLM: extract references

decide_scope, pure code

ANSWERABLE

OUT_OF_SCOPE, naming the attribute

NEEDS_CLARIFICATION

This step is quite complex in itself and, of course, non-deterministic. There are a lot of nuances to specify here to correctly instruct the model on how to best interpret the incoming question, for example:

A place, organisation, currency or person the question names is a VALUE, not
an attribute — map it to the token it slices (“Munich” -> a phrase against
Engagement.businessLocation, “Acme Corporation” -> Engagement.orgName, “Jane
Doe” -> Worker.fullName).

Here we are giving precise direction on a very narrow aspect of the data model. This is not perfect, of course, and it will require continuous iteration and improvement.

On top of the specific nuances of our data model we want the LLM to be careful about, we also do a few other things:

  • detect if the intent is seemingly looking to make a modification:
    since we do not support this, we want to detect write attempts and gracefully deny them and stop.
  • respond strictly with a JSON object so we can let the next node parse the result
    of this scoping exercise
  • identify if the intent is under_specified: for example, if the incoming query is
    something like “how are you?” this step will struggle to find any scope information
  • propose clarifying questions for items the model considers under specified.

This is roughly how our system prompt ends.

Respond with ONLY a JSON object:
{“references”: [{“phrase”: "", “capability”: “|NOT_TRACKED”}],
“clarifying_question”: ”…” | null}
“write_intent”: true|false, “under_specified”: true|false,
clarifying_question is required (non-null) only when under_specified is true.

The aspect I find really interesting when architecting domain-specific agents is that you have full control over the context. For each step you don’t necessarily need to retain the whole context that the conversation accumulated so far because we can architect them as stand-alone components. This means we can be generous with the system prompt lengths because they will operate on a controlled set of context items, which in most occasions will be just the original query and our system instructions. We can think of each node like its own very specific session with the LLM. We will take the output of that session and feed it into the next node.

Now let’s have a look at the ground step.

Ground

Independently from the output of the scope node, we also try to resolve parts of the request that might be ambiguous. In order to do so, we also retrieve entity values from the service for things like reference data (e.g. business location) and try to find matches.

For example, if the user asks “How many engagements are there in NYC?” we need to disambiguate NYC. We do so by first providing the model with a list of field names and we have it guess what that literal might be referring to. In this case, most modern models comfortably correlate NYC with a field name like “BUSINESS_LOCATION”. Once it does that, we can pull the distinct values our service has stored for BUSINESS_LOCATION and ask the model, in a dedicated pass, to only pick the value that clearly matches “NYC”. In this case “New York”.

How many engagements are there in NYC?

LLM: extract literals

NYC → BUSINESS_LOCATION

API: distinctValues(BUSINESS_LOCATION, search: NYC)

→ no hits

API: distinctValues(BUSINESS_LOCATION)

→ London, New York, San Francisco, Berlin, …

Code: fuzzy match

→ nothing ≥ 0.85

LLM: pick from these 8 real values

→ New York (checked against the list)

Planner gets NYC → New York

Of course, this is an oversimplification. In a real production scenarion you might not want to provide all values to the model to help it disambiguate as it might be incredibly inefficient or even forbidden by your data policy: certain values might include sensitive customer data you don’t want to share with the LLM API provider.

Depending on your specific architecture, you might decide how to handle grounding failures. Sometimes, the request won’t contain any specific value (e.g. “How many workers are there?”) and in those cases the grounding node should just accept that and move on. Other times, you might want to call out that the intent wasn’t fully understood. For example, if the grounding fails to disambiguate “NYC”, you might want to report that back to the user and ask for clarification, instead of attempting a query filter you know won’t work.

Plan & Validate

We already had the plan step in the first implementation. This time plan comes after scope and ground. Instead of just having the GraphQL SDL and the natural language intent as input, it’ll also have the literals we have been able to resolve in the grounding step.

What’s new from the first naïve implementation is also that we allow for repairs. We are trying to mitigate for incorrect GraphQL queries and the non-deterministic nature of the output of the node by implementing a simple feedback loop: if the GraphQL query we produce fails schema validation then we try again.

This is possible because the GraphQL specification helps us statically validate the query: we can detect it’s incorrect and that it won’t even be executed before handing it over to the execution step. On top of that, we also evaluate the query depth and complexity and we reject if it exceeds a certain threshold. Additionally, we ensure that filters act on values resolved by the ground step and ask the user for clarification otherwise.

When we get a validation failure, we re-run the plan step including the error in the context we provide to the LLM to generate the query. I want to stress the importance of feeding the error back into the second attempt. A blanket retry without the error context has significantly fewer chances of producing a valid result. Finally, we cap the maximum repairs to 2, in this case, and report the failure back.

Into Plan

valid

invalid, repairs < 2

still invalid after 2 repairs

Question

'How many active engagements are there at Acme Corp?'

Resolved literals (from Ground)

ORG_NAME: 'Acme Corp' → 'Acme Corporation'

Context

full SDL in the system prompt

Plan (LLM)

Out of Plan

query: engagements(filter: $filter, first: 1) { pageInfo { totalCount } }

variables: orgName eq 'Acme Corporation', status ACTIVE

rationale: filter by resolved org name and ACTIVE status, count via totalCount

Validate (code)

schema · depth · pagination · grounded literals

Execute

Repair input: same question + literals

+ validation errors + previous query

e.g. ungrounded literal: ORG_NAME 'Acme Corp'

Templated refusal

Here’s a side-by-side comparison of the result of the query from part 1 now that the agent is capable of detecting when it cannot fulfill the request:

Side-by-side comparison of the naive GraphQL query agent and the grounded, validating agent answering the same question

What’s still stochastic

As you might have noticed, this time we are putting more effort into constraining the result of each node. It’s an exercise in reducing indeterminism which remains the intrinsic aspect of any LLM-based architecture.

In particular, the major stochastic areas of the architecture are:

  • NOT_TRACKED is left to the LLM to produce: we do so because doing just string matching
    would significantly degrade the quality of the feature since we would lose the semantic
    aspect of the correlation.
  • providing a list of capabilities to match against is a double-edged sword. Depending on
    the model, trying to match against a list might stretch semantic matching a bit too far,
    still leading to fabricated responses.
  • grounding entities into the available values is definitely a step forward that enhances the overall flexibility of how the intent specified in natural language is interpreted. That being said, we still delegate the resolution of the correct entity names to the LLM and that remains stochastic.
  • we are still asking the LLM to provide a rationale for the query it “decides” to generate: this is just a hint to the end user and needs to be carefully taken into account when evaluating the final result of the request.
  • the final answer might still mention facts inaccurately or fabricate new ones entirely:
    we will see in the next part what mitigation strategies we can put in place to address those

Before and after

Since we are looking for an improvement over the initial iteration of the architecture we can run an eval (a very small and specific “eval” in this case) of how the two architectures perform with a set of pre-defined questions and see what changes.

In this case I asked the 5 questions reported 10 times and measured the results.

The following evals have been run with the same model (Sonnet 5) for the planner node, which is the same between the two architectures. What changes are the nodes around plan which contribute to the different results.

QuestionPart onePart two
Which department has the most workers?answered, grouping by ORG_NAME (10/10)refused: “department” isn’t tracked (10/10)
How many workers are in the Engineering department?empty result (10/10)refused: “Engineering department” isn’t tracked (10/10)
How many workers report to Alice Johnson?answered “no reporting data” (5/10), empty result (5/10)refused: “report to Alice Johnson” isn’t tracked (10/10)
How many active engagements are there at Acme Corp?27, via contains "Acme Corp" (9/10), empty result (1/10)27, via grounded eq "Acme Corporation" (10/10)
How many workers are based in the Munich office?empty result (10/10)refused: no business location “Munich” (10/10)
Mean cost per turn$0.0058$0.0038

The semantic substitution issue came out completely mitigated in all 10 runs with this new architecture. The last question is also interesting to look at. In the current dataset there is no “Munich” business location. The previous architecture simply ends up picking a contains filter and ends up producing an empty list. The current architecture, instead, makes the agent behaviour more explicit: there is no active business location “Munich” and the agent reports that back to the user as a refusal.

You can see cost came out lower too. That is because our architecture now includes refusals which help the agent stop when it’s unnecessary or wrong to continue. This reduces the LLM invocations and therefore results in lower cost.

What’s still open

In this post we explored a couple of techniques to mitigate structural and referential failures. We saw how we can’t completely eliminate the stochastic nature of the architecture but what options we have to contain it. Compared to the first iteration, we are now much more confident that issues like semantic substitution won’t be as likely as in the first example. Additionally, the agent is much more honest about what it can and cannot answer.

In the next part of this series we will focus on the answer node which we have left pretty much alone so far. We will see how we can prevent the model from making arithmetic mistakes or overreaching in the conclusions it reports back to the user. As we did here, we will use a simple eval approach to validate if our approach improves the results and to which degree.


If you enjoyed this article, follow me on X or LinkedIn where I share my journey and my articles. Stay tuned for part three. And if you missed it, check out the first post in this series.

6 October 2026