In the first post I replaced the LLM supervisor in a Quarkus LangChain4j
car-rental assistant with a System One decision model. A custom Planner asked one Choice question
(“which specialist?”) and called that agent, and an OutputGuardrail asked a yes/no (Noul) question
to check each reply. I could only test the hosted Jev path against a stub or an alternative solution (Laya) at the time,
because TypeSafe had paused signups.
I now managed to get an API key, so I figured I should test it all out again with the real API. Fortunately, it worked on the first try. At least for routing a clear-cut request to a single agent. The downside is that I noticed ambiguous requests did not get handled so well since the routing I created would only forward the request to one sub agent. To fix that I changed how the router asks its questions, and I added a fan-out step to the planner. Along the way I ran into a few things about guardrails and the agentic framework that I didn’t expect. And I also ran the same evaluation against Laya, where single-intent routing held up and the fix for multi-intent requests… did not.
Connecting to the real API
With the new API key in place I tested the same scenario and.. it worked! I did write against the Typesafe docs so that was the expectation all along, but we all know how things go.
I then sent the example requests from the first post again, and Jev picked the right specialist every time, with a probability of 0.98 or higher.
The specialist agents also have a reply guardrail, which checks each answer before it goes back to the customer. After an agent writes its reply, the guardrail sends Jev the customer’s question together with that reply and asks a yes/no question: “Does the drafted reply directly address the customer’s request?” Jev answers with a probability between 0 and 1. Below 0.4, the guardrail rejects the reply and the agent has to write a new one. Above 0.6, the reply goes through. In between, it goes through but gets flagged for review.
Here is what that looked like for two replies to “Will it rain in Lisbon next Tuesday?”. The first is the weather agent’s real answer. I wrote the second by hand to check that the guardrail rejects a bad reply.
| Reply | Jev’s score | Guardrail |
|---|---|---|
| “Based on the current weather forecasts, it is expected to rain in Lisbon next Tuesday. If you plan to drive, be prepared for potentially wet and slippery road conditions…” | 0.96 | passes |
| “Our SUVs start at $65 per day, and you can add a second driver for $12 per day.” | 0.01 | rejects, and the agent tries again |
All the real replies scored between 0.96 and 0.98.
While checking the logs I also found two bugs of my own that the stub had hidden. Only the weather and reservation agents had the guardrail, even though the post said every reply was checked. The general agent’s prompt never included the customer’s message, so it answered greetings fine and everything else blind. The Laya runs in the first post didn’t catch this because “Hi there!” was the only general request I sent :D.
Ambiguous scenarios
Each example request from the first post was about one thing: the weather in Lisbon, the price of an SUV, or a booking. But questions often ask about several things in one message, while the router I built can only send a message to one specialist. So I tried a few requests that mix topics, which I’ll call multi-intent requests from here on, to see what the router does with them.
The first was “Is it worth upgrading to a convertible for a weekend in Nice if it might rain?” A good answer needs three specialists: reservation for the upgrade, cost for whether it’s worth the money, and weather for the rain. Jev’s Choice question (“which specialist should handle this request?”) gave weather 0.59, general 0.25 and oddly cost was only 0.02, so the router sent the request to the weather agent. The weather agent wrote a reasonable answer about rain, and the guardrail passed it at 0.96, since it did address the request.
A more explicit two-part question made it worse. For “How much extra is a convertible, and will the
weather in Nice be good enough next weekend?”, the Choice picked general at 0.44 with a confidence
of 0.26, and gave weather 0.04.
This behavior follows the definition of the “Choice”. The options are mutually exclusive and their probabilities add
up to one, so a request that needs two specialists either ends up split between options or lands
on the catch-all. In both ambiguous requests, the runner-up was general. Picking the top two
options would have added the general agent each time, but never the cost agent the question needed.
One yes/no question per specialist
Jev can however take several questions about the same input in one request. So besides the Choice, the router now asks one Noul per specialist:
public static Map<String, JevQuestion> questions() {
Map<String, JevQuestion> questions = new LinkedHashMap<>();
questions.put(ROUTE_QUESTION, JevQuestion.choice("Which specialist should handle this customer request?", ROUTE_CRITERIA));
NEEDS_TOPICS.forEach((route, topic) -> questions.put(NEEDS_PREFIX + route, JevQuestion.noul(
"Does answering this customer request fully require input about " + topic + "?")));
return questions;
}
Each Noul is judged independently, so several can be yes at once. For the two-part convertible question, Jev said yes to weather (0.91), cost (0.85) and reservation (0.69), in the same call that produced the unsure Choice. Four questions take the same 0.3 s as one.
My first policy kept the Choice as a fast path: if its confidence was 0.8 or higher, route on the Choice alone, and only use the Nouls otherwise. It seemed like the cautious option, but it was wrong, which I only found out by measuring.
Measuring it
I wrote 28 labeled requests in four groups. There are 8 clear single-intent requests, 7 with
indirect wording (“Will I need snow chains to drive to Chamonix in January?”), 8 that need two
specialists, and 5 general or off-topic ones. Where a label is debatable, the eval accepts an
alternative. “Can I bring my dog in the rental car?” can go to general or reservation.
The eval is an opt-in test in the project. It sends each request to Jev once, with the same questions the router asks, and then applies every policy to the stored answers, so comparing policies costs no extra calls:
./mvnw test -Dtest=RoutingEvalTest -Drouting.eval=true # needs TYPESAFE_API_KEY
| Policy | single (8) | indirect (7) | multi-intent (8) | general (5) | total |
|---|---|---|---|---|---|
| Choice only | 8 | 7 | 0 | 5 | 20/28 |
| Choice alone if confidence ≥ 0.8, else Nouls | 8 | 7 | 5 | 5 | 25/28 |
| Nouls first, Choice when no Noul passes | 7 | 6 | 8 | 5 | 26/28 |
The fast path missed three multi-intent requests, and in each of them the Choice was confident about a single specialist:
| Request | Choice | Nouls (reservation / weather / cost) |
|---|---|---|
| Book me a car for Saturday and tell me what it’ll cost with full insurance. | reservation, 0.97 | 0.84 / 0.06 / 0.95 |
| Should I pick up the car Friday or Saturday given the storm forecast? | reservation, 0.91 | 0.67 / 0.81 / 0.06 |
| Can I change my booking to a 4×4, and will I need one for the weather in the Highlands? | reservation, 0.99 | 0.92 / 0.77 / 0.30 |
Choice confidence says how sure the model is about the label, assuming there is exactly one label. It says nothing about whether the request has one intent. In the first post I suggested using low confidence to trigger a clarifying question, and on Laya the mixed requests did come back with confidence below 0.1. On Jev, some of them came back at 0.97. Since the Nouls arrive in the same response, the fast path saved nothing, and dropping it gained three requests.
The Nouls-first policy fails differently. Both of its misses are single-intent cost questions where the reservation Noul also passed, such as “Is the young-driver surcharge included in the quote you sent me?” (reservation 0.84, cost 0.91). The reply then includes a reservation paragraph nobody asked for. It’s slower than necessary, but the answer is still right, and I prefer that to a missing answer. The threshold mattered less than I expected. Values from 0.3 to 0.5 gave the same multi-intent results, and 0.5 scored best overall. At 0.5, 10 of the 28 requests fanned out.
The policy is a few lines, and it reads only probabilities and the Choice’s value:
List
<String> routes = needs.entrySet().stream()
.filter(e -> e.getValue() > policy.fanOutThreshold())
.sorted(Map.Entry.<String, Double>comparingByValue(Comparator.reverseOrder()))
.map(Map.Entry::getKey)
.toList();
String mode = switch (routes.size()) {
case 0 -> rawChoice == null ? MODE_FALLBACK : MODE_CHOICE;
case 1 -> MODE_NEEDS;
default -> MODE_FAN_OUT;
};
if (routes.isEmpty()) {
routes = List.of(normalize(rawChoice));
}
Greetings and off-topic questions end up in the case 0 branch, where no specialist is needed and
the Choice picks general. The Choice is still useful there. It’s just no longer the first thing
the router trusts.
Twenty-eight requests is a small set, and I wrote the labels myself. I could probably remove the two misses by rewording the reservation Noul, but tuning wording against the same set I measure with would only make the table look better. A larger set from real traffic is the next step.
The same eval on Laya
The first post ended with the idea that you could build on the open model and only pay for the hosted one if it does better on your data. With the eval in place I could actually check that, so I ran the same 28 requests against Laya, with the same checkpoint as in the first post:
./mvnw test -Dtest=RoutingEvalTest -Drouting.eval=true -Drouting.eval.backend=laya
| Policy | Jev multi-intent (8) | Jev total | Laya multi-intent (8) | Laya total |
|---|---|---|---|---|
| Choice only | 0 | 20/28 | 0 | 18/28 |
| Choice alone if confidence ≥ 0.8, else Nouls | 5 | 25/28 | 1 | 19/28 |
| Nouls first, Choice when no Noul passes | 8 | 26/28 | 1 | 19/28 |
Laya matched Jev on single-intent routing. Its Choice picked the right specialist for all 15 single and indirect requests, though with much lower confidence: never above 0.60, where Jev was at 0.88 or higher for most of them. It did send two off-topic questions, such as “Do you offer jobs for drivers?”, to the reservation agent.
The per-specialist questions are where it fell short. They only work when the yes/no answers clearly separate the specialists a request needs from the ones it doesn’t. Jev’s answers do, mostly above 0.7 for a needed specialist and below 0.3 for the others. Laya’s sit in a narrow band. The second topic of a two-part request scored as low as 0.19, while a specialist the request didn’t need at all scored up to 0.38 (reservation, for the jobs question). No threshold can tell those apart, and only one of the 28 requests fanned out.
The routing still behaved sensibly, because when no yes/no answer passes, the Choice decides. Laya ended up where single-Choice routing is and never worse than it. It was also faster, at under 0.1 s per call on my laptop against about 0.3 s for Jev over the network.
I didn’t tune the question wording or the threshold for Laya, because doing that against the same 28 requests would mostly fit the test set. The first post already found yes/no answers to be Laya’s weak spot, and this run confirms it for routing. For now, routing a request to one specialist works on both models, and fan-out only seems to work reliably on Jev.
Fan-out in the planner
Once the router can return several specialists, the planner has to call them and combine the results. The flow for “Book me a car for Saturday and tell me what it’ll cost” looks like this:

After each agent finishes, the framework calls the planner’s nextAction(), and the planner decides
what runs next from what just finished. The list of pending specialists lives in the request’s
AgenticScope, not in a field on the planner, so concurrent requests don’t share it:
@Override
public Action nextAction(PlanningContext context) {
AgentInvocation previous = context.previousAgentInvocation();
AgenticScope scope = context.agenticScope();
List
<String> routes = scope.readState("routes", List.of());
if (routes.size() <= 1 || MergeAgent.class.equals(previous.agentType())) {
return done(previous.output());
}
// Fan-out: after each specialist, collect its reply; then call the next specialist, or merge.
if (!isCollector(previous.agentType())) {
return call(agents.get(SpecialistRepliesCollector.class));
}
List
<String> pending = new ArrayList<>(scope.readState(PENDING_ROUTES, List.<String>of()));
if (pending.isEmpty()) {
return call(agents.get(MergeAgent.class));
}
String next = pending.remove(0);
scope.writeState(PENDING_ROUTES, pending);
scope.writeState(SpecialistRepliesCollector.CURRENT_ROUTE, next);
return call(pickSubagent(next));
}
Every specialist writes its answer to the same reply key in the scope, so each one would overwrite
the previous reply. A small collector step runs after each specialist and adds its reply, labelled
with the route ([cost], [reservation]), to a list that the merge agent reads. The merge agent is
one more LLM call. It
gets the original request and all the replies, and its prompt tells it to cover every part of the
request, drop repetition and add no facts of its own. Without it, the customer gets two separate
answers that each ask for the same booking details.
Routing to several specialists has an obvious cost. A single route took 2 to 4 s end to end, and the fan-outs I ran took about 6 to 12 s, because each specialist and the merge is another LLM call. The Jev part stays at one call either way.
One guardrail question didn’t fit both cases
Adding fan-out broke the guardrail. The question was “Does the drafted reply directly address the customer’s request?”, and two kinds of reply now failed it for the wrong reason. A friendly answer to “Hi there!” scored 0.38 and was retried until it ran out of retries, because a greeting has no request to address. A reply that covered only the weather half of a weather-and-price question scored 0.24, and in a fan-out that is exactly what each specialist is supposed to write.
I tried three wordings against live Jev on five test replies:
| Reply | “directly address the request” | “relevant, on-topic response” | “answers at least one thing the customer asked or said” |
|---|---|---|---|
| Friendly reply to a greeting | 0.36 | 0.72 | 0.55 |
| Off-topic reply | 0.01 | 0.01 | 0.02 |
| One specialist’s part of a fan-out | 0.24 | 0.39 | 0.98 |
| Good reply | 0.93 | 0.93 | 0.97 |
| Evasive non-answer | 0.03 | 0.14 | 0.08 |
No single wording handled every row. The third one never rejected a reply that should pass, and it still caught the off-topic and evasive ones. But it would also pass a merged answer that dropped half the question, which is the failure the original wording catches. So there are now two guardrails that share one implementation and differ only in the question. The specialists use the relevance question, and the merge agent uses the original completeness question. The greeting now scores in the uncertain band and passes with a flag for review.
Neither guardrail checked the actual facts. When I asked the general agent about opening hours at Lisbon airport, it said the desk was open 24 hours, which it made up, and the guardrail scored the reply 0.97 because it was on topic. A yes/no question about relevance can’t catch invented details.
Smaller things worth knowing
One Jev call hung for 30 s before the app gave up and let the offline stub judge the reply instead, so I lowered the timeout to 5 s. Calls normally take 0.3 s, so that leaves plenty of room. And when a guardrail ran out of retries, the API returned an HTML error page. It now returns a JSON error that says the reply was rejected.
What I’d do next
LangChain4j merged an experimental DecisionModel API
with a TypeSafe integration on 29 September. It wasn’t in a release yet when I wrote this, so I
haven’t switched to it. It covers most of what I wrote by hand, including the decision client and
the question types, and it can point at other servers that speak the same API, such as Laya. The PR
also says to rely on probabilities instead of confidence, because each provider computes
confidence differently.
With regards to the custom routing planner, after showing Mario Fusco from the LangChain4j team my implementation of this decision model pattern, he has been working on creating a built-in pattern in LangChain4j, so likely my custom solution can be replaced by an “official” LangChain4j solutions soon! It should also feature calling sub agents in parallel, something I struggled to do with my custom planner. The PR is already available: https://github.com/langchain4j/langchain4j/pull/6561
For Laya, the next thing to try is a checkpoint or question format suited to multi-label decisions, measured on a larger request set than the one I wrote by hand. The eval runs against any backend that speaks the protocol, so trying another model is a config change.
On Jev, the routing now answers every part of a two-part question, at the cost of an extra LLM call per specialist. I made most of these changes only after measuring on real requests, and running Laya taught me not to assume a policy tuned on one decision model carries over to another.
Links
- Source code: https://github.com/kdubois/quarkus-langchain4j-jev-laya, including the eval set and the Jev and Laya eval reports
- Part 1: Routing Agents with Jev and Laya
- LangChain4j
DecisionModelPR: https://github.com/langchain4j/langchain4j/pull/6469 - LangChain4j custom agentic patterns: https://docs.langchain4j.dev/tutorials/agents/#custom-agentic-patterns
- Jev / TypeSafe System One docs: https://docs.typesafe.ai
- Laya: https://github.com/NandhaKishorM/laya