At the end of the previous post, I wrote that my custom routing code could probably be replaced by a built-in LangChain4j pattern soon that Mario Fusco and I had discussed after my initial blog post. On 2 October LangChain4j merged the resulting DecisionRouterPlanner, which selects and runs sub-agents with a DecisionModel.
Time for me to try the obvious follow-up: remove the custom planner and run the same car-rental assistant with LangChain4j’s new DecisionRouterPlanner. While I was at it, I also moved the project to the new DecisionModel API, another feature recently added to LangChain4j. I also found out about yet another Jev alternative called “Kev”, which, based on my name, I of course had to try out as a third decision backend alongside Jev and Laya.
With this refactor, the router is now smaller and can call several specialists in parallel. At the same time, the application still decides what probability is good enough, what to do when no specialist qualifies, and how to check the final answer.
I also upgraded the application to the brand new Quarkus 4 and Quarkus LangChain4j 2.0 Beta. The new versions use LangChain4j 1.21, which contains both the DecisionModel API and the new agentic routing pattern. Since both Quarkus dependencies are beta releases, this is still fairly new territory, but this POC was the perfect opportunity to give them a first spin!
Why I needed a custom planner
The first version of the poc app I wrote asked one Choice question about which specialist should handle a given request. It worked for requests about one topic, but it could only return one answer, while a user asking about both the weather and the price of a convertible needed multiple agents to provide a comprehensive answer.
In the second version, the router asked one independent yes/no question per specialist, so several answers could pass the threshold. That meant a request could fan out to multiple agents, e.g. to get reservation, weather and cost agent responses at the same time.
This updated implementation solved the routing problem, but I had to build a custom planner implementation myself. My “JevRouter” created the questions and applied the threshold and a JevRoutingPlanner called the selected agents one by one, and a small collector copied their answers into a map before the merge agent could use them.
With the custom planner API I was using, though, I couldn’t reliably continue to the merge step after several parallel branches finished. So even though it would’ve been nice to run the agents in parallel, I had to resort to a sequential execution, even when two specialists had no reason to wait for each other.

Routing with DecisionRouterPlanner
LangChain4j now has two pieces that fit this use case. DecisionModel is the common interface for typed questions and answers. And the DecisionRouterPlanner uses that interface to route an agent call.
The planner has two modes. Without a threshold, it asks a Choice question and calls the most likely agent. With a threshold, it asks one independent yes/no question for every sub-agent and calls every agent whose probability reaches that threshold. The selected agents then run in parallel, and the planner returns their outputs as a map keyed by agent name.
The threshold mode matches the policy that performed best in my earlier Jev evaluation.
The complete router in the updated project now looks like this:
public interface SpecialistRouter {
@PlannerAgent(
name = "specialistRouter",
description = "Selects the specialists needed to answer a car-rental customer request",
outputKey = "specialistReplies",
subAgents = {
ReservationAgent.class,
WeatherAgent.class,
CostAgent.class,
GeneralAgent.class
})
Map<String, String> route(String request);
@PlannerSupplier
static Planner planner() {
ActiveDecisionClient decisionModel =
Arc.container().select(ActiveDecisionClient.class).get();
double threshold = ConfigProvider.getConfig()
.getOptionalValue("routing.fan-out-threshold", Double.class)
.orElse(0.5);
return new DecisionRouterPlanner(decisionModel, threshold);
}
}
The agent names and descriptions supply the routing criteria. For a request such as “Book an SUV and tell me the price with full insurance,” the planner asks whether the reservation agent, the weather agent, the cost agent or the general agent should handle it.
If reservation returns 0.9 and cost returns 0.9, both agents run. Their replies come back in a map like this:
{
"reservation": "...",
"cost": "..."
}
The subsequent merge agent receives this map and then writes one customer-facing answer.
Running the specialists in parallel
Fan-out starts several independent branches from one input. In this application, one customer request can start several specialist agents.
Consider “How much extra is a convertible, and will the weather in Nice be good enough next weekend?” The cost agent can answer the first part while the weather agent answers the second. There is no data dependency between them, so running cost first and weather second only added latency in my previous version.

The DecisionRouterPlanner owns the full fan-out operation. It starts the selected agents in parallel, waits for their results and returns the completed map. The surrounding sequence can then invoke the merge agent as a normal next step:
@SequenceAgent(
name = "tripAdvisor",
description = "Routes a customer request to specialists and combines their replies",
outputKey = "reply",
subAgents = {
SpecialistRouter.class,
GeneralFallback.class,
MergeAgent.class
})
String planTrip(String request);
I still define the sequence around the router, but I no longer manage each parallel branch or copy its reply into shared state.
Handling an empty routing result
The built-in planner returns an empty map when no agent reaches the threshold, so I made a separate conditional step that calls the general agent which is basically a catch-all for requests that aren’t relevant for any of the specialized agents:
@ActivationCondition(GeneralAgent.class)
static boolean needsGeneral(
Map<String, String> specialistReplies) {
return specialistReplies == null || specialistReplies.isEmpty();
}
The logs record the two decisions separately. An empty result means the decision model did not activate any route, after which the application chose general as its fallback. I can therefore tell whether the model selected general or the workflow supplied it.
Jev, Kev and Laya behind one interface
The provider-neutral DecisionModel API also removed the backend-specific routing code I had initially created. The framework creates a named model for each System One-compatible endpoint, and the application selects one with decision.backend.
Quarkus LangChain4j now also has a quarkus-langchain4j-typesafe extension. In the previous version I had written my own REST clients and request mapping for Jev and Laya. The extension now handles the /v1/systemone protocol and creates a named DecisionModel bean for each configured endpoint.
I configure one named model for Jev, one for Kev and one for Laya. ActiveDecisionClient selects the requested bean and keeps the existing stub fallback. The extension owns the HTTP calls and protocol mapping, while the application still owns backend selection and fallback behavior. The application also keeps an adapter for the guardrails because they still use the older PoC client API.

The configuration names each backend and its endpoint:
decision.backend=jev
quarkus.langchain4j.jev.decision-model.provider=typesafe
quarkus.langchain4j.typesafe.jev.base-url=https://api.typesafe.ai
quarkus.langchain4j.kev.decision-model.provider=typesafe
quarkus.langchain4j.typesafe.kev.base-url=http://localhost:8009
quarkus.langchain4j.laya.decision-model.provider=typesafe
quarkus.langchain4j.typesafe.laya.base-url=http://localhost:8100
Jev is the hosted service. Kev is another open model that can expose the same /v1/systemone endpoint locally. Laya still runs through the small compatibility sidecar from the earlier posts. The router just receives a generic DecisionModel, so the code stays the same no matter which implementation we choose.
Trying Kev, because why not
I started with Kev’s 0.8B model because its base model is only about 1.8 GB. The first few requests looked promising because it selected the agents I expected, but the full 28-request evaluation changed that picture.
At the default 0.5 threshold, Kev only routed 9 of the 28 requests correctly. Its best result among the thresholds I tested was 18/28 at 0.7: 7/8 clear single-intent requests, 4/7 indirectly worded requests, 3/8 multi-intent requests and 4/5 general requests. The median decision call took 49 ms on my laptop.
The model often activated an extra agent. It sent a simple Lisbon rain question to both reservation and weather, and sent an SUV price question to reservation and cost. Raising the threshold filtered out many of these extra routes, but it also dropped agents that multi-intent requests needed. None of the thresholds from 0.3 to 0.7 handled both cases well.
I tested the guardrails separately with 24 labeled request-and-reply pairs. Kev scored 11/12 on the relevance check, which only asks whether a specialist answered at least one part of the request. It scored 9/12 on the stricter completeness check. The misses included two partial answers that it accepted as complete. It also rejected one complete two-part answer with a score of 0.390, just below the 0.4 retry threshold.
This evaluation also showed what the guardrails do not check. Kev accepted a direct answer about opening hours even though the test supplied no evidence that those hours were correct. Relevance and completeness can catch missing or unrelated answers, but factual accuracy needs a separate check against the source data.
The 0.8B model is fast and easy to run locally, but these results do not support using it as a direct replacement for Jev in this application. It may fit a narrower router that picks one agent, while reliable fan-out and completeness checks need a stronger model or more work on the questions and thresholds. Kev recommends its 4B model as the normal starting point, so perhaps I should try that version and see if it works better.
Recording the routing decision
The Kev evaluation did expose a problem: when the application handles a request incorrectly, the final reply does not tell me whether the router chose the wrong agents or one of those agents produced a bad answer. I need to see which routes the decision model activated.
LangChain4j provides a DecisionModelListener for this. My listener receives the same four yes/no answers as DecisionRouterPlanner and records every route that meets the configured threshold. It also records whether the result used one route, fan-out or the general fallback. The listener only records the decision, so the planner remains responsible for choosing and running the agents.
The REST response includes this routing information with the final reply. For the reservation-and-price example, it looks like this:
{
"request": "Book an SUV for Saturday and tell me the price with full insurance.",
"reply": "...",
"route": "reservation",
"routes": ["reservation", "cost"],
"routingMode": "fan-out",
"backend": "jev",
"model": "jev-latest",
"live": true
}
Keeping application policy outside the planner
The three classes that implemented the old routing workflow contained 386 lines. The new router and fallback contain 68. The difference in line count is less important than the ownership change, because LangChain4j now maintains the parallel fan-out and output collection.
Several pieces correctly remain in the application:
- The
0.5activation threshold is configuration because it needs to be measured against this application’s requests. - The general fallback is part of the customer experience.
- The specialist guardrail checks whether each partial answer is relevant, while the merge guardrail checks whether the final reply covers the whole request.
- The routing evaluation remains the way to compare Jev, Kev and Laya on the same labeled cases.
I ran the normal test suite against the deterministic decision stub. All 25 tests passed, including a fan-out case that activates reservation and cost and a fallback case where every probability is below the threshold. I also started the application and checked both paths through the REST API.
I still need to rerun the same evaluation against Jev and Laya. The built-in planner follows the same independent yes/no policy as the previous version, but its exact question wording comes from the agent names and descriptions. Since wording can move model probabilities, I would measure again before claiming that the old thresholds transfer unchanged.
The first two versions were partly about whether a decision model could route agents at all. I now get that routing behavior from LangChain4j and can spend the time on the questions that remain specific to this application: which model works best, where to set the threshold, and whether the final answer is actually useful.
Links
- Source code: https://github.com/kdubois/quarkus-langchain4j-jev-laya
- Part 1: https://www.kevindubois.com/2026/09/23/routing-agents-with-jev-and-laya-adding-system-one-decisions-to-quarkus-langchain4j/
- Part 2: https://www.kevindubois.com/2026/10/01/routing-with-jev-and-langchain4j-agentic-part-2-the-real-api-multi-intent-requests-and-fan-out/
- LangChain4j
DecisionModelAPI: https://github.com/langchain4j/langchain4j/pull/6469 - LangChain4j
DecisionRouterPlanner: https://github.com/langchain4j/langchain4j/pull/6561 - Jev / TypeSafe System One docs: https://docs.typesafe.ai
- Kev: https://github.com/jaredpalmer/kev
- Laya: https://github.com/NandhaKishorM/laya