What We Learned Using Jev to Choose Agents in an AI App
What we learned about context, confidence, and choosing the right agents with Jev.
Ask an AI app to “write a LinkedIn post,” and a small question appears before the writing even begins: who should help?
A Content agent seems obvious. A Research agent could find supporting facts. A Viral agent might sharpen the hook. A Scheduler could publish the result.
All reasonable possibilities. But the user might simply want a draft before lunch.
As AI apps gain more agents and tools, deciding which ones to use becomes a problem of its own. We explored using Jev to make that decision. Along the way, we learned that a useful integration needs clear questions, relevant context, and a reliable way to handle uncertainty.
The interesting work started after the first API call.
Give selection its own job
We placed Jev before the model that writes the plan.
Jev chooses the participating agents. The planning model then writes their tasks. The application checks whether those tasks are possible and properly ordered.
This division fits TypeSafe’s introduction, which describes Jev as a model for structured decisions. Its answers give code something concrete to work with: a choice, a probability distribution, and a confidence score.
Separating selection from writing also makes problems easier to understand. A poor team choice and a poor task plan need different fixes. If both happen inside one long generated response, it can be difficult to tell where things went wrong.
Ask who is needed, rather than who could help
A request can need several agents, so asking the model to choose just one would be too restrictive.
We used one “include or skip” Choice question per agent, sent together in a single request. The answers combine into a team. TypeSafe’s API reference documents how these questions and their structured answers work.
For a scheduling agent, the question is roughly:
Does this agent need to perform work to satisfy the request, given its supported actions and the available context?
That wording turned out to matter.
Ask whether an agent could help, and almost everyone has a case. Research could add background. Video could turn the idea into a script. Analytics could measure the eventual result.
Ask whether the agent contributes necessary work, and the decision becomes more closely tied to the user’s intent.
This also makes the agent descriptions important. A name such as “Growth agent” leaves plenty of room for interpretation. A description that lists specific actions gives the model clearer boundaries.
Context changes the team
Consider two requests:
Draft a LinkedIn post announcing our new product.
Draft and schedule a LinkedIn post announcing our new product next Tuesday.
The second request adds a clear responsibility. But the app still needs to know whether a publishing connection is available.
Now add another detail:
Use the launch notes I supplied.
That may reduce the need for additional research.
We found it useful to send the prompt alongside a compact summary of the facts that affect planning: the brand, desired outcome, channels, deadline, available connections, budget constraints, relevant past lessons, and research already gathered.
The aim is to give Jev enough context to judge whether an agent is needed. The application still checks whether the required accounts, actions, and data are available.
An agent can be relevant and unavailable at the same time. Keeping those two ideas separate makes both planning and the interface clearer.
It also keeps context focused. Credentials and unrelated customer records don’t help answer “Which agents are needed?” and don’t belong in the selection request.
Follow a request all the way through
Imagine a user asks:
Write and schedule a LinkedIn post for our product launch next Tuesday. Use the supplied launch notes.
An expected team might look like this:
Agent | Expected decision | Why |
|---|---|---|
Content | Include | The post needs to be written. |
Scheduler | Include | The user asked for scheduled publishing. |
Research | Skip | The supplied notes may contain enough information. |
Video | Skip | The request doesn’t include video work. |
Retention | Skip | There’s no customer win-back task. |
These are illustrative expectations, rather than measured model results.
After Jev answers, the application validates the response. If the team is accepted, the planner writes tasks within that team’s supported actions.
Then come the practical checks. The post must exist before it can be scheduled. The publishing channel must be connected. Any required approval must happen before execution.
A sensible team is a starting point. The final plan still has to make sense.
Uncertainty needs a destination
We started with a confidence threshold of 0.7. Every answer must meet that threshold before we accept the whole team, including answers that skip an agent.
That last detail creates a real tradeoff.
Suppose Jev confidently includes Content but is unsure whether Viral should participate. Under our policy, the entire selection goes to backup. We don’t accept part of the result and guess the rest.
In two early live requests, an uncertain choice triggered that backup path. In one, the Content decision cleared the threshold while the Viral decision was nearly split.
Those examples showed that the fallback worked. They also showed how cautious an “every answer must pass” policy can be.
We treat the threshold as a starting point to evaluate. TypeSafe’s confidence guidance explains that confidence comes from the answer’s probability distribution. A score of 0.7 should not be read as a guarantee of correctness.
A lower threshold might allow more selections through. Whether that improves the experience needs evidence from a broader set of requests.
We also use backup selection for missing configuration, timeouts, invalid responses, or an empty team. A bounded request time prevents selection from holding up planning indefinitely.
The interface makes this visible with a simple message:
Backup agent selection used.
Recording the reason helps us distinguish model uncertainty from an integration problem.
Make the rest of the app respect the choice
This was one of the most useful lessons from the implementation.
An application may already have helpful rules that add tasks, fill gaps, or repair a plan. Those rules can accidentally bring excluded agents back into the team.
We had to carry the selection through the entire planning process. The planner receives the chosen agents and their supported actions, and later task additions respect those boundaries.
Checks for capabilities, approvals, and prerequisites still apply. If the selected team cannot produce a usable plan, the application takes the visible backup path.
We also kept writing failure separate from selection failure. If Jev chooses a valid team but the planning model fails, a backup writer can still work with that team. There’s no need to discard a valid decision simply because a later component had trouble.
Using research and requesting research are different decisions
Research already gathered can help select agents and write a plan. That doesn’t automatically mean a Research agent needs another task.
If the available evidence is sufficient, the app can use it and skip further research.
This distinction is easy to miss in an interface. Showing “Research used in planning” separately from “Research selected for additional work” gives users a more accurate picture of what happened.
The same idea applies elsewhere: having an output available and needing someone to create a new output are different situations.
Test what happens after the answer
Testing taught us to look beyond whether the model returned a plausible team.
We covered simple requests, requests needing several agents, and cases where an agent lacked a required connection or data source. We also checked uncertain answers, incomplete responses, service failures, and missing task prerequisites.
Some of the most valuable checks came after selection. Could the planner quietly add an excluded agent? Could a publishing task appear before its content was ready? Would a backup writer respect an already accepted team?
Following the whole path helped catch problems that a successful API response would miss. Running those checks alongside the application’s existing checks helped us assess whether the integration was ready to ship.
It also clarified the limits of testing. Automated tests can verify that the application follows its rules and handles failures as intended. Assessing selection quality requires realistic requests with expected teams, followed by comparison with the model’s actual choices.
Both kinds of evidence matter.
The pattern could extend to tools
Although we used Jev to select agents, the same approach could help choose tools.
Describe each tool’s purpose and requirements. Ask whether it is needed for the request. Validate the answer before execution.
A search tool may be needed for current information. A publishing tool needs an appropriate connection and permission. A data lookup needs a clear reason to retrieve that information.
We haven’t evaluated this integration as a general tool selector, but the design offers a useful starting point.
What we learned is that structured selection works best when the application gives the model a precise decision to make and takes responsibility for what follows.
The next time someone asks for a LinkedIn draft, the app should assemble a useful team without turning a small writing task into a company-wide project.
For more detail, see the TypeSafe introduction, API reference, and Jev with coding agents.