I’m very forgiving of an agent that’s solely speaking. If it will get a draft or abstract flawed, I crash out at my display and ask once more, and since nothing outdoors the chat window has modified, a retry is all it prices me.
However the temper modifications as soon as the agent will get a instrument. A flawed reply can now imply a despatched e-mail or a moved cost, and the same old repair, which is placing a second mannequin in entrance to verify the primary, begins to really feel like hiring an intern to oversee an intern.
One of many easiest guardrail instances I wrote down was additionally the one which bothered me most.
A buyer will get charged twice for a $12 buy and asks an AI agent for a refund. The agent picks the proper instrument and the proper buyer, then prepares this:
Nothing in that decision seems damaged. The shopper ID is legitimate, amount_cents is an integer, and the refund instrument’s schema accepts it.
The 2 duplicate fees on the account, ch_1 and ch_2, had been 1,200 cents every, so the quantity is off by an element of 100, which is what changing {dollars} to cents twice seems like.
I’m not actually anxious concerning the dramatic failure the place a mannequin utterly loses the plot.
The failures that hassle me probably the most are the extraordinary ones, the place the motion seems cheap, and the error solely turns into apparent after one thing actual has occurred.
My first response to this one was: simple. Put a tough $500 refund restrict in entrance of the instrument and transfer on.
Then I modified the unhealthy quantity from $1,200 to $120. That sits underneath the cap, so the restrict I had simply reached for by no means fires, and the shopper nonetheless will get ten instances what they authorised.
That small change is what pulled me into TypeSafe AI’s Jev mannequin.
TypeSafe launched Jev on September 15 as the primary of what it calls System One Fashions, a category of fashions constructed to make quick, structured choices that software program can use immediately.
It’s nonetheless in early entry, for what that’s value. As an alternative of producing one other paragraph, Jev takes software state and a bounded query, then returns a typed probabilistic reply.
TypeSafe’s launch publish frames this as a unique job from a traditional chat mannequin: much less era, extra decision-making. Guardrails for LLM inputs and outputs are on its checklist of supposed makes use of, which is adjoining to what I would like right here.
Jev has really been out for a number of weeks now, which in AI time makes it virtually classic, so sure, I’m a bit of late to this. A part of that’s all the way down to a benchmark I attempted to get working and by no means fairly managed, which I’ll get to in a second.
I had deliberate a small benchmark with dozens of artificial instrument calls, however I by no means bought a clear run by means of the gateway I had entry to, and I definitely didn’t need to move off half-working experiments as exact numbers.
What’s left is the half I discovered extra fascinating anyway: the place a mannequin like Jev ought to sit in an agent system, and what it shouldn’t be trusted to resolve.
The bug is typically semantic, not structural
Schema validation already handles a helpful class of failures. If send_email() expects a recipient checklist and an attachment, I can be certain that each fields exist. If issue_refund() expects an integer variety of cents, I can reject a foul worth earlier than it even will get close to the cost service.
It doesn’t catch this:
The consumer solely requested to ship the bill to Alice. Each addresses will be actual, the attachment can exist, and the schema passes with out grievance.
All the pieces checks out besides whether or not that is what the consumer really requested for.
A JSON schema can’t clear up that, and I additionally don’t need one other big immediate whose job is to clarify, in 600 tokens, why a refund is perhaps suspicious.
At execution time, the applying principally wants a choice it will possibly route on, and that’s the place Jev begins to look much less like one other mannequin and extra like a guardrail element.
My first sketch gave Jev an excessive amount of energy
My first model was embarrassingly clear: the agent proposes an motion, Jev says enable, overview, or block, and that’s it. It appeared good in a diagram, and actually I disliked it nearly instantly.
If Jev decides whether or not to authorize a refund, I’ve simply moved a permissions drawback into one other probabilistic mannequin, which is not a lot of a security structure.
So I flipped the order. Laborious guidelines go first, and Jev solely sees the messy instances that stay.
If refunds above $500 all the time want a human, I don’t want a mannequin’s opinion on them.
The identical goes for guidelines like “manufacturing backups can’t be deleted autonomously” or “this agent can’t entry payroll information.” These belong in code or permissions.

A left-to-right flowchart in three colours. On the left, an agent proposes a instrument name, which first meets a inexperienced field labeled “Laborious guidelines in plain code,” masking issues like refunds over $500 and guarded information. If a rule journeys, an arrow goes as much as an amber field labeled “Rule tripped,” which sends the decision straight to an individual with no mannequin concerned. If the decision passes, it goes to a blue field labeled “Jev,” which reads the request and the proposed name and returns enable, overview or block with a confidence. Three arrows go away Jev. The primary goes to an amber “Human overview” field, for low confidence or for cash, deletion and exterior sends. The second goes to a purple “Blocked” field, for calls that battle with the request. The third goes to a inexperienced “Execute” field, for a assured enable on a low-risk instrument. A legend on the backside marks inexperienced as deterministic code, blue as Jev (probabilistic), and amber as an individual.
The semantic layer comes after that. A tough integration with the official Python SDK might seem like this, and I’d deal with it as a sketch somewhat than examined code:
The 0.90 is a beginning assumption, not a quantity I’d ship as a result of it appeared good in an article. What issues is the division of duty.
The applying owns the exhausting boundary, Jev handles the fuzzy judgment inside it, and a low-confidence determination falls again to an individual as an alternative of pretending uncertainty is autonomy.
For refunds, I’d nonetheless deal with even a assured enable as a suggestion at first, which is the place the tool-specific guidelines additional down are available.
One caveat alone sketch. TypeSafe’s docs advocate small, single-purpose questions mixed in code over one broad judgment, and a three-way Alternative that folds the entire coverage into the state is nearer to the broad form.
A tighter model would ask separate sure/no questions, akin to whether or not the quantity matches what the shopper authorised and whether or not the shopper matches the request, and let extraordinary code flip these solutions into enable, overview, or block. I stored the one query right here as a result of it’s simpler to learn.
That ordering isn’t my invention. TypeSafe has a group playground with a tool-router instance constructed the identical manner: a plain key phrase rule blocks dangerous requests earlier than any mannequin is known as, and something delicate nonetheless wants express approval.
It’s a mock that routes between graph nodes somewhat than judging instrument arguments, however the order is the purpose.
The clearest line I discovered on this comes from the group kedi-typesafe LangChain integration, which says a constructive Jev evaluation ought to by no means exchange your personal instrument approval or coverage checks. That was in all probability probably the most helpful factor I learn whereas working by means of this.
The boring edge instances are those I care about
A guardrail that blocks “ship our personal API key to an unknown e-mail deal with” is beneficial, nevertheless it doesn’t inform me a lot. I care concerning the instances that look cheap for the primary two seconds.
Take this request:
Let the suppliers know the Q3 invoices are prepared.
The agent prepares one e-mail to 214 exterior contacts. The instrument is correct and the motion broadly matches the request, however I’d not let it fireplace mechanically.
The blast radius modified the choice, which is why I’d keep away from one common secure=True query for each instrument. A documentation search and a mass exterior e-mail are usually not the identical sort of danger, even when each are legitimate actions.
One other one:
Delete the exported CSV after confirming the add succeeded.
The agent factors delete_file() on the right CSV, however nothing within the state reveals the add ever succeeded. The goal is okay. The lacking prerequisite is the issue.
And yet another:
Ship the pricing sheet to our authorised associate.
The associate e-mail is right, and the attachment is:
A recipient allow-list won’t prevent there, and neither will checking the file extension. The guardrail wants sufficient context to see that this explicit file doesn’t belong on this motion.
That’s the sort of determination I’d give Jev: does this proposed motion nonetheless make sense subsequent to what the consumer really requested for?
Why not simply use one other LLM?
You’ll be able to, and I don’t suppose Jev makes that sample out of date. A powerful LLM can examine a proposed motion, cause concerning the coverage, and return a structured determination.
If that infrastructure already exists and the latency is suitable, I’d not rewrite a working security layer simply because a brand new mannequin launched.
The narrower mannequin is interesting for a sensible cause. The applying doesn’t want a mini essay each time an agent desires to learn a file. It wants enable, overview or block, plus sufficient likelihood info to resolve whether or not to belief the route.
TypeSafe’s present API exposes three determination primitives: Alternative, Noul (a sure/no likelihood) and Rating. The official Python SDK returns typed views for them somewhat than making you parse generated prose. The SDK quickstart is refreshingly small.
There may be additionally a sensible techniques argument.
The agent already depends on a generative mannequin to plan and choose a instrument, so placing a second massive mannequin in entrance of each execution means one other immediate to take care of, one other latency hop, and one other place for output dealing with to go flawed. Jev simply does much less, and for this job which may really be a bonus.
I practically made the arrogance threshold look smarter than it’s
At one level my instance had one clear quantity:
Then I pictured the identical threshold guarding each search_docs() and issue_refund() and deleted it. If a documentation search is flawed, the agent can recuperate. If a refund is flawed, cash strikes. If a mass e-mail is flawed, the recall button is usually ornamental.
I’d begin with tool-specific guidelines and preserve them conservative. That is pseudocode, not a whole integration:
Then I’d log what Jev needed to do subsequent to what the human ultimately selected, and solely loosen up something after sufficient actual visitors. “Make the agent extra autonomous” isn’t mechanically an enchancment right here.
I’d somewhat be aggravated by a number of further overview requests within the first month than discover out what the error price means with an actual buyer hooked up.
Typed output isn’t the identical factor as being proper
TypeSafe talks about Jev avoiding hallucinations as a result of the output house is outlined upfront, and I’d phrase that declare fastidiously. And structurally I get the argument. I imply, if the one choices are enable, overview and block, the mannequin can’t invent a fourth route known as refund_and_email_everyone, and this system is aware of the doable outputs earlier than inference.
TypeSafe is upfront about what that assure covers: its launch publish says the 0% type-error determine in its charts isn’t an empirical measurement, as a result of schema matching is assured by building.
That solely covers the form of the output. The mannequin can nonetheless choose enable when the proper reply is block, and that could be a unhealthy determination somewhat than a damaged output.
Typed output removes one sort of failure. It doesn’t take away mannequin error, lacking context, weak insurance policies, or unhealthy software design. The mannequin will get a vote, not the keys to the constructing.
···
Remaining ideas and takeaways
I began by asking whether or not Jev might make brokers safer with out placing one other LLM in entrance of each instrument name. I feel the reply is sure, if “safer” means one thing pretty particular.
Jev seems helpful within the hole between “the agent desires to do that” and “the applying is about to let it occur.” That could be a slim position, and I see that as a power.
I’d nonetheless preserve exhausting limits on cash, deletion, secrets and techniques, and permissions, with an individual concerned wherever a mistake is pricey.
What I’d hand to Jev is the half that’s exhausting to jot down as an if assertion.
Does the motion nonetheless match the request?
Did the agent quietly widen the scope?
Is a situation lacking?
The $12 refund is mundane, which is why I prefer it. If a small determination layer provides the applying yet another probability to catch that earlier than cash strikes, a file disappears, or an e-mail reaches 214 folks, that’s sufficient for me to take the concept significantly.
I don’t want Jev to be one other mind within the agent. I’d somewhat have or not it’s a really choosy gate.

