Decision checks inside your coding agents, set up and supported by deemwar.
jevx puts small, local yes/no checks inside Claude Code, Codex and your own agents: guardrails before a command runs, checks on claims an agent makes, and memory of what your team already decided. It runs on your machines, on CPU. deemwar installs it, wires it into your agents, trains it on your data and supports it.
What we measured
Our own tests, on held-out sets, with a small local model next to hand-written rules. Numbers, not promises; the limits are below.
Prompt injection caught
42 of 60 injected instructions caught by the local model, against 20 of 60 for a keyword rule. False alarms: 1 of 56 clean inputs (the rule: 0).
Dangerous commands
Destructive shell commands in formats our rules had never seen: the model caught 16 of 20, the rules 1. Zero false alarms on 20 safe look-alikes.
Claim checks with memory
A 0.4B local model checking claims against our docs: AUC 0.86 with retrieved notes, 0.30 without, on 68 held-out claims.
- Rules first, the model for the long tail. On familiar patterns the model alone adds false alarms, so we never run it alone.
- Secrets stay rules-first: for secret detection the model gave no significant gain over rules.
- The same person wrote parts of the test sets and the rules; we use novel formats to limit that bias, but your data is the real test. That is what a pilot is for.
What deemwar does
jevx is free and open source. We charge for people's time: installation, support and training.
Installation
On developer machines and in CI, for Claude Code, Codex and your own agents. Pinned versions; the checks run on your machines.
Guardrails in your agents
Command checks before execution, prompt-injection checks on what agents read, claim checks on what they report: wired into your hooks and pipelines.
Recipes and training on your data
Custom checks for your decisions, and a model tuned on your own command history, tickets or docs, measured on a held-out set before it ships.
Support
Upgrades, tuning when a check misfires, and a person to ask when an agent does something surprising.
How an engagement works
Installation, support and training, scoped in writing before anything starts. Rate on request.
A 20-minute call
Which agents you run, what worries you (commands, injected instructions, wrong claims, lost context), and where the checks would sit.
A scoped pilot
jevx installed on one team or one pipeline, one or two checks tuned on your data, and the numbers measured on your own held-out examples.
A retainer, if it earns one
Rollout to more teams, new recipes as your agents change, upgrades and support. Cancel any month.
Talk to us
Tell us what your agents do and what you want checked. A person replies, usually within a working day. Or email [email protected].