
Last week I wrote about running Claude Code on local models with vLLM and my Nvidia Pair router. That setup works great, but I left out an embarrassing little detail: for days, my local AI kept interrupting itself to ask permission, and the permission-seeker kept falling asleep on the job. This post is about fixing that with a Claude Code plugin (a "mod") that hands the safety decisions to a dedicated decision model.
Here is the thing about auto mode. Claude Code has a safety classifier that looks at every tool call which your permission rules don't already cover and decides whether to just run it. And that classifier, in my setup, was a model call going to the exact same machine that was already busy writing the code. So whenever GLM was in the middle of a long reasoning stretch at xhigh effort, the classifier's request would sit in the queue, time out, and I'd get this:
GLM-5.3-Flash-EXL3[1m] is temporarily unavailable (timed out), so auto mode cannot determine the safety of Write right now. Wait a moment and then try this action again.
Read-only operations skip the classifier, so browsing kept working — but every file write, every network call, every anything-mutating got gated. During one particularly busy afternoon I watched it block the same mkdir fifteen times in a row while I sat there like a traffic cop waving at an empty intersection. The AI fixing the AI, blocked by the AI. Not the good kind of irony.
But wait — it gets better. The classifier is just a model call. Which means it doesn't have to be that model, or even that machine.
Enter Clef: Cloudflare's 27B decision model, fine-tuned from Qwen3.8-27B, that turns a state and a schema of typed questions into decisions. No autoregressive decode — every option of every question scored in a single forward pass. That's exactly what a permission classifier is: one question, asked constantly, needing an answer in milliseconds, not a paragraph. It runs happily on a second box via Ollama's /v1/systemone endpoint. Their own docs use tool call moderation as the example:
{
"model": "clef",
"state": "send_email(to=\"all-customers\", subject=\"FINAL NOTICE: account will be suspended today\")",
"questions": {
"harm": { "type": "noul", "instructions": "Could this tool call cause harm?" }
}
}
which answers, in one pass:
{ "model": "clef", "answers": { "harm": { "type": "noul", "noul": 0.951 } } }
(That's a classic "FINAL NOTICE" mass email scoring 95% harm. Nice.)
Now, how do you swap a classifier inside Claude Code? Turns out the plugin ("mod") system is exactly the right wrench. Mods are folders of TypeScript function hooks that hot-reload into a running session, and the one that matters here is tool.check — it fires after the permission rules and before the mode settles an ask, and its answer is the last word. The built-in classifier only gets consulted when the final verdict is still "ask". So a tool.check hook that turns asks into confident allow/deny answers doesn't fight the built-in classifier. It simply never lets it speak.
So I had Claude build it — clef-guard. On every uncovered ask in auto mode it sends the tool call (plus a little conversation context) to clef with two questions: a choice verdict (allow / deny / uncertain) and a noul harm check, then applies a simple policy: allow needs P(allow) ≥ 0.7 and P(harm) < 0.5; deny needs P(deny) ≥ 0.7 or P(harm) ≥ 0.5; anything else falls through to the stock pipeline, same as if the mod weren't there. Every failure — clef unreachable, garbage reply, below threshold — also falls through. The mod can never fail into allow-all; its worst case is stock behavior. It ships 14 tests, a /clef command that prints the last verdict with its probabilities, and an enforce-off shadow mode where clef annotates the status line but never decides. There is a build quirk worth documenting too: the engine's scanner only lets $ (its API handle) flow through same-file top-level functions, which is why the HTTP transport lives in the entry file.
Here is what /clef reports in a live session:
clef-guard (enforce: true)
endpoint: http://192.168.86.86:11434/v1/systemone
transport: curl, model: clef
thresholds: allow>=0.7, deny>=0.7, harm>=0.5
mode: auto (as of 2026-10-05T02:38:22.249Z)
last: 2026-10-05T02:39:03.114Z Bash -> allow in 284ms — clef: allow (p=0.94, harm=0.01, conf=0.92)
And to my delight, it validated and passed the full test suite (14 tests: decision policy, unreachable-clef fallback, rule pass-through) on the first push.
While I was in there, I also formalized the other trick from the setup: one model wearing four hats. Claude Code picks different model aliases for main work, subagents and the small fast model, and each alias can carry its own reasoning effort. The trick is that Claude Code keys effort settings on the resolved model ID, so the proxy serves suffixed model IDs off one checkpoint:
"env": {
"ANTHROPIC_BASE_URL": "http://127.0.0.1:8888",
"ANTHROPIC_AUTH_TOKEN": "any-dummy-token",
"ANTHROPIC_DEFAULT_FABLE_MODEL": "GLM-5.3-Flash-EXL3",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "GLM-5.3-Flash-EXL3",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "GLM-5.3-Flash-EXL3-sonnet",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "GLM-5.3-Flash-EXL3-haiku",
"CLAUDE_CODE_AUTO_MODE_SERVER": "0",
"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1"
},
"modelSettings": {
"GLM-5.3-Flash-EXL3": { "effortLevel": "xhigh" },
"GLM-5.3-Flash-EXL3-sonnet": { "effortLevel": "medium" },
"GLM-5.3-Flash-EXL3-haiku": { "effortLevel": "low" }
},
Same weights, three effort levels. Measured on the same prompt: 173 output tokens at low, 253 at medium, 399 at xhigh. A soft knob, not a hard cap — but it nudges hard, and it means the heavy thinking happens only where I asked for it.
The end result: the safety check is now the fastest call in the house. One forward pass on a machine whose only job is deciding, while the GPU thinks about code instead of thinking about whether it's allowed to think about code. Still feels odd that the bottleneck in local AI turned out to be the permission slip — but that's local AI for you: you become the datacenter, and the datacenter argues with itself.