Judgment calls as one line of Python.
“Is this comment spam?” “Which team owns this ticket?” “How urgent is it?” gut asks small, fast, cheap models instead of a frontier LLM: TypeSafe's Jev first, and any small model after it. Every yes/no answer is YES, NO, or UNSURE — and UNSURE is where a person takes over.
pip install gutfeel
import gut
gut.configure(backend=gut.JevBackend()) # or just set TYPESAFE_API_KEY
if gut.likely(comment, "is spam"):
hide(comment)
Type to wake the model: 96 MB, once, then cached. Shown with ask_human=True so you can see the band.
Between a regex and a frontier model, there is a lot of room.
Most judgment calls in a codebase are small: is this spam, does this mention a refund, is this urgent. A regex is too blunt for them. A frontier LLM is far more than they need. gut lives in the middle.
regex
- speed
- Microseconds.
- cost
- Free.
- parsing
- None needed.
- brittleness
- Breaks on the first rephrasing. Reads characters, not meaning.
small models
- speed
- About 0.1 s a question for the NLI model on a laptop CPU.
- cost
- Free when local; cheap per call with Jev.
- parsing
- None. You get YES / NO / UNSURE, an Enum member, or a rung.
- brittleness
- Good at concrete claims. Weak at taste. Says UNSURE rather than guess.
frontier LLM
- speed
- Seconds, over the network.
- cost
- Per token, and it adds up at volume.
- parsing
- Free text or JSON you parse and hope about.
- brittleness
- Understands nearly anything. Overkill for “is this spam”.
Three questions cover almost every judgment call.
Is it? Which one? How much? Each is a function, each returns a typed answer, and none of them asks you to parse anything.
- Team.BILLING
- Team.PLATFORM
- Team.MOBILE
You never write a threshold. You say how careful to be.
Three words set the posture: lean, stakes and ask_human. gut works out where the boundaries go. Drag the needle, change the words, and watch the hatched band move.
| stakes | lean none | lean "yes" | lean "no" |
|---|---|---|---|
| low | 0.40–0.60 | 0.20–0.40 | 0.60–0.80 |
| medium | 0.25–0.75 | 0.125–0.625 | 0.375–0.875 |
| high | 0.10–0.90 | 0.05–0.85 | 0.15–0.95 |
The real model, running on your machine.
This is gut's ZeroShotBackend model, loaded into this page. No key, no server. Pick a preset or write your own.
- model
- deberta-v3-xsmall-zeroshot
- size
- 70M parameters, 8-bit
- download
- 87 MB + 8.7 MB tokenizer
- runs on
- WASM, in this tab
- billing—
- a bug in the app—
- a feature request—
- account access—
This is the 70M model running on your machine. Jev, and larger models, are sharper; the code does not change.
One line of code. Six places to run it.
gut is built around TypeSafe's Jev, a model that answers typed questions directly, with probabilities. It is also model-agnostic: the same line runs on a laptop CPU, an open decision model on Ollaya, a small local LLM or any OpenAI-compatible server.
- JevBackendgutfeel[jev]TypeSafe's Jev. Answers typed questions directly with probabilities. Set
TYPESAFE_API_KEYand go, or reach it through OpenRouter withJevBackend.openrouter().JevBackend.ollaya()asks an open decision model on your own machine instead. - ZeroShotBackendgutfeel[local]A 70M NLI model on a laptop CPU, about 0.1 s a question. Free, offline, and hard to talk round.
- TransformersBackendgutfeel[local]A small local LLM such as Qwen3-0.6B, for when you want a language model without a server.
- OpenAICompatibleBackendgutfeelAnything that speaks the OpenAI API: Ollama, vLLM, llama.cpp, or OpenAI itself.
- CascadegutfeelCheap model first. Only the answers it is unsure of go on to the next.
- FakeBackendgutfeelFor tests. No model, answers you script.
recipe_or_memoir.py, where the NLI model settled 6 of 8 judgments and Qwen3-0.6B took the other 2. Here Jev takes them.gut.configure(backend=gut.Cascade(
gut.ZeroShotBackend(), # free and local: settles the obvious
gut.JevBackend(), # sees only what the first could not
))
Many subjects, one call. And async twins.
gut.each() batches a question over a list. Every function has an async twin: alikely, aclassify, arate.
spam = gut.each(comments).likely("is spam")
# many subjects, one call
decision = await gut.alikely(email, "...")
# alikely, aclassify, arate
Evidence, from the examples folder.
Real output from scripts in the repository. Printed as it came out, including the miss.
Your coding agent can learn it too.
The big model thinks; the small one decides, fast. From a shell, gut filter judges a thousand files, commits or log lines in one command, so an agent never has to read them all. The skill teaches it when to reach for that, and how to write gut in code.
# teach your coding agent the skill
npx skills add Kungie/gut --skill gut
# then it can judge at scale, with a budget
git ls-files | gut filter "retries failed requests" \
--read-files --max-cost 0.50
- Start with no arguments.A plain
gut.likelyfirst; add posture when the results ask for it. - Never invent a threshold.Say lean, stakes or ask_human. gut places the boundaries.
- One claim per question.“Is spam and is rude” is two questions.
- Do arithmetic in Python.Ask the model for the judgment, not the sum.
- For untrusted text, prefer ZeroShotBackend.It was moved by none of six injection attempts.
Let the big model delegate the small calls.
gutfeel-mcp hands any MCP client, such as Claude Code, Claude Desktop or Cursor, gut's judgments as tools, so the cheap calls go to a small model instead of a big one.
likelyYES, NO or UNSURE on one claim, with its probability.classifyOne of the options you give it.rateA rung on the scale you give it.eachOne question over up to a thousand texts.
claude mcp add gut --env TYPESAFE_API_KEY=your-key -- uvx --from "gutfeel[mcp]" gutfeel-mcp
{
"mcpServers": {
"gut": {
"command": "uvx",
"args": ["--from", "gutfeel[mcp]", "gutfeel-mcp"],
"env": { "TYPESAFE_API_KEY": "your-key" }
}
}
}
Contents.
Ten chapters, read in order if you are new. Every code sample in them is run by the test suite.
Small models are small.
“Lists ingredients”, “asks for a refund”, “asks to reschedule a meeting”. Things a careful reader could point to in the text.
Judging the quality of writing, or rating on fine scales, is beyond a 70M model. Say so in code with ask_human=True.
Every number on this page came from one set of examples. Try your own questions on your own data before you trust it.