Judgment calls as one line of Python.

“Is this comment spam?” “Which team owns this ticket?” “How urgent is it?” gut asks small, fast, cheap models instead of a frontier LLM: TypeSafe's Jev first, and any small model after it. Every yes/no answer is YES, NO, or UNSURE — and UNSURE is where a person takes over.

pip install gutfeel

Jev first, any small model after it · every code sample in the docs is run by the test suite.

moderation.pylisting 1
import gut
gut.configure(backend=gut.JevBackend())  # or just set TYPESAFE_API_KEY
 
if gut.likely(comment, "is spam"):
    hide(comment)
UNSURE
p = ?
0.25.5.751
?

Type to wake the model: 96 MB, once, then cached. Shown with ask_human=True so you can see the band.

§1The gap

Between a regex and a frontier model, there is a lot of room.

Most judgment calls in a codebase are small: is this spam, does this mention a refund, is this urgent. A regex is too blunt for them. A frontier LLM is far more than they need. gut lives in the middle.

regex

speed
Microseconds.
cost
Free.
parsing
None needed.
brittleness
Breaks on the first rephrasing. Reads characters, not meaning.

small models

speed
About 0.1 s a question for the NLI model on a laptop CPU.
cost
Free when local; cheap per call with Jev.
parsing
None. You get YES / NO / UNSURE, an Enum member, or a rung.
brittleness
Good at concrete claims. Weak at taste. Says UNSURE rather than guess.

frontier LLM

speed
Seconds, over the network.
cost
Per token, and it adds up at volume.
parsing
Free text or JSON you parse and hope about.
brittleness
Understands nearly anything. Overkill for “is this spam”.
§2Three questions

Three questions cover almost every judgment call.

Is it? Which one? How much? Each is a function, each returns a typed answer, and none of them asks you to parse anything.

specimen Ayes / no
likely
Is it?
gut.likely(ticket, "is a bug report")
returns one of
YESNOUNSURE
specimen Bwhich one
classify
Which one?
gut.classify(ticket, Team)
returns a member of your Enum
  • Team.BILLING
  • Team.PLATFORM
  • Team.MOBILE
specimen Chow much
rate
How much?
gut.rate(ticket, ["can wait", "this week", "right now"])
returns a rung of your scale
can waitthis weekright now
§3The band

You never write a threshold. You say how careful to be.

Three words set the posture: lean, stakes and ask_human. gut works out where the boundaries go. Drag the needle, change the words, and watch the hatched band move.

lean
stakes
ask_human
UNSURE
p = 0.62
0.25.5.751
← NOYES →
UNSURE
gut.likely(email, "the customer threatens to cancel", ask_human=True) → gut.UNSURE
stakeslean nonelean "yes"lean "no"
low0.40–0.600.20–0.400.60–0.80
medium0.25–0.750.125–0.6250.375–0.875
high0.10–0.900.05–0.850.15–0.95
§4Playground

The real model, running on your machine.

This is gut's ZeroShotBackend model, loaded into this page. No key, no server. Pick a preset or write your own.

The model is not loaded yet.
model
deberta-v3-xsmall-zeroshot
size
70M parameters, 8-bit
download
87 MB + 8.7 MB tokenizer
runs on
WASM, in this tab
Loads when you scroll here or type in the hero, then stays in your browser's cache. Nothing you type leaves this browser.
presets
hypothesis: “This text lists ingredients.”
lean
stakes
ask_human
UNSURE
p = ?
0.25.5.751
?
gut.likely(text, "lists ingredients", ask_human=True)

This is the 70M model running on your machine. Jev, and larger models, are sharper; the code does not change.

§5Jev first, any model

One line of code. Six places to run it.

gut is built around TypeSafe's Jev, a model that answers typed questions directly, with probabilities. It is also model-agnostic: the same line runs on a laptop CPU, an open decision model on Ollaya, a small local LLM or any OpenAI-compatible server.

  1. JevBackendgutfeel[jev]TypeSafe's Jev. Answers typed questions directly with probabilities. Set TYPESAFE_API_KEY and go, or reach it through OpenRouter with JevBackend.openrouter(). JevBackend.ollaya() asks an open decision model on your own machine instead.
  2. ZeroShotBackendgutfeel[local]A 70M NLI model on a laptop CPU, about 0.1 s a question. Free, offline, and hard to talk round.
  3. TransformersBackendgutfeel[local]A small local LLM such as Qwen3-0.6B, for when you want a language model without a server.
  4. OpenAICompatibleBackendgutfeelAnything that speaks the OpenAI API: Ollama, vLLM, llama.cpp, or OpenAI itself.
  5. CascadegutfeelCheap model first. Only the answers it is unsure of go on to the next.
  6. FakeBackendgutfeelFor tests. No model, answers you script.
Cascade: ZeroShotBackend, then JevBackend Eight subjects enter the local NLI model. Six are settled there. The two it is unsure of flow on to Jev. 8 subjects ZeroShotBackend 70M NLI · CPU · local 01 UNSURE settled here: 3 YES, 3 NO JevBackend sees only the 2 6 settled 2 settled or UNSURE, if Jev is too
fig. 1Drawn with eight subjects, as in recipe_or_memoir.py, where the NLI model settled 6 of 8 judgments and Qwen3-0.6B took the other 2. Here Jev takes them.
cascade.pylisting 2
gut.configure(backend=gut.Cascade(
    gut.ZeroShotBackend(),   # free and local: settles the obvious
    gut.JevBackend(),        # sees only what the first could not
))
§6Many at once

Many subjects, one call. And async twins.

gut.each() batches a question over a list. Every function has an async twin: alikely, aclassify, arate.

batch.pylisting 3
spam = gut.each(comments).likely("is spam")
# many subjects, one call
handler.pylisting 4
decision = await gut.alikely(email, "...")
# alikely, aclassify, arate
gut.each()5.2 s
a for loop15.9 s
§7Exhibits

Evidence, from the examples folder.

Real output from scripts in the repository. Printed as it came out, including the miss.

$ python examples/recipe_or_memoir.py
Ingredients: 200 g flour, 2 eggs…→Straight to the recipe. A rare and beautiful thing.
Every autumn my grandmother's kitchen smelled of apples…→Recipe found, after a life story. Scroll on.
Let me tell you about the summer of 2009.→No recipe. Only a memoir.
exhibit AThe NLI model settled 6 of 8 judgments; Qwen3-0.6B took the other 2.
$ python examples/meeting_or_email.py
Choose the launch date (30 min)…→Go. There is something to decide, and deciding needs people.
FYI: the office wifi password is changing (30 min).→Email. Reply with 'Could you send this over instead?'
Sync (60 min). Let's sync.→Ask for an agenda. Nobody, including the model, can tell what this is for.
exhibit BThe third answer at work: when the invitation says nothing, UNSURE is the honest reply.
$ python examples/commit_roast.py
wip→Two words or fewer. Future you, reading git blame at 2 a.m., is not amused.
Fix off-by-one in pagination…→fix · A bug went to live on a farm upstate.
Bump httpx to 0.28…→feature (it was a chore; small models are small)
exhibit CCommit messages, classified and gently judged. One miss, left in on purpose.
§8For agents

Your coding agent can learn it too.

The big model thinks; the small one decides, fast. From a shell, gut filter judges a thousand files, commits or log lines in one command, so an agent never has to read them all. The skill teaches it when to reach for that, and how to write gut in code.

terminallisting 5
# teach your coding agent the skill
npx skills add Kungie/gut --skill gut
 
# then it can judge at scale, with a budget
git ls-files | gut filter "retries failed requests" \
  --read-files --max-cost 0.50
  1. Start with no arguments.A plain gut.likely first; add posture when the results ask for it.
  2. Never invent a threshold.Say lean, stakes or ask_human. gut places the boundaries.
  3. One claim per question.“Is spam and is rude” is two questions.
  4. Do arithmetic in Python.Ask the model for the judgment, not the sum.
  5. For untrusted text, prefer ZeroShotBackend.It was moved by none of six injection attempts.
§9MCP server

Let the big model delegate the small calls.

gutfeel-mcp hands any MCP client, such as Claude Code, Claude Desktop or Cursor, gut's judgments as tools, so the cheap calls go to a small model instead of a big one.

  • likelyYES, NO or UNSURE on one claim, with its probability.
  • classifyOne of the options you give it.
  • rateA rung on the scale you give it.
  • eachOne question over up to a thousand texts.
terminal · Claude Codelisting 6
claude mcp add gut --env TYPESAFE_API_KEY=your-key -- uvx --from "gutfeel[mcp]" gutfeel-mcp
json · Claude Desktop, Cursor, any MCP clientlisting 7
{
  "mcpServers": {
    "gut": {
      "command": "uvx",
      "args": ["--from", "gutfeel[mcp]", "gutfeel-mcp"],
      "env": { "TYPESAFE_API_KEY": "your-key" }
    }
  }
}
Read the MCP chapter →
§10Docs

Contents.

Ten chapters, read in order if you are new. Every code sample in them is run by the test suite.

  1. Getting startedread →
  2. Backendsread →
  3. Knowing when it doesn't knowread →
  4. Asking everything at onceread →
  5. Asyncread →
  6. Exact costsread →
  7. Caching and observabilityread →
  8. Command lineread →
  9. MCP serverread →
  10. Honest limitationsread →
§11Honest limitations

Small models are small.

Good at concrete claims.

“Lists ingredients”, “asks for a refund”, “asks to reschedule a meeting”. Things a careful reader could point to in the text.

Weak at taste.

Judging the quality of writing, or rating on fine scales, is beyond a 70M model. Say so in code with ask_human=True.

Measure on your data.

Every number on this page came from one set of examples. Try your own questions on your own data before you trust it.