angel291592/Intent-Router

★ 500⑂ 56

Intent compiler for AI agents — converges vague requests into typed IntentSpec contracts (probe, ask, or halt before routing), the input layer for routers and typed-decision models like Jev & Laya

About angel291592/Intent-Router

angel291592/Intent-Router is an open-source project on GitHub, mainly written in Python. Intent compiler for AI agents — converges vague requests into typed IntentSpec contracts (probe, ask, or halt before routing) It currently holds 500 stars and 56 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Today's Trending board.

GitHub Repository Details

Repository angel291592/Intent-Router · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

Intent-Router

English | 简体中文

An intent compiler for AI agents.

Turns a vague request into a typed IntentSpec: it looks up what it can, asks only what it can't, calls out a request that contradicts itself, admits two readings, or rests on a premise your own repo already disproves, and refuses to emit while the intent is still underspecified. When the work gets handed off and carried out, it checks the result against the spec it wrote — one constraint at a time, not a vibe check.

What that buys you:

engineering — the gap between the top 1% of prompt writers and everyone else stops deciding what you get back. Say it the way you'd say it to a capable teammate, "add caching to the user API", and the skill writes the briefing a senior engineer would have written first. You stop coaxing the model and start putting it to work. reads what your repo, ticket system or docs already answer, so the forty-six-question interview becomes the one question that genuinely needs your judgment — and in the delivery comparison, every run with the skill asks exactly that one question and 3 of 3 delivered caches state their failure policy, while none of the 5 bare deliveries do (report). .intent/.intent.yaml before the reply that carries it, whose probed fields carry evidence pointers and which records its own state: still waiting on a question, settled, or already carried out. A later session reads that file before probing anything — an answered question is never asked again, a settled contract is implemented as written — and the result of checking the delivered work against it is written back into the same file. So the next agent, session or teammate starts from what you actually decided instead of zero.

License: MIT Release CI Stars Works without installing anything

What "installed" changes, in the diff itself

Same weak request, same fixture, same permissions, one line apart — bare vs. with the skill (N=5 bare, N=3 skill). Excerpted verbatim from each run's own diff.patch, the part that decides what happens when the cache itself fails:

Bare — no failure path exists to read. A downed Redis takes the request down with it:

const hit = await getCache(key);
if (hit !== null) {
  return hit;
}
const value = await load();
if (value !== null) {
  await setCache(key, value);
}
return value;

With the skill — asked the one question no file could answer, then wrote the answer into the code:

async function readCache(key: string): Promise<{ hit: boolean; value: T | null }> {
  try {
    return { hit: true, value: await getCache(key) };
  } catch {
    return { hit: false, value: null };
  }
}

Every one of the 5 bare runs shipped the first shape; every one of the 3 skill runs shipped the second. Neither line was written by hand for this README — both come straight out of the runs that produced the numbers in Evals.

          "add caching to the user API"       "this customer is furious — sort it out"
                        │                                         │
                        ▼                                         ▼
 ┌────────────────────────────────────────────────────────────────────────────────────┐
 │  parse  →  resolve  →  typecheck  →  emit                    one engine, one spec  │
 ├────────────────────────────────────────────────────────────────────────────────────┤
 │  PROBE   deps, routes, git history, ADRs   │  order + ticket log, entitlements,    │
 │          — you are not interrupted         │  the policy in force                  │
 │                                            │                                       │
 │  ASK     serve stale, or serve uncached?   │  refund, or replace?                  │
 │          the one answer no file holds      │  the one that can't be undone         │
 └────────────────────────────────────────────────────────────────────────────────────┘
                        │                                         │
                        ▼                                         ▼
              IntentSpec  ──────►  your planner / agent / workflow / human

Same three passes. Same stopping predicate. Only the places it looks change.

What a real run looks like — one sentence in, it reads the repo itself and asks only what no file can answer. From the measured run of 2026-09-23 behind the numbers below: 8 of the 10 cases the suite held that day, 0 hallucinated evidence.

https://github.com/angel291592/Intent-Router/blob/HEAD/Intent-Router loads automatically and probes the repo: package.json, src/routes/users.ts, src/cache/redis.ts and more

https://github.com/angel291592/Intent-Router/blob/HEAD/The one question it asks — fail open or fail closed — with A/B options and a recommendation

https://github.com/angel291592/Intent-Router/blob/HEAD/The emitted IntentSpec: 6 unknowns found, 4 resolved by probe, every field carrying an evidence pointer

No install, no API key, no dependencies — it's a skill:

npx skills add angel291592/Intent-Router

Manual install, or an agent the installer doesn't know: Quick start.

---

Contents

---

How it compares

Intent-Router is the layer between a vague request and the work: look it up, ask, or act — then check what got built. Your harness's own plan mode already prefers looking things up over asking, and grill-me converges beautifully and states the right principle — but neither one checks a stated premise against the evidence, and neither one checks the finished work against what was decided. spec-kit keeps everything and makes you buy its whole workflow to get it. Kiro's specs come closest on paper — requirement-level conflict analysis, a correctness check after delivery — at the cost of a full three-file spec workflow and an IDE. Jev and Laya make the same bet this skill makes — typed output instead of prose — one layer further down: both turn an already-clear input into a typed, calibrated decision in a single pass, and neither one converges a vague request into that clear input to begin with. Intent-Router is that upstream layer as a single skill: no workflow to adopt, no files to keep in sync by hand, no IDE, no separate decision service to run.

Read the row gaps, not the checkmarks. Every cell below is sourced to the tool's own docs or repo — quoted cells are verbatim, and an unquoted ✅/⚠️/❌ still traces to a specific page.

| | Converges vague input | Doesn't ask what it can look up | Uses evidence against contradictions & bad premises | Never guesses past irreversible | Cross-session artifact | Checks delivery against intent | Fires automatically | |---|---|---|---|---|---|---|---| | Intent-Router | ✅ | ✅ PROBE | ✅ ambiguous / conflict / premise | ✅ HALT | ✅ .intent/*.yaml | ✅ Intent check | ✅ | | Your harness's native plan mode + question tool | ✅ asks | ⚠️ varies — Codex's plan template states the rule outright; Claude Code and Cursor explore but document none | ❌ not documented on any of the three | ⚠️ an approval gate, not a risk-based one | ✅ Claude Code, Cursor save a plan file; Codex's plan lives in the reply | ❌ not documented | ⚠️ Claude Code's tool is model-invoked; Cursor only suggests; Codex needs an explicit switch | | grill-me | ✅ rounds/frontier | ✅ "don't ask the user for anything you could look up yourself" | ❌ not documented | ✅ "the decisions are the user's" | ❌ stateless by design | ❌ not documented | ❌ manual (/grill-me) | | spec-kit /clarify | ✅ category scan, ≤5 questions | ⚠️ excludes questions "already answered", not questions the repo could answer | ⚠️ a separate command, /analyze, flags conflicting requirements | ⚠️ flags unresolved items as "Deferred", no reversibility test | ✅ writes into spec.md | ✅ a separate command, /converge | ❌ manual, one /speckit-* command at a time | | Kiro specs | ✅ "asks clarifying questions if needed" | ⚠️ "codebase exploration", no stated rule | ⚠️ analyze-requirements finds inconsistencies across the requirement set (IDE/CLI only) | not documented | ✅ three files per spec | ✅ property-based "Correctness" checks (IDE only) | ❌ manual (Shift+Tab / /plan) | | Jev (typed decisions) | ❌ needs clear input already | — | — | — | — | — | ✅ | | Laya (open-source System 1) | ❌ needs a formed state to classify | — | — | — | — | — | ✅ |

Jev's and Laya's rows are mostly —, not ❌: they are the System 1 decision layer, called once a request is already well-formed — a different job from converging one, so most of these columns do not apply to them. Jev is a closed API; Laya is the open-weights, Jev-wire-compatible alternative you can run locally. The IntentSpec this skill emits is the shape of input both of them assume already exists, which is why both show up again under Backends, as the planned optional L2.

What your harness already does — and what this adds

To be fair: a good harness natively prefers looking things up over asking, and offers a recommended default when it does ask. In the delivery comparison, all 5 bare runs read docs/adr/0007-no-inproc-cache.md on their own and avoided the trap it warns about — that habit is real, and it is not this skill's contribution. Nothing else on this page should be read as claiming it is.

What it adds is the part a conversation cannot hold:

context, open a new session, switch models, and they are gone. An IntentSpec is a file-level contract, so the next agent, session or teammate starts from it rather than zero. underspecified collapse into the same "I need more information". Here degraded is an operations signal and underspecified is your next step; merging them hides a backend outage behind what looks like a clarifying question. default, mention it, keep going". When an inferred value touches an irreversible boundary, the skill refuses to ROUTE — it HALTs and names the field instead. dependency or a removed file that contradicts what you asked for is worth one question before the work starts; liking a different approach better is never grounds to raise one. conversation, the skill reruns its own spec against what was actually delivered and reports one line per constraint — met, with a pointer into the code, or not met.

---

The problem

A compiler doesn't guess the address of an undefined symbol.
Your agent shouldn't guess your intent.

When you hand an agent "refactor this module", "look into our churn" or *"sort this customer out"*, exactly one of two things happens.

It guesses. You get 400 lines of confident work built on an assumption you never made. You read it, realize the premise is wrong, and throw it away. The agent was never wrong about how to do the job — it was wrong about what you meant, and it found that out only after spending your tokens.

Or it interviews you. This is better, and it's what /grill-me-style skills do well. But the bill is real: the grill-me docs call "forty-six questions across four rounds an ordinary session." And when it's over, the skill is explicitly stateless — *"it writes no files and leaves no workspace behind. The only thing it leaves is a sharper version of the idea, in your own head."*

So you pay twice. Once in questions, and again because nothing downstream can read the answers. The next agent, the next session, the next teammate — all start from zero.

Both failures share one root cause: nobody decided whether the missing information was worth asking a human for.

Half those forty-six questions already have answers — in your package.json, your router file, your ADRs. Or, in the next queue over, in the order record, the account's entitlements, the refund policy currently in force. A compiler resolving an undefined symbol doesn't interrupt you to ask where malloc lives; it goes and looks in the libraries it was given. That's the missing step, and it isn't a step about code.

---

How it works

Three passes, in compiler order.

1. Parse — request → candidate structure. The raw request becomes a draft IntentSpec: a normalized action, its objects, explicit constraints, and a list of unknown fields. Nothing is invented; anything the model had to fill in itself is tagged source: inferred so you can veto it in one glance.

2. Resolve — the part everyone skips. This is context engineering as a routing decision: every unknown is classified before anything reaches you:

| State | When | What happens | |---|---|---| | PROBE | The answer exists somewhere it can reach | It reads it. Code, deps, config, version history, tests, CI, docs — or a ticket log, an entitlement table, a policy document, your prior notes, an API. You are not interrupted. | | ASK | The answer only exists in a human's head — a preference, a trade-off, an irreversible boundary | One question, with two concrete options and a recommended default. |

The rule that makes this work, and the one you should hold it to:

If an objective answer exists and you have a way to reach it — PROBE. Never ASK.
ASK is reserved for answers that live in a person's head, or decisions that can't be walked back.

3. Typecheck & emit — a decidable stopping condition. Interviews that stop "when it feels done" run to forty-six questions and drift into a full context window. Intent-Router stops on a predicate instead:

sufficient  ⟺  every required field is filled
            ∧  no inferred field touches an irreversible boundary

Not sufficient, and out of ASK budget? It does not emit. It halts and names the field that is still open — the same way a compiler refuses to link rather than picking an address at random.

Halts carry a cause, and the two kinds are never merged:

problem.

Every clarify/route library I found collapses these into one FALLBACK. That's how you end up with a router that silently sends 100% of traffic to the default handler for a week because one backend was down, and nothing in your logs says so.

---

The four decision states

| State | Meaning | Emits | |---|---|---| | ROUTE | Sufficient, one clear target | IntentSpec + target | | PROBE | Missing info is lookup-able | (internal — loops back to resolve) | | ASK | Missing info needs human judgment | one question, two options, a default | | HALT | underspecified \| degraded | a named open field, or an ops alert |

Others state it as a principle — grill-me's grilling skill says *"finding facts is your job, never the user's."* Intent-Router makes it a state: every probe leaves an evidence pointer, every run reports resolved_by_probe / unknowns_found, and stopping is a predicate, not a feeling.

---

Two walkthroughs

Same engine, same artifact, two different worlds. The first is the domain the full-suite eval numbers belong to; the second shows what changes when there is no repository in sight — which is: the probe surfaces, and nothing else.

A. In a codebase — "add caching to the user API"

Parse finds four unknowns: which cache backend, which endpoints, what TTL policy, what happens when invalidation fails. Resolve classifies them before saying a word to you:

PROBE  cache backend      → package.json: ioredis@5; src/cache/redis.ts exists   ✓ resolved
PROBE  which endpoints    → src/routes/users.ts: 3 GET handlers                  ✓ resolved
PROBE  TTL policy         → src/cache/redis.ts:12 — repo convention is 300s      ✓ resolved
ASK    invalidation fails → not in the repo. It's a trade-off. Irreversible.

One question reaches you:

On invalidation failure, which way should it fail?
A (recommended) — serve uncached. Slower, always correct.
B — serve stale. Fast, can be wrong for up to 300s.
*Why you and not me: this is a product decision about whether your users may see stale data. It
isn't in your code, and it's expensive to reverse once clients depend on it.*

Emit: unknown: [], so it compiles. resolved_by_probe: 3, asked: 1.

And it caught something neither a guesser nor a griller would: ADR 0007 says in-process caching was already tried and reverted. That constraint is in the spec with a file pointer — not because you remembered it, but because probing is cheap and memory isn't.

B. In a support queue — "this customer is furious — sort it out"

Not a line of code involved. The same four categories of unknown appear, and most of them are still lookups:

PROBE  what actually happened → order #8821: delivery 12 days late, carrier marked lost  ✓
PROBE  what they're entitled to → Pro plan, 30-day guarantee — 9 days of it left         ✓
PROBE  what we already offered → ticket #4402: 10% credit, declined by the customer      ✓
ASK    refund or replacement  → policy permits both. Issuing one forecloses the other.
Refund or replacement?
A (recommended) — replace, expedited. Keeps the subscription; carrier claim covers cost.
B — full refund. Ends the dispute today, likely ends the account with it.
*Why you and not me: both are permitted, the customer's preference decides, and a refund
issued cannot be walked back.*

resolved_by_probe: 3, asked: 1 — and the agent never asked the customer to re-explain what happened, because the ticket already said. "Is the purchase inside the warranty window?" is a record, not an opinion; asking it is the bug this thing exists to remove.

Honesty note: walkthrough A is the domain covered by the eval suite and the
screenshots above. Walkthrough B's domain is measured on its own terms —
support-delegate-irreversible passed in the 2026-09-23 subset run —
while the full-suite numbers still belong to coding. See Where it works.

---

The artifact

One file, three consumers: a human reviews it, an agent executes it, an auditor replays it.

intent: add_caching
objects: [GET /api/users, GET /api/users/:id, GET /api/users/:id/prefs]
constraints:
  • source: explicit
text: don't change response shape
  • source: probed # found in package.json + src/cache/redis.ts
text: use the existing Redis client, not a new dependency evidence: src/cache/redis.ts:12
  • source: probed # found in version history + docs/adr/0007.md
text: in-process caching was tried and reverted in #412 — don't reintroduce evidence: docs/adr/0007-no-inproc-cache.md
  • source: asked
text: on invalidation failure, prefer correctness (serve uncached) over availability unknown: [] # empty ⟹ sufficient decision: state: ROUTE target: implement confidence: 0.88 resolution: unknowns_found: 4 resolved_by_probe: 3 # diagnostic, not a target asked: 1 trace: [...] # replayable

Outside code the fields don't change — only what an evidence pointer names. A file path becomes a record identifier or a document section (record:tickets/4402, doc:returns-policy#eu); the requirement that everything probed carries one does not relax.

The full field list is fixed by intentspec.schema.json (JSON Schema draft 2020-12), with a worked example of each state in schema/examples/.

Every spec is also saved to .intent/.intent.yaml before the reply that carries it — the next agent, session or teammate starts from that file, and it sits next to the diff it produced, so "what was this PR actually trying to do" has an answer that isn't archaeology. Only that one file is written, nothing is committed, and nothing is written at all when the request is fully specified and the skill stays silent; whether .intent/ belongs in version control is your project's decision. Meanwhile resolved_by_probe / unknowns_found still gives you something to hold the tool to.

---

Where it works

The three passes and the sufficiency predicate are domain-independent. What changes between domains is one thing only: where an objective answer can be found. Status is stated the same way this project states harness compatibility — measured, specified, or neither.

| Domain | PROBE reaches | The question worth asking | Status | |---|---|---|---| | Coding agents | repo, deps, version history, tests, CI, ADRs | irreversible technical trade-offs | ✅ measured — eval suite, 17 cases | | Support & service triage | ticket history, order and event logs, entitlements, the policy in force | refund vs. replace, when both are allowed and one forecloses the other | ✅ measured — 1 case (support-delegate-irreversible), subset run | | Research & analysis | prior notes, previous reports, the source allow-list, cached retrievals | depth vs. breadth, when the deliverable changes shape | 📋 specified | | Ops & data work | schema, dashboards, last run's output, deploy and incident history, retention policy | may a backfill rewrite historical rows | 📋 specified | | Multi-role assistants | the candidate registry, attachment metadata, conversation history, user tier and locale | which of two genuinely overlapping specialists | 📋 specified | | Your domain | whatever you've given it access to | — | write the probe surfaces, it compiles |

measured = at least one case from that domain passed in a published report in evals/reports/; each row states how many cases that is. specified = probe su

GitHub Stars & Activity

500Stars
56Forks
0Open issues
PythonLanguage

GitHub Popularity

GitHub stars500
Forks56
Open issues0
Primary languagePython
License-
Stars gained today0
Created-
Last pushed-

Trending History

Weekly boardrank #98 · ▲ 0 stars

Related AI Projects

1

Shubhamsaboo / awesome-llm-apps

Python★ 140,502⑂ 20,640▲ 139 stars
→
2

harry0703 / MoneyPrinterTurbo

Python★ 127,914⑂ 20,013▲ 598 stars
→
3

ComposioHQ / awesome-claude-skills

Python★ 76,303⑂ 8,898▲ 345 stars
→
4

debpalash / VoiceStudio

Python★ 51,263⑂ 5,697▲ 1,395 stars
→
5

PostHog / posthog

Python★ 40,080⑂ 3,464▲ 31 stars
→
6

VectifyAI / PageIndex

Python★ 38,407⑂ 3,326▲ 543 stars
→
7

sgl-project / sglang

Python★ 36,699⑂ 9,254▲ 39 stars
→
8

alirezarezvani / claude-skills

Python★ 27,153⑂ 3,837▲ 138 stars
→

More AI Rankings