redhat-et/ripwire
The ripgrep of AI context: a zero-dependency C++23 CLI + MCP server for coding agents. Find what you want without reading the repo, then check you built what you meant — blast radius, tests-to-run
About redhat-et/ripwire
redhat-et/ripwire is an open-source project on GitHub, mainly written in C++. The ripgrep of AI context: a zero-dependency C++23 CLI + MCP server for coding agents. It currently holds 2,408 stars and 151 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).
Project Overview
AI Homed tracks it on the Today's Trending board.
GitHub Repository Details
README
Rip'n Fast. Fewer Tokens. Better Code.
The ripgrep of AI context. A map before your agent reads the repo — and a check on what it writes.
Ranked, deterministic call graph: what to touch, what it breaks, which tests to run. On the edit: blast radius, tests that reach it, eleven quality kinds reporting only what got worse, forgotten co-changes, fields read and written, names that resolve more than one way.
Just want to use it? Install it with the one line below, then start each coding session by telling your agent to use it, for example: "Use ripwire on this repo." That is all most people need: the install also teaches your agent when to reach for each command.
Want every detail? The reference guide near the bottom covers install, commands, output format, exit codes and limits. You do not need it to get started.
Field report: what ripwire contributed to a large multi-agent coding engagement — written by Claude Fable 5.0, the frontier model orchestrating ~20 coding agents over two days on a ~1,500-file C++/Metal codebase. Click for the full report.
For a single developer, this tool is a good lookup accelerator. For an orchestrated fleet, it's load-bearing: it
halved the research spend, twice redirected tasks before wasted work, prevented at least one silent-divergence shipped
bug, and turned code-quality hygiene from a hope into a per-task mechanical gate. Whole-workflow ~2×; per-lookup
10–25×; and two moments where one call was worth more than the rest of the session's tooling combined.
> — the report's bottom line
Text version of the report
Field report: what ripwire contributed to a large multi-agent coding engagement
*Context, genericized: one orchestrating session directing ~20 sequential/parallel coding agents over a ~1,500-file C++/Metal codebase across two days — a deep architecture audit, then a 14-task feature wave (new subsystems, measurement infrastructure, a search-archive migration), ending in a verified all-on release flip. Every agent was instructed to lead with ripwire for orientation and to close with its quality gates.*
The headline numbers
- Roughly half the total token spend of the audit/research phase, saved. This is the operator's whole-phase
- Two tasks had their direction changed by a single call. One task was chartered to activate a steering behavior
--callers query returned zero production callers, redirecting the
task to the real gap (the data provider that behavior needed had never been wired anywhere). Another task
refactored a widely-shared computation; --edit-check flagged 5 of 6 call sites as incompatible — sites a text
search had missed — that would otherwise have silently diverged from the canonical path the day the feature was
enabled, surfacing months later as an unexplainable visual bug.
- A dozen-plus real code-quality defects fixed, not waived, across ~15 implementing agents — all caught by
--quality-delta at each agent's "I think I'm done" moment: a 480-token duplicated routine, a fourth private copy of
a shared RNG utility, a hand-duplicated cost function, a pair of near-identical functions with a flipped sign
(correct at one boundary, quietly wrong at the other), test-fixture duplication across sibling suites, and functions
that had quietly absorbed a second job. None of these would have failed a test; all of them are the sediment that
rots a codebase under high-velocity multi-agent development. The gate made removing them routine instead of heroic.
Where the value concentrated
1. Orientation and recall (the token win). Ranked-signature task orientation and section-granular document recall
meant agents started from answers, not file dumps. Recall over planning docs was the orchestrator's single
most-used verb — the audit's entire framing came from two calls.
2. The contract checker (the shipped-bug preventions). Beyond the 5-of-6 catch above, --edit-check gave
near-free proof-of-contract on every public-symbol change across the wave — the kind of verification that
otherwise simply doesn't happen at agent speed.
3. **The quality gate as a protocol. The measurable effect wasn't any single catch — it was that fifteen different
agents, none sharing context, all converged on the same "fix or justify with a written reason" discipline, with an
acknowledgments ledger that survived across tasks. Unattended orchestration usually leaks quality; here the
leak-check was mechanized.
4. Persistent notes. Agents left gotcha notes pinned to symbols mid-wave (a caller-wiring law, an arming-rule
invariant), which later agents' lookups surfaced automatically — cheap institutional memory between contexts
that never met.
The honest boundary
The engagement's deepest findings did not come from the tool.** A structural capacity ceiling, a mis-derived
constant, a floating-point re-association drift of 1 ulp, and a seed-keying bug were all found by *bespoke
measurement the agents built* — baseline worktree diffs over full data fills, multi-arm attribution sweeps,
funnels. The tool is a floor for honesty and orientation, not a substitute for measurement design.
- Known false-positive mode, handled by its own documentation: the contract checker's arity heuristic
- For broad, common-word conceptual queries, plain grep-and-read still occasionally won — consistent with the tool's
Bottom line
For a single developer, this tool is a good lookup accelerator. For an orchestrated fleet, it's load-bearing: it halved the research spend, twice redirected tasks before wasted work, prevented at least one silent-divergence shipped bug, and turned code-quality hygiene from a hope into a per-task mechanical gate. Whole-workflow ~2×; per-lookup 10–25×; and two moments where one call was worth more than the rest of the session's tooling combined.
*The model's own report of one engagement, on a version before 0.5; not a controlled measurement. Controlled measurements are in docs/EVALS.md.*
Fifty years of software-engineering results, and research from last month. 49 repositories and 71 papers folded — McCabe (1976) through to seven published in the last two months — each row in docs/LINEAGE.md naming the lesson taken and the file it lives in
Beside those sits a labelled survey of 237 tools that contributed nothing and says so. The two sets are disjoint by construction, so they add rather than nest — a tool that gave a lesson is never counted twice.
Both halves are load-bearing, and they are doing different jobs. The settled results are what make the quality lens trustworthy: McCabe on complexity (1976), Halstead on volume (1977), Spärck Jones on term specificity (1972), Nagappan & Ball on churn. Fifty years of replication means those are not opinions, and a tool that measures your code should be built on the ones that survived.
The recent work is what makes it current: seventeen of the folded papers are from 2026, seven published in the last two months and three in the last thirty days (dates as of 2026-09-08; every row carries its arXiv id, so the claim is checkable rather than atmospheric). Retrieval for coding agents, context-compression cost, placebo-controlled localization — that literature is months old, not decades, and several rows were folded within weeks of the paper appearing.
Neither half alone would be enough. A tool built only on the classics would not know what an agent
needs; one built only on last month's preprints would have nothing underneath it. And the newest row
is a result that failed when it was tested here — which is the point of writing them down. All three counts are re-derived from that document's own tables by
test/readmedriftcheck.sh on every run, which fails if this page and those tables disagree, so the
claim cannot quietly drift. The row-by-row ledger is
docs/LINEAGE.md.
Languages: Rust · C++ · Objective-C/C++ · C · Metal · CUDA · Python · Go · Swift · TypeScript · JavaScript · Java · Ruby · PHP · Lua · Elixir · Dart · Kotlin · GDScript · Bash · C# · JSON · TOML · YAML · Markdown — see language support and limits.
Latest: 0.6.5 — TypeScript alias imports resolve, and fixes from Windows testers. Release notes · the presentation · the changelog — with thanks to the contributors named there; this release is largely theirs.
---
One process, no server — indexes this repository in 0.25 s using 6.6 MB, against 46.8 s and 391 MB for the graph-database MCP server it was measured against; warm queries answer in 197 ms to its 1,082 ms
Measured on 48 matched questions across django, webpack and this repository. Across all three,
ripwire indexes in 0.25–0.45 s and 6.6–16.5 MB against that server's 23–52 s and
391–623 MB. The full method, the wins named one by one and the losses included, is in
Against the leading graph-database code-context MCP server
and docs/EVALS.md.
One binary, offline — and the same line activates the skills for every agent it finds: Claude Code · Codex · Cursor · Windsurf · Gemini · opencode · aider
One self-contained binary on your own machine, offline, installed in one line — and the same line installs and activates the task-shaped skills that teach your agent when to reach for it, not just how, for every agent it finds on the machine. If your agent can run shell commands — Claude Code, Codex, Cursor, Windsurf, Gemini, opencode, aider — it is set up the moment the install finishes; the MCP server is the optional second interface. Install it and ask it something before you finish reading this page:
RIPWIRE_REPO=redhat-et/ripwire bash -c "$(curl -fsSL https://raw.githubusercontent.com/redhat-et/ripwire/main/scripts/install.sh)"
export PATH="$HOME/.local/bin:$PATH" # where it installed; the installer prints this line if you need it
cd your-repo
ripwire . --for=""
Every install route (prebuilt, from source, per-agent skills, hooks, the MCP server) is in INSTALL.md.
Reach for the CLI first — it is the cheaper interface. The MCP server is the optional second way in, and its convenience has a cost the shell pipe does not carry: its verb schemas sit in your agent's context every session, whether or not it calls them.
The goal: one question, one complete answer.
Terminality is the objective. Ask the codebase a question and the answer should carry everything you need — no follow-on grep, no three more whole-file reads to fill in what it left out. A call followed by three greps is the same search paid for twice: it does not save you tokens and it does not make the coding faster.
The two stair-steps that make it reachable — honest about what is missing, priced in what it spends
Two things make that reachable in practice, and neither is the destination. Answers are honest about their own limits — a count that cannot be a total is labelled a floor, a zero means "none found" and never "none exists", every truncation is disclosed — so an answer never looks more complete than it is, and the map never degrades the code by guessing. And an answer can be given a token budget, so what one costs is something you ask for rather than discover; where a complete answer will not fit, it says it went over rather than silently dropping the row you needed.
Those two are the stair-steps: honest about what is missing, priced in what it spends. The step they climb toward is a question fully answered in one call, which is not always trivial to reach — and where it is not, the output says so rather than pretending otherwise.
How we measure the climb. A complete answer to every code question is genuinely hard, and it is reached over many releases rather than in one. Until an answer can be complete, it owes you three things: it stays bounded in size, it is fully honest about what it does not know, and it gives you a clue where to look next. So every change is judged on more than "was it right":
- Complete answers: the share of questions answered so you can act without re-checking.
- Honest-partial rate: how often an incomplete answer says what it is missing, instead of
- False-confidence rate: wrong answers that claim to be complete — the worst case, held near zero.
- Next-clue usefulness: for partial and wrong answers, whether following the answer's first
Answering comes first; efficiency is how we get there. We aim for answers about as lean as the leanest tools, but a byte budget is never allowed to cost an answer. We have gone too far toward short answers before and cut information that was needed, so now every size limit has one of two jobs: a runaway guard far above typical answers (it catches an output bug, and says so when it trips), or a stair-step target where anything over it must be explained. A capped answer is always measured against the same answer uncapped, and the one that answers better wins.
Same answer, a fraction of the tokens — read this table first if your agent is on a budget
How these ten rows were measured — 2026-08-08, figures in ~tokens (≈ bytes/4), every ratio from a real run reproduced by the command in its row
Ten everyday moments, re-measured on this repository, 2026-08-08. Figures are ~tokens (≈ bytes/4);
every ratio comes from a real run, reproduced by the command in its row — raw byte counts and exact
commands in
docs/EVALS.md §5.
Ordered understand → navigate → review-the-change:
| Ask it | Command | ripwire | naive read | token savings |
| --- | --- | --- | --- | --- |
| "Orient me in this repo" | ripwire . | ~5.6K tok | ~20K–25K tok — read README.md (+docs/ARCHITECTURE.md) | 3.6×–4.5× |
| "Where is X handled?" | ripwire . --for="…" | ~2.1K tok | ~4.9K–20K tok — grep -rn src/, then read the file it points at | 2.3×–9.3× |
| "What do I already know?" | ripwire . --recall="…" | ~15K tok | ~445K tok — read all 119 markdown docs this repo carries | 29.2× |
| "Set me up for this task" | ripwire . --pack-task="…" | ~2.1K tok | ~16K–80K tok — read every relevant file, whole | 7.7×–37.7× |
| "Show me this one function" | ripwire . --expand=SYM --top-k=0 | ~260–16.5K tok body (+~5.7K for the ranked-neighborhood bundle) | ~43K–174K tok — read the whole file it lives in | 2.6×–670× |
| "Who calls this function?" | ripwire . --callers=SYM | ~580 tok | ~40K–52K tok — grep -rn SYM src/ (mostly noise), then open 2–3 files to sort real calls from mentions | 69.2×–89.1× |
| "Is it safe to change this?" | ripwire . --impact=SYM + --uses=SYM | ~1.3K tok | ~18K tok — open every direct-use file, whole | 14.4× |
| "I have a stack trace" | ripwire . --from-trace=FILE | ~1.4K tok | ~124K–298K tok — grep all 7 frame names, then open the innermost file(s) | 86.9×–208.6× |
| "I changed these files — tests? blast radius?" | ripwire . --situ | ~410 tok | ~3K–132K tok — git diff + grep -rn test/, then open the candidates | 7.3×–324.2× |
| "Review this PR/diff" | ripwire . --pr-context=REF | ~1.9K tok | ~4.8K–51K tok — git diff REF, then open the touched files | 2.6×–27.5× |
Same-correct-answer verification, and the honesty line these ratios come with
These aren't summaries that gamble with information. Each row is scored
same-correct-answer-or-it-doesn't-count, and both sides were checked, not assumed: orient surfaces
this repo's own pipeline files (ingest.cpp, graph.h, serialize.h) in the first screen, the same
three docs/ARCHITECTURE.md names as central; the --for row lands mcpStale
(src/mcpindex.h:633), the actual staleness check, 5th-ranked; --recall lands the container-rule
doc (AGENTS.md) that states, verbatim, the same "no std::map" rule CONTRIBUTING.md explains in
full; --pack-task names the same three touch points a human would — cachelint.h, mergeCachePack
(src/main.cpp:1787), the lintrules.h helpers it reuses; --expand --top-k=0 hands back the
requested function's complete, unmodified body — the ranked-neighborhood addition costs the same
~22.6 KB regardless of which function you ask for, confirmed on two (a fixed floor, not per-function
variance); --callers on langOfPath names its 2 real callers, the same ones a grep hit-list
buries under 5 files of comment-only mentions; --impact+--uses on coversOrEquals names the same
2 direct call sites --uses alone would, plus (disclosed) a transitive reach --uses doesn't cover
at all; --from-trace resolves all 7 frames of a real call chain by name to the same definitions a
per-frame grep would eventually find, mixed with call sites and comments; --situ on a 2-file diff
names the same 6 real test harnesses, 2 of which a filename grep across test/ cannot find even
after opening every one of its 41 candidates — a completeness gap, not just a byte one; --pr-context
surfaces co-change partners (test/regression.sh, src/main.cpp) a raw git diff has no way to
know were usually touched and weren't this time — Fowler's Shotgun Surgery checked rather than
merely named (the backtest is in docs/EVALS.md). The map ranks and discloses — it never paraphrases
your code — and every truncation is disclosed in the header.
The honesty line, made concrete: the same auto-selection behind the --expand row also runs the
other way. On a small file (pageRankDouble in src/pagerank.cpp, 5,559 B) the ranked bundle would
cost 27,916 B — nearly 5× more than the file — so ripwire serves the file itself instead, disclosed
as mode="whole-file" on the response, not silently.
docs/EVALS.md §7 lists that and the other counterexamples
this project publishes against itself.
Where those savings compound: an orchestrator that matches tasks to models — every lane it spawns starts cold on the same tree. The map is the one artifact that does not have to be rediscovered per agent, and the quality verbs hand a verdict back instead of a pile of files to re-read
The shape: plan the work, then run a loop that matches each task to the model that fits it. Every thread it spawns opens with an empty context on a repository it has never seen. Orienting an empty context to a large repository is the most repeated cost in the whole system, and the one this tool was built for — it is also the cost that grows with the size of the tree, which is why the pattern matters more the bigger the repository gets.
--for and --pack-task answer it in a single call, at the per-call rates in the table above,
instead of a grep-and-read tour that every lane pays over again from scratch.
The return path matters as much. --quality-delta, --test-gate and --edit-check answer *what did
I make worse*, which tests must run, did I change a contract — quantitative answers a lane can
hand back as a verdict, rather than a transcript the orchestrator has to read to find out what
happened.
What is and is not claimed here. Every figure on this page is a single-agent measurement. That the saving compounds with the number of cold orientations follows from the fixed-cost mechanism, but it is pre-registered and unrun — the reason the pattern is worth trying, not a result this project has published.
See the map — not just the numbers

Django's migration autodetector, coloured by complexity. Thresholds are fixed, so the colour means the same thing on every repo you point it at.
ripwire path/to/django/db/migrations --rank-by=rrf --top-k=120 --color-by=cx --html=map.html
![]() |
![]() |
The same graph, re-coloured by git churn. 76% of these nodes move to a differen
GitHub Stars & Activity2,408Stars 151Forks 0Open issues C++Language GitHub PopularityGitHub stars2,408 Forks151 Open issues0 Primary languageC++ License- Stars gained today2,411 Created- Last pushed- Trending HistoryMonthly boardrank #46 · ▲ 2,411 stars Related AI Projects1 2 3 4 5 6 7 8 More AI Rankings |


