redhat-et/ripwire

▲ 2,411 stars today★ 2,408⑂ 151

The ripgrep of AI context: a zero-dependency C++23 CLI + MCP server for coding agents. Find what you want without reading the repo, then check you built what you meant — blast radius, tests-to-run

About redhat-et/ripwire

redhat-et/ripwire is an open-source project on GitHub, mainly written in C++. The ripgrep of AI context: a zero-dependency C++23 CLI + MCP server for coding agents. It currently holds 2,408 stars and 151 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Today's Trending board.

GitHub Repository Details

Repository redhat-et/ripwire · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

CI Release Licence Standard Runtime dependencies Slides

https://github.com/redhat-et/ripwire/blob/HEAD/ripwire — the ripgrep of AI context

Rip'n Fast. Fewer Tokens. Better Code.

The ripgrep of AI context. A map before your agent reads the repo — and a check on what it writes.

Ranked, deterministic call graph: what to touch, what it breaks, which tests to run. On the edit: blast radius, tests that reach it, eleven quality kinds reporting only what got worse, forgotten co-changes, fields read and written, names that resolve more than one way.

Just want to use it? Install it with the one line below, then start each coding session by telling your agent to use it, for example: "Use ripwire on this repo." That is all most people need: the install also teaches your agent when to reach for each command.

Want every detail? The reference guide near the bottom covers install, commands, output format, exit codes and limits. You do not need it to get started.

https://github.com/redhat-et/ripwire/blob/HEAD/Trendshift: C++ Repository of the Week badge for redhat-et/ripwire

Field report: what ripwire contributed to a large multi-agent coding engagement — written by Claude Fable 5.0, the frontier model orchestrating ~20 coding agents over two days on a ~1,500-file C++/Metal codebase. Click for the full report.
For a single developer, this tool is a good lookup accelerator. For an orchestrated fleet, it's load-bearing: it
halved the research spend, twice redirected tasks before wasted work, prevented at least one silent-divergence shipped
bug, and turned code-quality hygiene from a hope into a per-task mechanical gate. Whole-workflow ~2×; per-lookup
10–25×; and two moments where one call was worth more than the rest of the session's tooling combined.
> — the report's bottom line

https://github.com/redhat-et/ripwire/blob/HEAD/The full field report: what ripwire contributed to a large multi-agent coding engagement — headline numbers, where the value concentrated, the honest boundary, and the bottom line

Text version of the report

Field report: what ripwire contributed to a large multi-agent coding engagement

*Context, genericized: one orchestrating session directing ~20 sequential/parallel coding agents over a ~1,500-file C++/Metal codebase across two days — a deep architecture audit, then a 14-task feature wave (new subsystems, measurement infrastructure, a search-archive migration), ending in a verified all-on release flip. Every agent was instructed to lead with ripwire for orientation and to close with its quality gates.*

The headline numbers

  • Roughly half the total token spend of the audit/research phase, saved. This is the operator's whole-phase
estimate — it includes everything the research agents did, ordinary file reads and all, which makes it conservative rather than cherry-picked. The per-call factors underneath are much steeper: a single doc-recall call served the relevant sections of a 164KB planning document in ~6K tokens (~25×), and agents that led with the tool ran ~30–40% leaner on tool-call counts than agents doing raw read fan-outs over comparable questions.
  • Two tasks had their direction changed by a single call. One task was chartered to activate a steering behavior
the team believed was in production; one --callers query returned zero production callers, redirecting the task to the real gap (the data provider that behavior needed had never been wired anywhere). Another task refactored a widely-shared computation; --edit-check flagged 5 of 6 call sites as incompatible — sites a text search had missed — that would otherwise have silently diverged from the canonical path the day the feature was enabled, surfacing months later as an unexplainable visual bug.
  • A dozen-plus real code-quality defects fixed, not waived, across ~15 implementing agents — all caught by
--quality-delta at each agent's "I think I'm done" moment: a 480-token duplicated routine, a fourth private copy of a shared RNG utility, a hand-duplicated cost function, a pair of near-identical functions with a flipped sign (correct at one boundary, quietly wrong at the other), test-fixture duplication across sibling suites, and functions that had quietly absorbed a second job. None of these would have failed a test; all of them are the sediment that rots a codebase under high-velocity multi-agent development. The gate made removing them routine instead of heroic.

Where the value concentrated

1. Orientation and recall (the token win). Ranked-signature task orientation and section-granular document recall meant agents started from answers, not file dumps. Recall over planning docs was the orchestrator's single most-used verb — the audit's entire framing came from two calls. 2. The contract checker (the shipped-bug preventions). Beyond the 5-of-6 catch above, --edit-check gave near-free proof-of-contract on every public-symbol change across the wave — the kind of verification that otherwise simply doesn't happen at agent speed. 3. **The quality gate as a protocol. The measurable effect wasn't any single catch — it was that fifteen different agents, none sharing context, all converged on the same "fix or justify with a written reason" discipline, with an acknowledgments ledger that survived across tasks. Unattended orchestration usually leaks quality; here the leak-check was mechanized. 4. Persistent notes. Agents left gotcha notes pinned to symbols mid-wave (a caller-wiring law, an arming-rule invariant), which later agents' lookups surfaced automatically — cheap institutional memory between contexts that never met.

The honest boundary

  • The engagement's deepest findings did not come from the tool.** A structural capacity ceiling, a mis-derived
constant, a floating-point re-association drift of 1 ulp, and a seed-keying bug were all found by *bespoke measurement the agents built* — baseline worktree diffs over full data fills, multi-arm attribution sweeps, funnels. The tool is a floor for honesty and orientation, not a substitute for measurement design.
  • Known false-positive mode, handled by its own documentation: the contract checker's arity heuristic
over-counts on defaulted trailing parameters (one task saw "incompatible=18" that were all fine). The tool's own trust-calibration notes say to verify by compiling in exactly this case, and agents that followed them lost nothing.
  • For broad, common-word conceptual queries, plain grep-and-read still occasionally won — consistent with the tool's
own guidance that it shines on specific technical asks.

Bottom line

For a single developer, this tool is a good lookup accelerator. For an orchestrated fleet, it's load-bearing: it halved the research spend, twice redirected tasks before wasted work, prevented at least one silent-divergence shipped bug, and turned code-quality hygiene from a hope into a per-task mechanical gate. Whole-workflow ~2×; per-lookup 10–25×; and two moments where one call was worth more than the rest of the session's tooling combined.

*The model's own report of one engagement, on a version before 0.5; not a controlled measurement. Controlled measurements are in docs/EVALS.md.*

Fifty years of software-engineering results, and research from last month. 49 repositories and 71 papers folded — McCabe (1976) through to seven published in the last two months — each row in docs/LINEAGE.md naming the lesson taken and the file it lives in

Beside those sits a labelled survey of 237 tools that contributed nothing and says so. The two sets are disjoint by construction, so they add rather than nest — a tool that gave a lesson is never counted twice.

Both halves are load-bearing, and they are doing different jobs. The settled results are what make the quality lens trustworthy: McCabe on complexity (1976), Halstead on volume (1977), Spärck Jones on term specificity (1972), Nagappan & Ball on churn. Fifty years of replication means those are not opinions, and a tool that measures your code should be built on the ones that survived.

The recent work is what makes it current: seventeen of the folded papers are from 2026, seven published in the last two months and three in the last thirty days (dates as of 2026-09-08; every row carries its arXiv id, so the claim is checkable rather than atmospheric). Retrieval for coding agents, context-compression cost, placebo-controlled localization — that literature is months old, not decades, and several rows were folded within weeks of the paper appearing.

Neither half alone would be enough. A tool built only on the classics would not know what an agent needs; one built only on last month's preprints would have nothing underneath it. And the newest row is a result that failed when it was tested here — which is the point of writing them down. All three counts are re-derived from that document's own tables by test/readmedriftcheck.sh on every run, which fails if this page and those tables disagree, so the claim cannot quietly drift. The row-by-row ledger is docs/LINEAGE.md.

Languages: Rust · C++ · Objective-C/C++ · C · Metal · CUDA · Python · Go · Swift · TypeScript · JavaScript · Java · Ruby · PHP · Lua · Elixir · Dart · Kotlin · GDScript · Bash · C# · JSON · TOML · YAML · Markdown — see language support and limits.

Latest: 0.6.5 — TypeScript alias imports resolve, and fixes from Windows testers. Release notes · the presentation · the changelog — with thanks to the contributors named there; this release is largely theirs.

---

https://github.com/redhat-et/ripwire/blob/HEAD/No API key. No embeddings. No index server. No daemon.

One process, no server — indexes this repository in 0.25 s using 6.6 MB, against 46.8 s and 391 MB for the graph-database MCP server it was measured against; warm queries answer in 197 ms to its 1,082 ms

Measured on 48 matched questions across django, webpack and this repository. Across all three, ripwire indexes in 0.25–0.45 s and 6.6–16.5 MB against that server's 23–52 s and 391–623 MB. The full method, the wins named one by one and the losses included, is in Against the leading graph-database code-context MCP server and docs/EVALS.md.

One binary, offline — and the same line activates the skills for every agent it finds: Claude Code · Codex · Cursor · Windsurf · Gemini · opencode · aider

One self-contained binary on your own machine, offline, installed in one line — and the same line installs and activates the task-shaped skills that teach your agent when to reach for it, not just how, for every agent it finds on the machine. If your agent can run shell commands — Claude Code, Codex, Cursor, Windsurf, Gemini, opencode, aider — it is set up the moment the install finishes; the MCP server is the optional second interface. Install it and ask it something before you finish reading this page:

RIPWIRE_REPO=redhat-et/ripwire bash -c "$(curl -fsSL https://raw.githubusercontent.com/redhat-et/ripwire/main/scripts/install.sh)"
export PATH="$HOME/.local/bin:$PATH"      # where it installed; the installer prints this line if you need it
cd your-repo
ripwire . --for=""

Every install route (prebuilt, from source, per-agent skills, hooks, the MCP server) is in INSTALL.md.

Reach for the CLI first — it is the cheaper interface. The MCP server is the optional second way in, and its convenience has a cost the shell pipe does not carry: its verb schemas sit in your agent's context every session, whether or not it calls them.

The goal: one question, one complete answer.

Terminality is the objective. Ask the codebase a question and the answer should carry everything you need — no follow-on grep, no three more whole-file reads to fill in what it left out. A call followed by three greps is the same search paid for twice: it does not save you tokens and it does not make the coding faster.

The two stair-steps that make it reachable — honest about what is missing, priced in what it spends

Two things make that reachable in practice, and neither is the destination. Answers are honest about their own limits — a count that cannot be a total is labelled a floor, a zero means "none found" and never "none exists", every truncation is disclosed — so an answer never looks more complete than it is, and the map never degrades the code by guessing. And an answer can be given a token budget, so what one costs is something you ask for rather than discover; where a complete answer will not fit, it says it went over rather than silently dropping the row you needed.

Those two are the stair-steps: honest about what is missing, priced in what it spends. The step they climb toward is a question fully answered in one call, which is not always trivial to reach — and where it is not, the output says so rather than pretending otherwise.

How we measure the climb. A complete answer to every code question is genuinely hard, and it is reached over many releases rather than in one. Until an answer can be complete, it owes you three things: it stays bounded in size, it is fully honest about what it does not know, and it gives you a clue where to look next. So every change is judged on more than "was it right":

  • Complete answers: the share of questions answered so you can act without re-checking.
  • Honest-partial rate: how often an incomplete answer says what it is missing, instead of
sounding confident.
  • False-confidence rate: wrong answers that claim to be complete — the worst case, held near zero.
  • Next-clue usefulness: for partial and wrong answers, whether following the answer's first
suggested next step reaches what the question needed, within one hop. Graded blind.

Answering comes first; efficiency is how we get there. We aim for answers about as lean as the leanest tools, but a byte budget is never allowed to cost an answer. We have gone too far toward short answers before and cut information that was needed, so now every size limit has one of two jobs: a runaway guard far above typical answers (it catches an output bug, and says so when it trips), or a stair-step target where anything over it must be explained. A capped answer is always measured against the same answer uncapped, and the one that answers better wins.

Same answer, a fraction of the tokens — read this table first if your agent is on a budget

How these ten rows were measured — 2026-08-08, figures in ~tokens (≈ bytes/4), every ratio from a real run reproduced by the command in its row

Ten everyday moments, re-measured on this repository, 2026-08-08. Figures are ~tokens (≈ bytes/4); every ratio comes from a real run, reproduced by the command in its row — raw byte counts and exact commands in docs/EVALS.md §5.

Ordered understand → navigate → review-the-change:

| Ask it | Command | ripwire | naive read | token savings | | --- | --- | --- | --- | --- | | "Orient me in this repo" | ripwire . | ~5.6K tok | ~20K–25K tok — read README.md (+docs/ARCHITECTURE.md) | 3.6×–4.5× | | "Where is X handled?" | ripwire . --for="…" | ~2.1K tok | ~4.9K–20K tok — grep -rn src/, then read the file it points at | 2.3×–9.3× | | "What do I already know?" | ripwire . --recall="…" | ~15K tok | ~445K tok — read all 119 markdown docs this repo carries | 29.2× | | "Set me up for this task" | ripwire . --pack-task="…" | ~2.1K tok | ~16K–80K tok — read every relevant file, whole | 7.7×–37.7× | | "Show me this one function" | ripwire . --expand=SYM --top-k=0 | ~260–16.5K tok body (+~5.7K for the ranked-neighborhood bundle) | ~43K–174K tok — read the whole file it lives in | 2.6×–670× | | "Who calls this function?" | ripwire . --callers=SYM | ~580 tok | ~40K–52K tok — grep -rn SYM src/ (mostly noise), then open 2–3 files to sort real calls from mentions | 69.2×–89.1× | | "Is it safe to change this?" | ripwire . --impact=SYM + --uses=SYM | ~1.3K tok | ~18K tok — open every direct-use file, whole | 14.4× | | "I have a stack trace" | ripwire . --from-trace=FILE | ~1.4K tok | ~124K–298K tok — grep all 7 frame names, then open the innermost file(s) | 86.9×–208.6× | | "I changed these files — tests? blast radius?" | ripwire . --situ | ~410 tok | ~3K–132K tok — git diff + grep -rn test/, then open the candidates | 7.3×–324.2× | | "Review this PR/diff" | ripwire . --pr-context=REF | ~1.9K tok | ~4.8K–51K tok — git diff REF, then open the touched files | 2.6×–27.5× |

Same-correct-answer verification, and the honesty line these ratios come with

These aren't summaries that gamble with information. Each row is scored same-correct-answer-or-it-doesn't-count, and both sides were checked, not assumed: orient surfaces this repo's own pipeline files (ingest.cpp, graph.h, serialize.h) in the first screen, the same three docs/ARCHITECTURE.md names as central; the --for row lands mcpStale (src/mcpindex.h:633), the actual staleness check, 5th-ranked; --recall lands the container-rule doc (AGENTS.md) that states, verbatim, the same "no std::map" rule CONTRIBUTING.md explains in full; --pack-task names the same three touch points a human would — cachelint.h, mergeCachePack (src/main.cpp:1787), the lintrules.h helpers it reuses; --expand --top-k=0 hands back the requested function's complete, unmodified body — the ranked-neighborhood addition costs the same ~22.6 KB regardless of which function you ask for, confirmed on two (a fixed floor, not per-function variance); --callers on langOfPath names its 2 real callers, the same ones a grep hit-list buries under 5 files of comment-only mentions; --impact+--uses on coversOrEquals names the same 2 direct call sites --uses alone would, plus (disclosed) a transitive reach --uses doesn't cover at all; --from-trace resolves all 7 frames of a real call chain by name to the same definitions a per-frame grep would eventually find, mixed with call sites and comments; --situ on a 2-file diff names the same 6 real test harnesses, 2 of which a filename grep across test/ cannot find even after opening every one of its 41 candidates — a completeness gap, not just a byte one; --pr-context surfaces co-change partners (test/regression.sh, src/main.cpp) a raw git diff has no way to know were usually touched and weren't this time — Fowler's Shotgun Surgery checked rather than merely named (the backtest is in docs/EVALS.md). The map ranks and discloses — it never paraphrases your code — and every truncation is disclosed in the header.

The honesty line, made concrete: the same auto-selection behind the --expand row also runs the other way. On a small file (pageRankDouble in src/pagerank.cpp, 5,559 B) the ranked bundle would cost 27,916 B — nearly 5× more than the file — so ripwire serves the file itself instead, disclosed as mode="whole-file" on the response, not silently. docs/EVALS.md §7 lists that and the other counterexamples this project publishes against itself.

Where those savings compound: an orchestrator that matches tasks to models — every lane it spawns starts cold on the same tree. The map is the one artifact that does not have to be rediscovered per agent, and the quality verbs hand a verdict back instead of a pile of files to re-read

The shape: plan the work, then run a loop that matches each task to the model that fits it. Every thread it spawns opens with an empty context on a repository it has never seen. Orienting an empty context to a large repository is the most repeated cost in the whole system, and the one this tool was built for — it is also the cost that grows with the size of the tree, which is why the pattern matters more the bigger the repository gets.

--for and --pack-task answer it in a single call, at the per-call rates in the table above, instead of a grep-and-read tour that every lane pays over again from scratch.

The return path matters as much. --quality-delta, --test-gate and --edit-check answer *what did I make worse*, which tests must run, did I change a contract — quantitative answers a lane can hand back as a verdict, rather than a transcript the orchestrator has to read to find out what happened.

What is and is not claimed here. Every figure on this page is a single-agent measurement. That the saving compounds with the number of cold orientations follows from the fixed-cost mechanism, but it is pre-registered and unrun — the reason the pattern is worth trying, not a result this project has published.

See the map — not just the numbers

https://github.com/redhat-et/ripwire/blob/HEAD/ripwire --html on Django's migration autodetector: 120 symbols, 183 call edges, arrows pointing caller to callee, nodes coloured by cyclomatic complexity on a five-stop scale running deep blue, mid blue, amber, orange, pale yellow, module outlines drawn as translucent regions, and low-confidence call edges drawn with dashed shafts

Django's migration autodetector, coloured by complexity. Thresholds are fixed, so the colour means the same thing on every repo you point it at.

ripwire path/to/django/db/migrations --rank-by=rrf --top-k=120 --color-by=cx --html=map.html
https://github.com/redhat-et/ripwire/blob/HEAD/The same graph twice: above coloured by cyclomatic complexity, below by git commit count. Most nodes sit in a different colour band between the two. https://github.com/redhat-et/ripwire/blob/HEAD/A close crop showing solid and dashed call edges side by side; dashed shafts mark calls the resolver could not pin to a single target
The same graph, re-coloured by git churn. 76% of these nodes move to a differen

GitHub Stars & Activity

2,408Stars
151Forks
0Open issues
C++Language

GitHub Popularity

GitHub stars2,408
Forks151
Open issues0
Primary languageC++
License-
Stars gained today2,411
Created-
Last pushed-

Trending History

Monthly boardrank #46 · ▲ 2,411 stars

Related AI Projects

More AI Rankings