pedrohcgs/claude-code-my-workflow
A ready-to-fork Claude Code template for academics using LaTeX/Beamer + R. Multi-agent review, quality gates, adversarial QA, and replication protocols.
About pedrohcgs/claude-code-my-workflow
pedrohcgs/claude-code-my-workflow is an open-source project on GitHub, mainly written in HTML. A ready-to-fork Claude Code template for academics using LaTeX/Beamer + R. Multi-agent review, quality gates, adversarial QA, and replication protocols. It currently holds 1,643 stars and 3,091 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).
Project Overview
AI Homed tracks it on the Today's Trending board, currently at rank #46 with 11 new stars today.
GitHub Repository Details
README
My Claude Code Setup
Actively maintained. A summary of how I use Claude Code for academic work — slides, papers, data analysis, and more — packaged so you can fork it for your own research. See CHANGELOG.md for the latest changes.
Live site: psantanna.com/claude-code-my-workflow
A ready-to-fork foundation for AI-assisted academic work. You describe what you want — lecture slides, a research paper, a data analysis, a replication package — and Claude plans the approach, runs specialized agents, fixes issues, verifies quality, and presents results. Like a contractor who handles the entire job. Extracted from a production PhD course and extended by a growing community.
---
Quick Start (5–10 minutes, plus ~30 min for first-time installs)
Before you start: Claude Code, git and Python 3 are the minimum. Python 3 runs the hooks, the gate suite (./scripts/backtest.sh— 10 checkers) and the quality scorer, and is pre-installed on macOS/Linux. To run the includedHelloWorlddemos end-to-end you also need XeLaTeX (Beamer sample) and Quarto (Quarto sample). R and the GitHub CLI are recommended. Full list in Prerequisites below. Fastest path: clone first, then run./scripts/validate-setup.sh— it reports exactly what's missing with install links.
> Only need Python/R/markdown? You don't need XeLaTeX or Quarto. The agents, rules, skills, and orchestration patterns work for any text/code artifact. Skip theHelloWorlddemos and head straight to/data-analysis,/review-paper,/lit-review, or/review-r.
> Session 2 onwards: MEMORY.md (committed) collects generic[LEARN]entries that help all forkers; machine-specific notes accumulate in Claude Code's native auto memory (~/.claude/projects//memory/, machine-local, never committed). See.claude/rules/meta-governance.mdfor the distinction.
1. Fork & Clone
# Fork this repo on GitHub (click "Fork" on the repo page), then:
git clone https://github.com/YOUR_USERNAME/claude-code-my-workflow.git my-project
cd my-project
./scripts/validate-setup.sh # reports missing tools with install links
Replace YOUR_USERNAME with your GitHub username.
2. Start Claude Code and Paste This Prompt
claude
Using VS Code? Open the Claude Code panel instead. Everything works the same — see the full guide for details.
Avoid prompt fatigue. On Claude Code ≥ 2.1.283, new interactive sessions start in auto mode by default (classifier-gated — most actions run, risky ones prompt; on earlier versions this applied to Pro/Max/Team); where auto is unavailable, Manual mode prompts per risky tool call. If you still see too many prompts, toggle Auto-accept edits mode (a keybinding; see the permission modes section of the guide) or runclaude --permission-mode acceptEdits. Bypass mode skips permission prompts and safety checks (deny rules still apply), and Anthropic scopes it to isolated containers and VMs; it takes effect from the CLI flag,--settings, or user/managed settings (~/.claude/settings.json) — a bypass default in a project's.claude/settings.jsonis not honoured (the session starts in Manual). The template's.claude/settings.jsonsets no default mode (terminal sessions get the platform's auto mode); its.vscode/settings.jsoncarries two bypass keys, but the current VS Code extension ignores them (it readsclaudeCode.initialPermissionModeandclaudeCode.allowDangerouslySkipPermissionsonly from your VS Code user settings, and never reads the unprefixedallowDangerouslySkipPermissionskey), so VS Code sessions also start in auto unless your user settings say otherwise — set bothclaudeCode.keys there if you want bypass in VS Code — and the.claude/settings.jsonships broad catch-all allows (Bash(),Edit(),Write()— 7 wildcard rules, not a curated list), so in Manual mode almost nothing prompts; auto mode drops the blanketBash()and routes shell commands through its classifier. Working with restricted data? Add deny rules — see TROUBLESHOOTING.
Then paste the starter prompt from the guide, filling in your project details:
I am starting to work on [PROJECT NAME] in this repo. [Describe your project in 2–3 sentences.] I've set up the Claude Code academic workflow... Please read the configuration files and adapt them for my project. Enter plan mode and start.
The full guide has the complete starter prompt with all the details.
What this does: Claude reads all the configuration files, fills in your project name, institution, and preferences, then enters contractor mode — planning, implementing, and (within the skill you invoke) running the review + verify loop. You approve the plan, invoke a skill, and the skill handles the rest within its scope.
Heavily adapting CLAUDE.md for a non-academic project? Anthropic's built-in/initcommand will re-derive aCLAUDE.mdfrom your codebase as a starting point. The pre-shipped CLAUDE.md in this template already covers the academic setup — you only need/initif your fork diverges substantially (e.g., a Python/ML project that doesn't use LaTeX or Quarto).
3. Verify Your Setup
Before building real lectures, confirm your environment works:
./scripts/validate-setup.sh # Checks XeLaTeX, Quarto, Python, git, etc.
Then inside Claude:
/compile-latex HelloWorld # Compiles Slides/HelloWorld.tex to PDF
/deploy HelloWorld # Renders Quarto/HelloWorld.qmd to HTML
If both succeed, delete Slides/HelloWorld.tex and Quarto/HelloWorld.qmd and start on your real work.
---
How It Works
Goal-first, gate-enforced (the v2.0 shift)
You don't craft a perfect prompt — you state a goal and let the work loop toward it under gates. Specialist agents do the labor; enforcing gates decide when it's good enough; you adjudicate the disagreements they surface. Three things make that trustworthy:
- Real gates, not reminders. One command —
./scripts/backtest.sh— runs ten gates: surface-sync, skill integrity, model currency against the SSoT, link and anchor resolution, Agent Skills spec conformance, staleness (including source-vs-published divergence), repo hygiene, derived counts (enumerable claims re-counted from disk), ledger coverage (the qualification ledger and the checks that actually run must agree in both directions, and every hook declared in settings must exist and be invocable — a one-character path typo no longer disables a hook in silence), and a seeded hook battery (every active guard hook is re-fired against the failure it targets, alongside clean controls, on every run). A version-controlled pre-commit hook (run./scripts/install-hooks.shonce) runs it plus the quality check (≥80) on every commit (the hook battery only when a hook, its settings or the battery itself is staged; CI always runs everything) — bypassing the skill no longer bypasses the review. Agit-guardrailshook blocks destructive git (reset --hard,clean -f,push --force,add -A) and refuses a merge, rebase, or pull while the tree is dirty — reading the tree as it is rather than predicting what a chained command might do to it, sogit stash && git mergeis denied too and you run the two steps separately (ALLOW_DIRTY_MERGE=1if you mean it); like its sibling it is a textual check over the command line, so an op carried inside an interpreter, an alias, or a script is outside what it can see, and its docstring says which forms those are. Its siblingroot-of-trust-guarddenies the common shell write paths into the files that define the gates themselves (.claude/settings*.json,.claude/hooks/,.githooks/) — redirection,tee,cp/mv/rm, in-placesed— and, because it already unwrapsbash -candenv -Spayloads for those rules, it hands the unwrapped payload to the same destructive-git deny list, sobash -c 'git reset --hard'andbash -c 'git clean -fdx'no longer fall between the two hooks. Read that one as a tripwire, not a lock, and read the name as a filename rather than a claim: it is a best-effort textual scan that fails open on its own errors, and the files it watches stay replaceable through channels the template deliberately allows — anEdit/Write/MultiEdit, a branch switch, a clean merge, or a bug in the hook itself. What it buys is a change of channel — a change to a gate arrives as a reviewable diff instead of an invisible overwrite — not a guarantee that the gates cannot be disabled. Nothing here recovers anything either: the transcript andgit reflogare an audit trail and a commit-history aid, and neither holds the bytes of an uncommitted edit or an untracked file. The review runtime re-checks any reviewer-introduced "fatal" finding before it counts. - Every gate is qualified, and the ledger is itself a gate. Each one has been shown a planted defect and confirmed to go red, with recall and false-alarm rate recorded in
quality_reports/qualification/LEDGER.md— and that ledger is now load-bearing rather than aspirational: a registered check with no row there fails the build, and a row naming a checker that no longer exists fails too. Checks that have not been qualified are listed there by name as visible debt — because an unqualified check is not weak evidence, it is none. Run/vaccinateto qualify one. - A real orchestration runtime. Reviews fan out to forked specialist agents, reduce over a shared finding schema, judge with a hallucination gate, and loop until dry — see
orchestrator-protocol.md. - Ground truth as a process. A mismatch isn't always a failure: a defensible, named alternative is recorded as
EXPLAINEDand carried into your response-to-referees, while genuine errors stay fail-closed.
Contractor Mode
You describe a task. For complex or ambiguous requests, Claude first creates a requirements specification with MUST/SHOULD/MAY priorities and clarity status (CLEAR/ASSUMED/BLOCKED). You approve the spec, then Claude plans the approach and runs the right skill — or, for user-invoked skills such as /create-lecture, tells you which one to run (e.g. /create-lecture, /qa-quarto, /review-paper --adversarial). That skill implements the orchestrator runtime internally — implement, verify, review, fix, re-verify, score — and returns a summary when the work meets quality standards. Say "just do it" and it runs the full loop; commits still require an explicit /commit (which the pre-commit hook then gates).
Specialized Agents
Instead of one general-purpose reviewer, 18 focused agents each check one dimension. A representative sample:
- proofreader — grammar/typos
- slide-auditor — visual layout
- pedagogy-reviewer — teaching quality
- r-reviewer — R code quality
- domain-reviewer — field-specific correctness, slides (template — customize for your field)
- domain-referee / methods-referee / editor — manuscript peer-review pipeline (
/review-paper --peer)
/slide-excellence skill runs the slide-review agents in parallel; /review-paper --peer runs the paper-review pipeline. The same pattern extends to any academic artifact — manuscripts, data pipelines, proposals.
Adversarial QA
Two agents work in opposition: the critic reads both Beamer and Quarto and produces harsh findings. The fixer implements exactly what the critic found. They loop until dry — converging after two consecutive rounds surface no new issue (a 5-round cap is the fallback, not the primary stop). This catches errors that single-pass review misses.
Quality Review
Every artifact gets a score (0–100). Scores below threshold halt the workflow and surface the findings — the user decides whether to fix or explicitly override:
- 80 — commit threshold
- 90 — PR threshold
- 95 — excellence (aspirational)
Framing honesty: Thresholds are advisory at the harness level — the/commitskill runs quality checks and halts on failure. And as of v2.0, running./scripts/install-hooks.shonce installs a real pre-commit hook (.githooks/pre-commit) that runs the backtest gate suite (the hook battery only when a hook, its settings or the battery itself is staged; CI always runs everything) plus the quality (≥80) gate on every commit, so bypassing the skill no longer bypasses the review. Opt out per-commit withSKIP_QUALITY_GATE=1orgit commit --no-verify.
Context Survival
Plans, specifications, and session logs survive auto-compression and session boundaries. The PreCompact hook saves a context snapshot before Claude's auto-compression triggers, ensuring critical decisions are never lost. MEMORY.md accumulates learning across sessions, so patterns discovered in one session inform future work.
For forced compression (long pipelines, mid-plan handoffs), /compress-session (v1.9.0) distils the conversation into a structured note — decisions, next actions, and discarded-as-noise — instead of letting auto-compaction truncate. /promote-memory (v1.9.0) periodically harvests generic learnings from native auto memory to committed MEMORY.md via a five-critic council.
Verification Discipline (v1.7.0+)
Multiple complementary verification layers run before submission:
/verify-claims(v1.7.0) — Chain-of-Verification with a fresh-context verifier that cannot self-confirm because it has never seen the draft. v1.9.0 adds HIGH/MED/LOW-WARN severity tiers; HIGH-WARN findings (fabricated citation, numerical contradiction) fail the verification closed — the draft is never reported as verified;/commitdoes not read these verdicts, so resolving them before committing is on you./audit-reproducibility(v1.4.0; Stata coverage v1.9.0) — every numeric claim in the manuscript is cross-checked against the script output that produced it. v1.9.0 addspassport.yaml— a per-paper YAML state file with PASS/FAIL/STALE/UNVERIFIED status per claim./humanize(v1.9.0) — detect AI-voice tells (boilerplate transitions, hedging stacking, sycophancy) before submission. Read-only by design; auto-rewriting degrades quality./review-paper --peer --variance N(v1.9.0) — runs N referees with sampled dispositions and reports a decision distribution, not a point estimate. Motivated by AgentReview (EMNLP 2024) finding 37% of decisions vary purely from disposition sampling.
The Guide
For a comprehensive walkthrough, read the full guide (or see the source).
It covers: 1. Why This Workflow Exists — the problem and the vision 2. Getting Started — fork, paste one prompt, and Claude sets up the rest 3. The System in Action — specialized agents, adversarial QA, quality scoring 4. The Building Blocks — CLAUDE.md, rules, skills, agents, hooks, memory 5. Workflow Patterns — slides, research, reproducibility, presentation rhetoric, sequential adversarial audits, and more 6. The Ecosystem — extensions by clo-author, claudeblattman, MixtapeTools, autoresearch, ClaudeCodeTools, and a growing community 7. Customizing for Your Domain — creating your own reviewers and knowledge bases
2026 Features
The guide covers Claude Code's latest capabilities:
- Model lineup — the Fable tier (opt-in via
/model fableor thebestalias) is the top tier for long-horizon work; the Opus tier is the Claude Code default and the template's high-judgment tier. Current point versions, prices, effort defaults, and the provider-dependent alias table live in the single source of truth,model-versions.md— surfaces here stay tier-abstract so they cannot go stale, and the staleness gate fails the build when the SSoT's own expiry passes. - Effort levels —
/effortsets cost vs. thoroughness (low/medium/high/xhigh/max). Default effort differs by tier (per the model SSoT): the current Opus defaults tomedium— and itsmediummatches the prior Opus generation'shigh— while the Fable and Sonnet tiers default tohigh. Set effort explicitly, lower it before prompting for brevity, reservexhighfor extended exploration andultracode(xhigh + dynamic workflows) for the largest autonomous runs;maxEffortLevelcaps effort on every surface. /goal(v1.9.0; Anthropic May 2026) — keep working across turns until a fast model confirms the condition holds. Pairs with/commitquality gates for verified-end-state runs.claude agentsdashboard (v1.9.0; Anthropic May 2026) — single screen for parallel review work (/review-paper --peer,/slide-excellence).- Cost-Conscious Composition — prompt-cache TTL (5-min default on API keys; 1-hour automatic on Claude subscriptions), 70/20/10 model routing (Haiku/Sonnet/Opus),
/cost+/usagemonitoring, Agent SDK credit-pool split (2026-06-15). - Skill frontmatter —
effort,context: fork,agent,hooks,disable-model-invocation(v1.8.0+),disallowed-tools(the actual tool restriction —allowed-toolsonly pre-approves),paths(glob-scoped auto-activation), and dynamic content ($ARGUMENTS,!commandsyntax) - Permission modes — Manual (config value
default), Accept edits, Plan, Auto (classifier-gated; the built-in starting mode for new interactive terminal and VS Code sessions on Claude Code ≥ 2.1.283 — on earlier versions, on Pro, Max, and Team since 2026-08-14 — and available on Bedrock / Google Cloud / Foundry without an opt-in flag), Bypass - Hook handler types — command, prompt, and HTTP handlers with 20+ hook events; hooks see
effort.leveland$CLAUDE_EFFORT(Apr 2026 Week 19) - Advanced agent configuration — model, maxTurns, isolation, tool restrictions;
model-routing.mdrule codifies per-agent tier (v1.9.0) - Worktree base ref (v1.9.0; Anthropic Apr 2026) —
worktree.baseRefsetting controlsfresh(default; remote default-branch) vshead(local HEAD) for new worktrees - Built-in skills —
/fewer-permission-prompts,/team-onboarding,/autofix-pr,/powerup, Ultraplan,/loop(self-pacing) - Plugins —
/discover-pluginsfor third-party extensions - Mid-2026 additions (Jun–Sep 2026) —
/skill-doctor(what each skill costs in context and how often it fires),claude plugin eval(score a plugin against test cases vs a no-plugin baseline),maxEffortLevel(cap effort on every surface — a hard ceiling for a fixed grant budget), built-in output styles (Concise for everyday work; Learning, which leaves parts for you to write — pairs with/scaffold-exercises), subagents running in the background and fork mode on by default,/fork(continue a copy of the conversation in a background session),/doctor(setup checkup), andautoUpdatesChannel: "stable"for deadline weeks
Use Cases
| Academic Task | How This Workflow Helps |
|---------------|----------------------|
| Lecture slides (Beamer/Quarto) | Full creation, translation, multi-agent review, deployment |
| Research papers | Literature review, manuscript review, simulated peer review (/review-paper --peer [journal]), reviewer-disposition variance reporting (--variance N) |
| Data analysis | End-to-end R pipelines (/data-analysis) or Stata pipelines via stata-mcp (/stata-replication, v1.9.0), replication verification, publication-ready output |
| Monte Carlo simulations | Reproducible simulation studies (/simulation-study, v1.10.0) — parameterized DGP, estimator grid, bias/RMSE/coverage/size/power with Monte Carlo SEs, dedicated sim-reviewer review pass |
| Package development | R package release gate (/r-package-check, v1.10.0) — devtools::document() + tests + R CMD check --as-cran + CRAN-policy triage + r-package-reviewer (Stata / Python checks on the roadmap) |
| Replication packages | AEA-compliant packaging, reproducibility audit trails, passport.yaml claims provenance (v1.9.0) |
| Presentations | Rhetoric of decks principles, visual audit, cognitive load review |
| Research proposals | Structured drafting with adversarial critique |
| Preregistration | OSF / AsPredicted / AEA RCT Registry-ready document (/preregister --style) — full workflow in Pattern 16 |
| Manuscript submission discipline | /humanize (detect AI voice), /verify-claims HIGH-WARN fail-closed reporting (a fabricated citation is never reported as verified), reviewer-disposition variance |
Disciplines preloaded: Economics (top-5 journal profiles, R conventions) and Political Science (APSR / AJPS / JOP profiles, formal-theory + survey-experiment paper types, conjoint/cjoint conventions). Forkers extend for psych / sociology / public-health via journal profiles + paper types + discipline cards.
One repo, many project types
This workflow is designed as a single hub for an entire research program — not one paper at a time. The same CLAUDE.md, rules, agents, and quality gates serve courses and lectures, papers and referee reports, data analysis and replication packages, Monte Carlo simulation studies (/simulation-study + sim-reviewer), and the R package release gate (/r-package-check + r-package-reviewer) — all new in v1.10.0. On the roadmap: Stata / Python package checks (SSC / PyPI) and personal-productivity workflows. See .claude/references/v2.0-backlog.md for what's next.
---
What's Included
18 agents, 61 skills, 37 rules, 11 hooks (click to expand)
Agents (.claude/agents/)
| Agent | What It Does |
|-------|-------------|
| proofreader | Grammar, typos, overflow, consistency review |
| slide-auditor | Visual layout audit (overflow, font consistency, spacing) |
| pedagogy-reviewer | 13-pattern pedagogical review (narrative arc, notation density, pacing) |
| r-reviewer | R code quality, reproducibility, and domain correctness |
| tikz-reviewer | Merciless TikZ diagram visual critique |
| beamer-translator | Beamer-to-Quarto translation specialist |
| quarto-critic | Adversarial QA comparing Quarto against Beamer benchmark |
| quarto-fixer | Implements fixes from the critic agent |
| verifier | End-to-end task completion verification |
| `domain