abdiff is a powerful tool designed to assess whether edits to CLAUDE.md alter Claude's behavior. By comparing results before and after rule changes side by side, it provides clarity on the impact of modifications, ensuring users can confidently understand and manage their rules within Claude.
abdiff is a specialized tool designed to assess whether modifications made in the CLAUDE.md file influence the behavior of Claude. This utility offers a systematic approach to validate the effectiveness of rules added to CLAUDE.md, allowing users to run tasks before and after rule changes and compare the outputs side by side.
Key Learning Outcomes
Determine if a newly added rule alters Claude's functionality or the output.
Verify the significance of existing rules and whether they should remain in place.
Establish if a reference document is actively considered and followed by Claude or if it remains unused.
Compare two different phrasings of the same rule to see which version is consistently adhered to by Claude.
Edit the desired rule and ensure the changes remain uncommitted.
Execute the command /abdiff:abdiff <what you want to test> and respond to the questionnaire.
Access the report in .abdiff/<experiment>/report.html. Use !open <path> within Claude Code for easy navigation.
Demonstrations
Explore various experiments that showcase the capabilities of abdiff. Each demonstration links to an individual report detailing the experiment setup and the specific modifications made:
fluent-korean: Examines the impact of enabling a Korean output style on code-analysis documentation.
payment-docs: Analyzes whether Claude effectively reads and adheres to policy documents once introduced.
fastapi-architecture: Investigates if an architecture document influences the destination of new endpoint code.
doc-sentence: Assesses the effect of changing a specific number in an imported specification on the resulting code.
Important Considerations
The total number of runs executed is calculated as 2 × repeats × test cases, influencing timing and costs based on model and tasks.
Each run occurs in a temporary copy of the project. Note that while permission prompts are skipped, the copy is not a complete sandbox: Bash commands may interact with external files. For untrusted projects, consider utilizing allowlist mode as detailed in skills/abdiff/references/protocol.md.
User-level settings are not in effect; only project-contained files can act as variables.
Prerequisites include Claude Code CLI, git, and Python 3.9 or higher.
Comments
0Start the conversation
Share the first comment.