Tech • AI • Robotics • Game

VIDEO
ENFR

Building verification loops in Claude Code

7/10
AnthropicClaudeSeptember 25, 2026 at 11:13 PM3:07
Audio player
0:00 / 0:00

TL;DR

Claude Code can automate post-change verification by turning manual checks into reusable project skills, letting it test, fix, and re-test application behavior on its own.

KEY POINTS

Manual verification as the missing step

Automated checks such as tests, type checks, and linters can confirm code quality, but they do not prove a change behaves as intended in the product itself. Teams often still verify by hand, opening a web page, calling an API, or tapping through a mobile app to confirm the result.

Verification can be codified

Those manual review steps can be written into a project so Claude Code can execute them directly with tools such as a browser, a terminal, and an iOS simulator. When a check fails, the system can make another fix and rerun the same validation loop without waiting for human feedback.

The built-in verify skill as a starting point

A practical entry point is the built-in verify skill. On first use, it runs the application and checks the change inside the app itself, then saves the successful steps as a reusable skill within the project for future runs.

Teams are expected to extend the generated checks

The generated skill is meant to be expanded rather than treated as final. Developers can define when Claude should run a check, what action it should take after a failure, and what concrete evidence counts as a pass.

Measurable checks improve reliability

The more objective a verification rule is, the easier it is for the system to determine success. Examples include a performance budget, an accessibility checklist, or design-system rules that can be checked consistently rather than judged informally.

Layout shift is a strong example

On web applications, one recurring manual check is layout shift, where content jumps as a page loads. That behavior can be measured through a performance trace using Google Chrome DevTools MCP, which records layout shift as part of Core Web Vitals.

A UI example shows the full loop

In one example, Claude was asked to add a Like button to a page that also had an unresolved layout shift issue. Because the change affected the user interface, it automatically ran the verification skill, started the development server, opened the page, clicked the new button, and captured a screenshot as proof the feature worked.

Verification can catch adjacent problems

The same run also executed a performance trace and detected layout shift unrelated to the button itself. Claude then fixed the issue, reran the checks, and returned both a working feature and evidence that the page no longer shifted on load.

The payoff is fewer feedback rounds

Automating this verify-and-retry cycle reduces the need for a person to manually inspect the result and describe what needs to change. The outcome is a tighter development loop, fewer rounds of back-and-forth, and stronger proof that a change meets both functional and quality expectations.

CONCLUSION

The broader goal is to move routine validation from human observation into measurable project rules. As more checks are codified, Claude Code can operate more independently while producing changes that are easier to trust.

Ask a question

More from Anthropic