Code coverage for AI agents

AI writes the code. Tests prove it works.

Agents open more pull requests than your team can read line by line. OtterWise checks that each change has tests before it merges, with patch coverage, status checks, and line annotations on every PR. We never need access to your source code.

No credit card required for public repos.

Why it matters now

Agents write more code than anyone can read.

Code review used to catch most problems. With agents writing code, there is now too much of it to review closely. Tests run on every change, however many there are.

01 · Volume

More code than anyone can read.

Each developer ships more code with an agent, but review time stays the same. When nobody reads every line, tests have to catch what review misses.

02 · Refactors

Agents change many files at once.

One prompt can touch twenty files. Regressions land in the files nobody opened, and a failing test finds them even when the reviewer skims.

03 · Feedback

Tests are how agents check their work.

An agent runs the tests to see if its change broke something. Where there are no tests, it has nothing to check against.

04 · Plausibility

Wrong code that looks right.

Generated code is tidy and well named, so bugs are harder to see by eye. A test runs the code instead of trusting how it looks.

05 · Compounding

Untested code gets built on.

Agents extend the code that is already there. One untested module becomes the base for the next ten changes.

06 · Familiarity

Your team knows the code less well.

When an agent writes most of a module, nobody on the team knows it by heart. The tests are the record of what it should do.

The blind spot

Agents are very good at gaming coverage.

Ask an agent to raise coverage and it will. It writes tests that run every line and check nothing useful. The percentage goes up, and the number means less. OtterWise also reports the metrics that are harder to fake.

Branch coverage

Branch coverage shows which conditions only ever went one way, even when every line ran.

Patch coverage

Measures only the lines a PR changed. Years of well-tested code cannot hide one new untested change.

CRAP score

Combines complexity with coverage to flag complex code that the tests barely touch.

Line annotations

Changed lines with no tests get marked in the PR diff, so reviewers know where to look first.

You still need to read the tests. These metrics tell you which ones to read first.

Patch coverage

Check if this change is tested.

When agents merge many PRs a day, project coverage hardly moves on any one of them. Patch coverage measures only the lines each pull request adds or changes, so a PR with no tests cannot hide behind a high project number.

Included on every plan, with project, line, and branch coverage.

Guardrails for agents

Thresholds are policy. No tests, no merge.

Set a minimum patch coverage and make the status check required in GitHub. Every agent then follows the same rule: it can merge only when its changes have tests. Line annotations list the exact lines without tests, so the agent knows what to fix next.

Required status checks on GitHub
Separate minimums for project and patch coverage
Fail when coverage drops more than you allow

Trend tracking

Is AI adoption quietly lowering your coverage?

One snapshot will not show a slow decline. OtterWise records project and patch coverage on every branch over time and sends a weekly digest. You can see if agent-written code arrives with tests or without them.

Ready when you are

Give agents more autonomy. Keep the proof.

Patch coverage, status checks, and line annotations on every pull request, from people or agents. No access to your source code. Free for public repos.