How to Review AI-Generated Code Before You Merge

Your agent says the fix is done. Compare that claim with the diff and your own test run, then use a practical checklist to decide what needs another look.

Your coding agent says the bug is fixed. The test summary is green. You’re one click away from merging, but you still can’t explain what changed or which test proves the fix.

To review AI-generated code, start with three things: the agent’s claims, the actual code changes, and a test run you can inspect. Compare them before deciding what to accept. A confident completion message is useful as a checklist of claims, but it isn’t evidence by itself.

Agent Receipt, our free browser tool, helps with that comparison. Paste the agent’s message, your test output and the diff. It checks claims against the supplied lines. It does not run your code, inspect your whole repository or certify that a change is safe to merge.

Start by turning “done” into specific claims

A message such as “fixed the checkout and all tests pass” bundles several questions together. Did the intended file change? Did the relevant test run? Did it pass? Does the original checkout problem behave differently now?

Separate those questions. For each claim, write down what would support it:

  • “I updated the total calculation.” Find the changed calculation in the diff, then read what it does.
  • “All tests pass.” Inspect the actual result, including skipped tests and errors that appear before the final summary.
  • “Nothing else changed.” Compare that claim with the complete changed-file list.
  • “The bug is fixed.” Reproduce the behavior that failed, using the changed version of the app.

This separates an easy-to-check fact, such as a file appearing in the diff, from a larger claim about behavior. Finding a file proves very little about whether its new code is correct.

A worked example: 58 tests that didn’t all pass

Open Agent Receipt and choose Try an example, then Make my receipt. The built-in example is fictional and is marked as an example on the receipt. It describes an agent claiming to fix a checkout rounding bug.

The message says “All 58 tests pass.” The supplied test summary says:

Tests 57 passed | 1 skipped (58)

The changed test is marked it.skip. That matters because the test meant to check the rounding case is present but does not execute. The receipt labels the all-tests-pass claim Contradicted.

Two other details change the review:

  • The agent says no other files changed, but the diff includes config/staging.env. That claim is also contradicted. The configuration change needs an explanation.
  • The agent claims a clean linter run, but the paste contains no linter output. That is No proof, not proof that linting failed. Some clean commands are quiet, so keep the command and exit result when checking them yourself.

The receipt also shows Proof shown for the narrow check that the named calculation file appears in the changes. It labels the broader claim that the checkout bug is fixed Can’t tell. A pasted diff cannot demonstrate a working checkout.

The useful next step is specific: remove the skip, run the relevant test again, inspect the configuration change, and try the original rounding case. Asking the agent “Are you sure?” would give it another chance to repeat the claim without supplying the missing evidence.

Read the diff, including what disappeared

A diff shows what was added and removed. Open it in your editor or Git client and compare its scope with the task you assigned. A small requested fix can still arrive with unrelated changes.

Start with the changed-file list. Then inspect the lines that decide behavior: conditions, calculations, permissions, error handling and calls to other services. If a line is unfamiliar, ask for an explanation tied to that exact line and check it against the surrounding code.

Pay particular attention to deleted or weakened tests. A failing test can disappear without the defect disappearing. GitHub’s guidance on reviewing AI-generated code explicitly calls out skipped or deleted tests, incorrect logic and invented APIs as things to inspect. It also recommends tests, static analysis and checking new dependencies.

For a dependency change, check the real package and its documentation before relying on the agent’s explanation. For a configuration change, establish why the new value belongs in this fix. If you cannot connect a change to the requested outcome, leave it unresolved until you can.

Check what the test result actually covers

A passing run answers a limited question: the checks that executed passed under those conditions. It does not show that every important behavior was tested.

Use your project’s documented test commands in the appropriate development or test environment. Keep the command, output and exit status together. Make sure you are looking at the version you intend to merge, rather than an old run from before the latest edit.

For a bug fix, inspect the test that is supposed to catch the bug. Ask what input it uses and what result it checks. Then choose a small set of cases that matters for the actual change:

  • The original failure: does the specific reported problem now behave as required?
  • A nearby boundary: what happens just below, at and above a relevant limit?
  • A failure path: what happens when an expected file, response or input is missing?
  • An existing behavior: does something that used to work still work?

These are prompts for choosing checks, not a claim that four tests are enough for every change. A text adjustment and a payment calculation deserve different levels of review. For consequential changes you don’t understand, get a qualified reviewer before merging.

A second AI can help suggest missing cases, but inspect its reasoning too. GitHub documents that Copilot reviews can miss problems or flag problems that aren’t there. Another model’s agreement should not become the final evidence.

Use Agent Receipt to organize the next question

Agent Receipt offers one combined paste box and a Paste separately option. Supply the completion message, the output from your test run and the code changes. If the combined paste is split incorrectly, use the separate fields.

Read each label together with its evidence and stated limit:

  • Proof shown: a supplied line supports the part being checked. This is narrower than “the code works.”
  • Contradicted: supplied evidence conflicts with the claim.
  • No proof: the required support is absent from the paste.
  • Can’t tell: the paste or the tool’s rules cannot settle the claim. Check it another way.

The tool cannot confirm that your test output came from the exact code in your diff. You still need to establish that connection. It also cannot inspect files you didn’t supply, authenticate the agent’s history or approve a release.

Use the receipt’s follow-up message as a starting point. Review it before sending it back to your agent. If you copy evidence into another AI service, check that you are allowed to share it there and remove secrets or private data.

A review prompt you can reuse

This prompt works best when you provide the actual requirement, changes and results. It asks for a review before further edits so you can inspect the findings first.

Review this change against the requirement below. Do not modify files yet.

For each claim in the completion message, identify the exact supplied evidence that supports or contradicts it. Separate a changed file, an executed test and verified user behavior. If evidence is missing, say what is missing.

Inspect the diff for unrelated changes and tests that were removed, skipped or weakened. Suggest the smallest useful next check for each unresolved issue. Do not claim you ran a command unless you actually ran it and can show the result.

Requirement: [what should happen]
Completion message: [agent’s message]
Changes: [diff]
My test run: [command, output, exit status and version tested]

For broader decisions before you act, browse our Free AI Prompts. Use a prompt to ask better questions, then verify the answers with evidence you can inspect.

Before you merge

  • I can explain the intended behavior and the important changed lines.
  • The changed-file list matches the task, or every extra change has a clear reason.
  • I inspected the relevant tests, including skips, removals and assertions.
  • The test output belongs to the version I am reviewing.
  • I checked the original problem and the important nearby cases.
  • I know which claims remain unresolved and who needs to review them.

If you want to practise spotting these gaps before reviewing your own project, play Diffective. For the broader risks of building by prompting, see Is Vibe Coding Bad?

For your next real “done” message, make an Agent Receipt. Use it to find the next question worth asking. Make the merge decision from the code, the actual checks and the behavior you can verify.

SoftDeveloper23
SoftDeveloper23

I’m the maker behind softDev23, building apps and exploring how AI and automation can make everyday work easier. I share practical guides and lessons from building in public: what worked, what broke, and what I’d do differently.

Follow along as I turn ideas into useful products, one experiment at a time.

Articles: 126

Leave a Reply

Your email address will not be published. Required fields are marked *