If you’re trying to figure out how to tell if code is AI generated, neat formatting and long comments won’t settle it. You can look for clues, but the strongest evidence comes from records of how the code was made. A detector score alone doesn’t prove who wrote it.
That’s frustrating if you wanted a quick yes or no. It also matters. A developer can write repetitive code, and an AI tool can produce a short function that looks completely ordinary.
There are two separate questions here. Did someone use AI? And can you trust the result? You may be unable to prove the first while still finding a clear answer to the second.
You can practice the second question in Diffective, a free debugging game. Each case asks you to check a coding agent’s claim against the evidence. Here’s how I’d work through both questions when reviewing actual code.
How to tell if code is AI generated using actual evidence
Start with the creation process, when you have access to it. An authorized chat record containing the generated function tells you more than the function’s indentation.
Useful evidence can include a contributor’s explanation, an AI tool’s recorded changes, or a saved conversation that matches the submitted code. Check what each record actually covers. A generated first draft doesn’t mean the final version was left untouched.
Version history can help you reconstruct when changes appeared. It cannot, by itself, tell you whether a person typed them, pasted them from a tutorial, or accepted an AI suggestion. A single large commit isn’t proof of AI use.
For a contractor or teammate, ask a neutral question about the process. Which tools did they use? What did they change afterward? What checks did they run? Keep the question separate from whether you like their coding style.
If you’re reviewing work against a school or workplace rule, use that rule and a fair review process. Don’t turn a guess about comments into an accusation.
Common signs of AI code are clues, not a verdict
Style can tell you where to look more closely. It gives you very little certainty about authorship.
| What you notice | Another possible explanation | A useful follow-up |
|---|---|---|
| Comments explain obvious lines | A beginner wrote them, or a team requires them | Ask why the approach was chosen |
| Names and formatting suddenly change | A formatter, template, or different contributor was involved | Compare the change with the project’s conventions |
| The solution adds unnecessary layers | A person overengineered it or adapted an example | Ask whether a smaller solution meets the requirement |
| A function calls a nonexistent method | Someone mistyped it or used outdated documentation | Check the installed library’s documentation |
| An explanation doesn’t match the code | The explanation may be stale or mistaken | Follow the actual inputs and outputs |
That last column is where the useful work happens. You can identify an invalid method or a misleading explanation without establishing who wrote the line.
Well-commented code can still be wrong. Messy code can still work. Neither appearance tells you whether AI was involved.
Can an AI code detector give you a reliable answer?
Treat a detector result as a signal that needs context. Its usefulness depends on what it was tested against and how closely that matches your code.
The 2025 Droid research found that the existing detectors it evaluated struggled to generalize beyond their training languages and coding domains. It also explored ways to improve detection. That supports caution, rather than a claim that detection can never work. Read the research.
Mixed authorship makes the question harder. A person might write a function, ask an assistant to simplify it, then rewrite the result. The CodeMirage benchmark includes original and paraphrased generated code to study detection under a wider range of conditions. See the benchmark.
Before relying on a detector, look for published evaluation details. Which languages and generators were included? How often did it flag human code incorrectly? Was it tested on edited code, or only untouched outputs?
A number labeled “AI likelihood” isn’t automatically a measured probability for your file. Without an explanation of how the score was calibrated, a precise percentage can imply more certainty than the evidence supports.
Don’t paste private client code, passwords, or access tokens into a public checker just to satisfy your curiosity. Use only a service approved for that material, with terms you understand.
Asking another chatbot doesn’t establish authorship
“Did AI write this?” can produce a confident explanation without a trustworthy answer. The chatbot may point to the same comments, naming patterns, and tidy structure that a human reviewer would notice.
A more useful request is to ask for specific problems in the code and evidence for each one. You can then check those claims. An explanation you can verify is more useful than a guess you can’t.
For example, ask which input breaks the function, which requirement it misses, or which library method needs checking. Don’t treat the resulting review as automatically correct either.
Check whether the code works, even when its origin is unknown
You can review a change without solving the authorship question first. Write down the expected behavior, inspect the changed lines, and run checks that could expose a mistake.
If you are deciding whether to rely on an AI-built project, my guide to the risks of vibe coding and what to check turns that review habit into a practical checkpoint.
GitHub’s own guidance recommends testing generated code, checking its fit with the project, and examining dependencies and unsupported assumptions. Those are useful review habits regardless of who produced the first draft. GitHub’s code-review guide.
Consider this made-up example. A developer says a report export now preserves every record. The test output says an export file was created successfully. That proves a file exists. It doesn’t prove the file contains every expected record.
A stronger check would compare the exported records with a small, known input set. Include a record with an empty optional field and one containing punctuation that needs escaping. Then open the result in the program people will actually use.
You still wouldn’t know whether AI wrote the exporter. You would know much more about whether the export claim holds up.
My review questions would be:
- What exact behavior was requested?
- Which changed lines implement it?
- What result would show that the change is wrong?
- Do the tests check that result, or only confirm that something ran?
- Which parts were not tested?
If you don’t understand a critical part of the code, ask someone who does before relying on it. A fluent explanation from the author or an assistant can guide your review, but it doesn’t replace checking the behavior.
For an agent’s completion report, Agent Receipt compares its claims with the test output and code changes you supply. It can help identify missing evidence, but it does not run the code or establish who wrote it.
Practice checking claims with Diffective
I built Diffective around the gap between an AI coding agent’s report and the evidence behind it. It’s a free browser debugging game you can play without an account.
A case gives you a task, a completion report with three claims, and code changes or terminal output. You pick the unsupported claim, select the evidence, and identify the trick. A guided tutorial explains how that works.
Diffective won’t tell you whether a stranger used AI. It’s designed to give you practice asking a more concrete question. Does this result support what the report says?
The Free AI Prompts page has reusable templates for checking claims and making requests clearer.
When you need to establish authorship, look for records and ask about the process. When you need to decide whether code is usable, inspect and test it. Keep those conclusions separate, and be willing to say “unknown” when the evidence doesn’t answer the first question.
Play Diffective and check the agent’s report. Start with the tutorial, then find the claim the evidence doesn’t support.



