Product
Scanner Reports AI Integration
Resources
Guides Research Security checklist Docs Status
Account
Sign in Scan your app free
Guides/Guide
Guide

From finding to fix: using AI to repair its own code

Why the same AI that helped create vulnerable code can still be useful for repairing it — once a separate security system has found the problem first.

There is an obvious contradiction at the heart of AI-generated software.

If an AI coding tool wrote the vulnerable code in the first place, why would you give the vulnerability back to AI and ask it to repair it?

For a non-coder, the question matters even more. If you built an application through Claude Code, Cursor, Lovable, Replit or another AI coding tool, you may not be able to inspect a security finding yourself, understand the vulnerable code and manually rewrite it. The coding AI may be the only practical way you have to change the application.

That sounds dangerous until you separate three very different jobs: building software, finding vulnerabilities and repairing known vulnerabilities.

An AI agent can struggle to notice an unstated security requirement while it is simultaneously designing an application, writing features, connecting a database and trying to satisfy a user's requests.

That does not mean the same system is incapable of changing a specific piece of code after another system has identified exactly what is wrong.

The difference is diagnosis.

Instead of asking an AI:

“Make my application secure.”

you can give it a concrete security finding:

“This endpoint allows an authenticated user to access another user's records because ownership is never checked server-side. The affected route is here, and this is the code path that causes it.”

The first problem asks AI to discover an unknown vulnerability somewhere inside an application.

The second asks it to solve a bounded engineering problem.

Research into AI-generated code, real-world vibe-coded applications and automated vulnerability repair increasingly suggests that this distinction matters enormously.

The emerging model is not AI builds → AI says everything is fine.

It is:

Build → independently scan → give the finding to the coding AI → test → rescan.

That is the idea behind Epherem's role in the development loop.

Epherem does not need to become the coding agent. It needs to show the coding agent — through the user — what looks wrong in the code, where, and how likely it is to be a real problem.

And there is a surprisingly strong research case for doing exactly that.

AI is much better at making code work than making it secure

AI coding has become extraordinarily good at producing software that appears to work.

That creates a dangerous illusion.

A broken feature tends to reveal itself. The button does nothing. The request fails. The page crashes.

A broken security property often does not.

A database query can return exactly the right information to the legitimate user while also allowing another user to retrieve it.

An API route can behave perfectly while missing an authorization check.

An application can successfully connect to its backend while exposing a privileged credential to the browser.

Everything the builder tests may work.

The vulnerability exists in what an attacker can do instead.

Veracode has been measuring this gap across generations of coding models. Its 2025 GenAI Code Security Report tested more than 100 models across Java, Python, C# and JavaScript and found that 45% of generated code samples failed its security tests. JavaScript failed 43% of the relevant tasks. Cross-site-scripting protection was particularly poor, with models failing 86% of applicable tests.5

The important part is what happened as the models improved.

Their ability to produce syntactically correct code became extremely strong. By Veracode's Spring 2026 testing, syntax correctness exceeded 95%.

Security did not follow it.

Across more than 150 models tested at that point, the average security pass rate remained around 55%. Veracode's subsequent 2026 report put the figure at roughly 56%.5

That is the distinction a vibe coder needs to understand: code that works is not the same thing as code that is secure.

And better coding models do not automatically make the distinction disappear.

We now have evidence from actual vibe-coded applications

Tests on isolated coding tasks can only tell us so much.

In June 2026, researchers published Understanding the (In)Security of Vibe-Coded Applications, one of the most important studies of this problem so far.

Instead of giving models artificial coding exercises, the researchers assembled 10,517 real-world applications developed using AI coding agents.

They then took a random sample of 200 publicly deployed web applications and performed an agent-assisted security audit followed by human validation.

The results were severe.

The researchers confirmed 1,471 vulnerabilities. 180 of the 200 applications — 90% — contained at least one exploitable vulnerability.

Among vulnerable applications, the median was seven vulnerabilities. And 76.7% of all confirmed vulnerabilities were rated High or Critical.1

The vulnerability distribution is important too.

Broken Access Control alone accounted for 36% of the vulnerabilities. Cryptographic failures and injection vulnerabilities made up another large share.1

These are precisely the kinds of issues that a non-coder may never notice during normal use.

Imagine an application where you can open /api/orders/123 and correctly see your order.

From the builder's perspective, the feature works.

The important security question is whether changing 123 to another customer's order ID still returns data.

That is not normally part of “does my app work?”

It is part of “does my app enforce its trust boundaries correctly?”

AI coding has dramatically reduced the amount of technical knowledge required to implement the first question.

It has not removed the second.

Many vulnerabilities are caused by requirements nobody said out loud

The vibe-coding study also investigated why these vulnerabilities appeared.

This may be even more important than the 90% headline.

The researchers grouped vulnerabilities into problems involving memory, objectives and security knowledge. About 49.9% of the vulnerabilities were attributed to knowledge defects.1

The single largest failure mode was what the researchers called hidden security rules, accounting for 43.9% of vulnerabilities.1

A hidden security rule is something the user never explicitly requested because it is normally assumed.

A user might ask:

“Let users upload profile pictures.”

They probably do not also say:

“Validate the file type server-side, prevent executable content from being served unsafely, enforce storage authorization, constrain file size and ensure one user cannot overwrite another user's files.”

Or:

“Build an endpoint to edit a project.”

They do not necessarily add:

“Verify on the server that the authenticated user owns or has permission to modify the project instead of trusting the project ID supplied by the browser.”

To an experienced security engineer, those requirements may be implicit.

To a non-coder prompting an AI, they may be invisible.

And to the coding agent, the immediate objective is usually to make the requested feature work.

This is one reason security cannot depend entirely on asking users to write better prompts.

The user cannot reliably specify security requirements they do not know exist.

Why “review your own code and make it secure” is not enough

A straightforward response would be to send the whole repository back to the coding AI and ask it to perform a security review.

There is a problem with that approach too.

SecureAgentBench, published in 2025, evaluates AI coding agents using 105 repository-level tasks based on real vulnerabilities. The tasks require agents to make realistic changes in large codebases while being evaluated for both functionality and security.

The best agent-model combination achieved a 15.2% correct-and-secure solution rate.

The study also found cases where agents produced code that worked functionally but remained vulnerable or introduced new vulnerabilities.

Most importantly, adding explicit security instructions did not significantly solve the problem.3

That does not mean AI cannot reason about security.

It means the task “implement this feature securely” is harder than it sounds.

The model has to understand the repository, determine which security requirements apply, locate every place affected by them, preserve existing behavior and write the implementation — all at once.

So instead of asking the coding model to be simultaneously developer, security auditor and final judge of its own work, there is a stronger architecture: separate diagnosis from repair.

A good finding fundamentally changes the AI's task

This is where the case for the Epherem workflow becomes much stronger.

Imagine handing an AI a 10,000-line repository and asking:

“Find the security problem.”

Now compare that with:

“There is a broken authorization check in this route. This is the affected code. User-controlled resource IDs reach the database lookup without an ownership check, allowing one authenticated user to request another user's record.”

These are not remotely equivalent tasks.

The second removes an enormous search problem.

Research into automated vulnerability repair has measured this directly.

PATCHEVAL, a benchmark published in late 2025, assembled 1,000 real vulnerabilities across 65 CWE categories in Go, JavaScript and Python. For 230 of them, researchers constructed executable environments so patches could be tested for both security and functionality.2

One experiment examined what happens when AI models are told where the vulnerability is.

Without localization guidance, models averaged 46.3 successful repairs across the tested set.

With precise localization, the average rose to 54.3.

Then the researchers deliberately made the location inaccurate.

When the supplied location was more than five lines away from the true vulnerable code, successful repairs fell to 42.7 — worse than giving the model no location at all.

2

That is a remarkably important result for a product designed around passing security findings into coding AI.

Correct context helps. Incorrect context actively hurts.

A security finding is therefore not merely something for the user to read.

It becomes part of the input to another AI system.

Finding quality directly affects repair quality.

Good guidance helps; bad guidance is disastrous

In August 2026, 1Password's Off-by-1 Labs published a large investigation into AI-generated vulnerability patches.

The researchers evaluated more than 6,000 patches generated for six recently disclosed, complex vulnerabilities.

One part of the study examined what happened when models received different kinds of initial guidance.

With correct guidance, the models achieved a 65.0% fix-success rate. With no guidance, they achieved 50.4%. With incorrect guidance, success collapsed to 15.2%.

4

That last number may be the most consequential.

The model did not reliably notice that it had been given the wrong diagnosis and recover.

It followed the premise it had been given.

This creates an extremely important design principle for AI-era security: the quality of the diagnosis matters as much as the ability of the model performing the repair.

This is also why a security product aimed at non-coders cannot simply maximize the number of warnings it produces.

If an incorrect warning is going to be handed directly to a coding agent, the cost of a false positive is no longer just annoyance.

The AI may start changing safe code because it has been confidently told that the safe code is vulnerable.

Precision matters. Evidence matters. Localization matters.

And uncertainty should remain uncertainty instead of being converted into a confident instruction.

Why Epherem points to the problem rather than prescribing the patch

There is an important consequence here for Epherem's design.

Epherem does not need to generate a pre-written fix prompt.

It does not need to tell the coding AI exactly how to rewrite the application.

That would move Epherem from diagnosis into implementation, where it has less context than the coding agent already working inside the user's repository.

Each finding in an Epherem report does include a suggested fix and a prompt for your coding agent, but treat both as a starting point: the finding and where it sits in your code are what matter, and your coding agent decides how to change the code.

Instead, the useful security finding answers a narrower set of questions:

What is wrong? Where is it happening? What evidence shows that it is happening? Why does it matter? How confident are we that the problem is real?

That is enough to transform the next AI interaction.

The coding agent can inspect the surrounding repository itself and decide how the application should be changed.

That division of responsibility is useful because security scanners and coding agents have different advantages.

The security system is explicitly looking for violations.

The coding agent has deeper context about how the application is supposed to function and can modify multiple related files when necessary.

The finding becomes the interface between them.

This is especially powerful for non-coders

Traditional security tooling was largely built around a developer being on the receiving end.

A scanner reports CWE-639: Authorization Bypass Through User-Controlled Key and assumes that somebody knows what to do next.

For a security engineer or experienced backend developer, that may be enough.

For someone who created an application through natural language, it may mean almost nothing.

That is a fundamental mismatch.

Vibe coding allows someone to create software without understanding every implementation detail, yet most application-security workflows still assume the person fixing vulnerabilities does understand those details.

The finding-to-AI workflow changes the user's job.

They do not have to look at a route handler and determine the correct authorization architecture themselves.

They need to understand the security issue at a high level, then give the diagnosed problem back to the coding system that already knows how to edit their application.

That is a much more realistic workflow.

The non-coder becomes the person connecting two specialized systems: the security system identifies the problem; the coding system changes the code.

This does not magically make the user a security engineer.

That is the point.

The catch: AI-generated fixes are not trustworthy by themselves

So far this sounds like an argument for AI remediation.

It is.

It is not an argument for autonomous AI remediation.

The same 1Password research provides the warning.

Only 26.0% of patches completely resolved the vulnerability without materially altering application behavior. Another 20.1% removed the vulnerability but changed application behavior in the process. The remaining 53.9% either failed to fully resolve the vulnerability, introduced another vulnerability, or did both.

4

Even patches classified as successful could be fragile. The researchers found that more than a third of successful categories contained security subtleties where the model had blocked a particular exploit input without necessarily eliminating the underlying weakness.4

That is why the workflow cannot end with: Epherem flagged it → AI says it fixed it → deploy.

The AI's statement that it fixed the vulnerability is not proof that the vulnerability is gone.

The code has changed.

That change has to be checked again.

The real system is a feedback loop

A safer model looks like this:

Build → Scan → Repair → Test → Rescan.

First, the coding agent creates or changes the application.

Then Epherem analyzes the resulting source code independently of the generation process.

When it flags a potential security issue, the user takes that finding back to their coding agent.

The coding agent investigates the affected code and implements a fix.

The affected feature is tested to make sure normal behavior still works.

Then the updated source is scanned again.

If the original finding survives, the repair was not sufficient.

If the change creates another detectable vulnerability, the new scan gets another chance to surface it.

If the finding no longer appears, the report shows that part of the code was checked, and the application still works, confidence in the repair has increased.

Not certainty. Confidence.

That distinction is important in security.

Rescanning matters because AI should not certify its own answer

Suppose you ask a coding agent:

“Fix this authorization vulnerability.”

It edits three files and replies:

“Done. The endpoint now verifies resource ownership before returning the record.”

That explanation can sound convincing.

But it was produced by the same system that produced the patch.

Asking it whether its own fix is correct provides weaker evidence than independently examining the result.

The point of rescanning is to inspect the new state of the application instead of trusting the model's description of what it changed.

This aligns with a much older principle in secure software development: security should be verified throughout the development lifecycle rather than treated as a one-time activity.

NIST's Secure Software Development Framework recommends integrating practices for identifying vulnerabilities, addressing their root causes and preventing recurrence throughout software development.8

The underlying idea predates vibe coding.

What changes with vibe coding is who — or what — performs each step.

Iteration makes AI repair substantially stronger

PATCHEVAL produced another result that supports the loop.

The researchers allowed models to try repairing vulnerabilities repeatedly while receiving feedback from security and functionality tests.

For Gemini 2.5, the number of successfully repaired vulnerabilities increased from 52 initially to 134 after iterative feedback, with improvements eventually beginning to plateau.

2

That is not proof that simply rescanning any application will reproduce the same improvement.

It does establish something broader: AI repair can become dramatically more effective when it receives external feedback about whether its previous attempt actually worked.

That is exactly what a scan-repair-rescan architecture is designed to provide.

Instead of one enormous prompt asking AI to understand and secure the entire application in one pass, security becomes an iterative process.

A system identifies a problem.

The coding AI attempts a repair.

The application is checked again.

If necessary, another iteration begins.

Security becomes a feedback loop rather than a promise in the original generation prompt.

Finding accuracy is part of remediation quality

Traditional static-analysis products have always cared about false positives because developers hate wasting time on alerts that are not real.

AI repair makes the stakes higher.

The PATCHEVAL experiment showed that inaccurate localization could perform worse than providing no location.

The 1Password study found that incorrect guidance could reduce fix success from 50.4% without guidance to just 15.2%.

Together, these results suggest a simple rule: bad security context can be worse than no security context.

That has major implications for Epherem.

A finding should not be treated as successful merely because a rule fired.

The useful output is not the largest possible pile of suspicious code.

It is the smallest defensible set of real security problems, supported by enough evidence that both the user and their coding AI are being sent in the right direction.

This is why verification matters before remediation begins.

It is also why uncertain cases should be allowed to remain uncertain.

A system that says “this needs review” when the evidence is ambiguous is more useful than one that confidently invents an explanation and sends an AI agent off to modify the wrong code.

For vibe coders, scanner precision is not just a product-quality metric.

It is part of the safety of the repair process itself.

Epherem’s free scan does not confirm its code findings. It shows everything its checks flag, in two groups: “likely issues”, where the scanner saw the risky code itself, first; then “needs checking”, where it looked for a protection and didn’t see it, or the match depends on context it can’t see. Only published advisories for your package versions are presented as known issues. That way uncertain findings stay marked as uncertain before they reach your coding agent.

The advantage of a different system inspecting the code

There is another reason this architecture makes sense.

The original coding agent may be carrying assumptions from the conversation that produced the software.

It knows what the user wanted.

That is useful for implementation.

It can also create blind spots.

Perhaps an authentication shortcut was introduced earlier because the user was struggling to get login working.

Perhaps the agent implemented a temporary workaround.

Perhaps a requirement was changed three sessions ago.

Perhaps a new route was added without propagating an authorization pattern used elsewhere.

The 2026 study of real vibe-coded applications found exactly these kinds of systemic failures: forgotten obligations, incomplete propagation of security requirements, demo-oriented design and security rules that were never made explicit.1

A fresh source-code analysis does not have to remember what the coding agent intended.

It examines what actually exists.

That is a useful separation.

Intent belongs to the coding agent. Evidence belongs to the security analysis.

Non-coders also need an external stopping rule

There is a subtler problem with using AI to build software you cannot personally inspect.

How do you know when the AI is right?

The 2025 Stack Overflow Developer Survey found that 84% of respondents were using or planning to use AI tools in development, yet more developers actively distrusted AI accuracy than trusted it: 46% versus 33%. Only 3% said they highly trusted AI output.

The largest frustration, reported by 66% of developers, was receiving AI solutions that were “almost right, but not quite.”7

Security makes “almost right” particularly dangerous.

An authorization fix that blocks nine unauthorized paths but leaves the tenth open is still vulnerable.

A secret that is removed from one frontend file but remains in another is still exposed.

An injection patch that filters the proof-of-concept payload but leaves the underlying unsafe operation reachable is not a robust repair.

For someone who cannot manually validate the implementation, the coding AI saying “fixed” cannot be the stopping rule.

There needs to be another signal.

That is another purpose of the rescan.

Why the workflow can never guarantee security

There is an important boundary to draw.

No source-code scanner finds every vulnerability.

No AI verification layer perfectly separates every true finding from every false one.

No coding agent reliably patches every vulnerability.

And a clean rescan does not prove that an application contains no vulnerabilities whatsoever.

Some problems depend on production configuration, runtime state, infrastructure or business logic that static source analysis may not fully understand.

OWASP explicitly treats automated testing and manual secure-code review as complementary. Human expertise remains particularly important for complex authorization models, business logic, cryptography, architecture and other context-heavy security decisions.9

That matters most for applications handling high-risk data or operations.

Payments, sensitive personal information, cryptographic systems, healthcare or financial data, complex permissions and other high-impact functionality deserve a higher level of scrutiny than an ordinary hobby application.

The realistic promise of this workflow is therefore not:

“Give Epherem your app and it becomes secure.”

It is:

“Give a non-coder a practical way to discover security problems, get them repaired using the development tool they already understand, and check the resulting code again.”

That is a much more defensible goal.

And it is still a significant change.

Finding vulnerabilities is only useful if someone can act on them

Application security has historically had a remediation problem.

Veracode's 2025 State of Software Security research reported that the average time to fix security flaws had risen to around 252 days, while the typical organization took roughly five months to fix half of its detected flaws.6

Those are enterprise figures and should not be interpreted as the expected remediation time for a vibe coder.

They illustrate a broader problem instead: finding vulnerabilities and fixing vulnerabilities are separate bottlenecks.

A scanner that produces findings nobody can act on only solves half the problem.

For experienced development teams, the missing bridge might be engineering time.

For a vibe coder, it can be technical ability itself.

That makes AI unusually valuable on the remediation side.

The same abstraction layer that allowed the user to create the application can become the interface through which they modify it.

The security tool does not need to teach the user how to program first.

It needs to communicate the problem accurately enough that the coding system can act on it.

The security finding becomes a new kind of interface

For most of software history, a vulnerability report was written for a human developer.

AI changes that.

A modern finding can be useful to two audiences simultaneously.

The human needs to know:

What happened? How serious is it? What could an attacker do? Where is the problem?

The coding AI needs almost exactly the same foundation:

What security property is broken? Which code is involved? What evidence demonstrates the problem? What behavior must remain protected?

Notice what is missing.

The security system does not necessarily have to prescribe the precise implementation.

That is why Epherem's job does not need to be generating copy-paste fix prompts.

The finding itself is the handoff.

A non-coder can take it into the coding environment they already use and say, effectively:

“Investigate and fix this finding without breaking the application's intended behavior.”

The coding agent can then read the actual repository, inspect related files and implement the change using its much broader project context.

This creates a clean division of labor:

Epherem determines what appears to be wrong. The coding agent determines how to change the application. Testing and rescanning determine whether that change actually improved the result.

The paradox disappears once the jobs are separated

So can AI be trusted to fix vulnerabilities in AI-generated code?

Not by itself.

That is the wrong question.

A better question is: can an AI coding system be useful for implementing repairs after another system has given it accurate security information, if the result is then independently checked?

The evidence increasingly points toward yes.

AI-generated software still has serious security weaknesses. Real-world vibe-coded applications have shown extremely high vulnerability prevalence. Generic instructions to “write secure code” are not enough. Autonomous patches frequently fail. And incorrect remediation guidance can make performance catastrophically worse.

But the other side of the evidence matters too.

Precise vulnerability localization improves AI repair.

Correct security guidance substantially outperforms incorrect guidance.

External test feedback can dramatically improve iterative repair.

And coding models are already capable of modifying large applications when they are given a concrete engineering objective.

Those findings lead to a different philosophy of AI security.

Do not ask AI to blindly trust itself.

Give it an external critic.

Less a replacement developer, more a feedback system

The most interesting consequence of vibe coding is not simply that AI writes more code.

It is that the traditional relationship between the person building software and the code itself has changed.

A non-coder can now own an application containing thousands of lines of TypeScript, database policies, authentication logic, server functions and third-party dependencies without being able to manually review most of them.

That creates what is effectively a comprehension gap between the person responsible for the application and the implementation they are shipping.

It would be unrealistic to solve that gap by requiring every vibe coder to become a security engineer.

A more scalable approach is to put security checks into the same abstraction-based workflow that allowed them to build the software in the first place.

The user asks AI to build.

Security analysis checks what was actually built.

The user gives concrete findings back to AI.

AI changes the implementation.

The resulting code is checked again.

That is not autonomous security.

It is AI-assisted remediation with independent feedback.

And that distinction matters.

From finding to fix

Vibe coding has removed an extraordinary amount of friction from software creation.

It has not removed the rules that make software safe.

Authorization still needs to happen on the server.

Sensitive credentials still cannot be exposed to the browser.

Untrusted input is still untrusted input.

Dependencies still acquire known vulnerabilities.

Database policies still need to enforce the application's intended boundaries.

The user may never see any of those implementation decisions.

But attackers can.

That is why the security layer cannot disappear just because the programming layer became invisible.

It has to become easier to use too.

Epherem's role in that model is deliberately narrow.

It analyzes the source code and surfaces potential security problems that someone who does not know how to review a codebase is unlikely to discover alone.

The user can take those findings back to Claude Code, Cursor, Lovable, Replit or whichever coding agent they already use.

That agent can repair the code.

The application can then be tested and scanned again.

No single participant has to do everything.

The coding AI builds. Epherem checks. The coding AI repairs. The result is checked again.

The same AI that wrote insecure code does not suddenly become trustworthy because it has been asked nicely to make the application secure.

What changes is the problem it has been given.

Instead of being asked to discover an unknown security flaw somewhere inside thousands of lines of code, it receives a concrete, evidence-backed issue to investigate and repair.

Research suggests that distinction is consequential.

Precise guidance can improve repair.

Bad guidance can make it dramatically worse.

And external feedback can make repeated attempts much more successful.

That makes the quality of the security finding the critical bridge between AI-generated software and AI-assisted repair.

Find the problem accurately.

Let the coding agent work on the bounded problem.

Test what it changed.

Scan the result again.

For non-coders building real software with AI, that may be a far more practical security model than expecting either the user or the coding AI to get everything right the first time.

Research referenced

  1. 1
    Understanding the (In)Security of Vibe-Coded Applications — Deng, Fan and Meng, 2026.
  2. 2
    PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities — Wei et al., 2025.
  3. 3
    SecureAgentBench: Benchmarking Secure Code Generation under Realistic Vulnerability Scenarios — Chen et al., 2025.
  4. 4
    Why AI-generated vulnerability patches still require expert human review — 1Password Off-by-1 Labs, August 2026.
  5. 5
    2025 GenAI Code Security Report and 2026 GenAI Code Security Report — Veracode.
  6. 6
    2025 State of Software Security — Veracode.
  7. 7
    2025 Stack Overflow Developer Survey — Stack Overflow.
  8. 8
    Secure Software Development Framework (SSDF) Version 1.1 — NIST.
  9. 9
    Secure Code Review Cheat Sheet — OWASP.

Scan first, then hand it over

Epherem’s free scan lists likely issues and findings that need checking in your code, with the file, the line, what someone could do, a suggested fix and a prompt, ready to take back to your coding agent. It also names what it couldn’t check.

Scan your app free