Controlled tests showing what Pallos catches, flags for review, and cannot establish. These two reproducible cases use a pinned commit of OWASP Juice Shop, an intentionally insecure training app. They are not customer results or a full-repository benchmark.
This is a selected rule demonstration, not a measured detection rate. Missed findings elsewhere and false-positive rate have not been measured. Source evidence cannot verify deployed exploitability.
Known issue / expected behavior: Avoid a high-priority injection finding because the expression is generated from fixed operators; a low-priority review note would be enough.
Pallos result: Dynamic code execution needs review. Classification: Partial within this two-case sample.
The expression is assembled from generated numbers and a fixed list of arithmetic operators. Pallos did not see an obvious user-controlled source in this file, so it does not label this as a confirmed injection.
Limit: This is a source-level observation, not a runtime test. Other files or configuration could change the trust boundary.
Known issue / expected behavior: Identify the visible user-data flow into eval and give a concrete fix direction.
Pallos result: User-derived value reaches dynamic code execution. Classification: Correct within this two-case sample.
The code assigns `user.username` to `username`, derives `code` from that value, then passes `code` to `eval`. The route only takes this branch when the named training challenge is enabled.
Limit: Static analysis shows the source-level flow; Pallos did not run the app or verify the challenge's deployed configuration.
What this demonstrates
Confidence follows the evidence.
The same broad rule can produce a review signal or a higher-confidence finding depending on visible source context. Severity describes potential impact; confidence describes the evidence found. Neither establishes runtime exploitability.