Audit History
llm-evaluation - 7 audits
Version comparison
Capability and finding changes across audited versions, newest first.
| Version | Date | Result | Review items | Change vs previous |
|---|---|---|---|---|
| v7 Latest | Jul 7, 2026, 08:07 AM | No confirmed findings | 0 | No capability change |
| v6 | Jul 7, 2026, 08:07 AM | No confirmed findings | 0 | External commands Network access |
| v5 | Jul 1, 2026, 12:28 AM | 1 confirmed | 0 | Network access External commands |
| v4 | Jan 17, 2026, 08:09 AM | No confirmed findings | 0 | No capability change |
| v3 | Jan 17, 2026, 08:09 AM | No confirmed findings | 0 | External commands |
| v2 | Jan 4, 2026, 04:38 PM | No confirmed findings | 0 | No capability change |
| v1 | Jan 4, 2026, 04:38 PM | No confirmed findings | 0 | Baseline |
Jul 7, 2026, 08:07 AM
The static external-command findings are markdown Python code fences in SKILL.md, not executable Ruby or shell backtick use. The certificate/key and network reconnaissance findings are false positives from ordinary evaluation text and resource names. No prompt injection attempt, data exfiltration intent, or malicious behavior was found.
Risk Factors
⚙️ External commands (23)
Jul 7, 2026, 08:07 AM
The static external-command findings are markdown Python code fences in SKILL.md, not executable Ruby or shell backtick use. The certificate/key and network reconnaissance findings are false positives from ordinary evaluation text and resource names. No prompt injection attempt, data exfiltration intent, or malicious behavior was found.
Risk Factors
⚙️ External commands (23)
Jul 1, 2026, 12:28 AM
Static analysis reported many high and medium findings, but review shows they are false positives from markdown code fences, metric labels, form scale labels, and dictionary key iteration. The skill is documentation-only, with one low-risk caution: example LLM-as-judge code would send evaluation inputs and outputs to an external API if users copy it into their own application.
Confirmed security concerns (1)
Static false positives ignored (4)
These static matches were dismissed by semantic review or matched schema-only tokens, so they are shown for transparency but do not drive the quality score.
Risk Factors
🌐 Network access (2)
Jan 17, 2026, 08:09 AM
This skill contains only static documentation (SKILL.md) with no executable files. All static findings are false positives: markdown code block backticks were misidentified as Ruby/shell command execution, and JSON metadata fields were misclassified as cryptographic issues. The skill provides evaluation guidance only with no data access, network activity, or command execution capability.
Risk Factors
⚙️ External commands (23)
Jan 17, 2026, 08:09 AM
This skill contains only static documentation (SKILL.md) with no executable files. All static findings are false positives: markdown code block backticks were misidentified as Ruby/shell command execution, and JSON metadata fields were misclassified as cryptographic issues. The skill provides evaluation guidance only with no data access, network activity, or command execution capability.
Risk Factors
⚙️ External commands (23)
Jan 4, 2026, 04:38 PM
Pure documentation skill containing no executable files or tooling. All content is static markdown with illustrative code examples. No data access, network activity, file writes, or command execution is possible. The skill provides evaluation guidance only.
Jan 4, 2026, 04:38 PM
Pure documentation skill containing no executable files or tooling. All content is static markdown with illustrative code examples. No data access, network activity, file writes, or command execution is possible. The skill provides evaluation guidance only.