📦

Audit History

test-driven-development - 9 audits

Version comparison

Capability and finding changes across audited versions, newest first.

VersionDateResultReview itemsChange vs previous
v9 LatestJul 9, 2026, 04:23 PM No confirmed findings0No capability change
v8 Jul 9, 2026, 04:23 PM No confirmed findings0No capability change
v7 Jul 6, 2026, 11:13 AM 1 confirmed0No capability change
v6 Jun 29, 2026, 10:54 PM No confirmed findings1No capability change
v5 Jan 17, 2026, 05:04 AM No confirmed findings0No capability change
v4 Jan 17, 2026, 05:04 AM No confirmed findings0External commands
v3 Jan 10, 2026, 01:32 PM No confirmed findings0No capability change
v2 Jan 10, 2026, 01:32 PM No confirmed findings0No capability change
v1 Jan 10, 2026, 01:32 PM No confirmed findings0Baseline

Jul 9, 2026, 04:23 PM

Static findings are false positives caused by Markdown fences, inline examples, and safe local npm test commands in SKILL.md. The network-reconnaissance hits are checklist wording about tests passing, and manual review of SKILL.zip found only the reviewed SKILL.md copy.

1
Files scanned
365
Lines analyzed
1
Review items
0
False positives ignored
Audited by: codex

Jul 9, 2026, 04:23 PM

Static findings are false positives caused by Markdown fences, inline examples, and safe local npm test commands in SKILL.md. The network-reconnaissance hits are checklist wording about tests passing, and manual review of SKILL.zip found only the reviewed SKILL.md copy.

1
Files scanned
365
Lines analyzed
1
Review items
0
False positives ignored
Audited by: codex

Jul 6, 2026, 11:13 AM

Static command findings were false positives from Markdown fences and inline examples, not executable Ruby or shell backtick evaluation. The zip archive contains only the same SKILL.md content, so the binary review finding is resolved. One semantic risk remains: the prose tells agents to delete code without requiring user confirmation.

1
Files scanned
365
Lines analyzed
2
Review items
0
False positives ignored

Confirmed security concerns (1)

Medium
Destructive Code Deletion Guidance
The skill tells agents to delete implementation code and start over when code was written before tests. Without an explicit confirmation step, this can cause loss of user work.
The deletion instruction is explicit and repeated, including "Delete means delete" and "Delete code. Start over with TDD." The risk is bounded by TDD context, so medium severity is appropriate.
Audited by: codex

Jun 29, 2026, 10:54 PM

AI review found no evidence of malware, data exfiltration, prompt injection, weak cryptography, or network reconnaissance in SKILL.md. Most static findings are false positives from Markdown code fences and example text. The only real risk factor is guidance to run local npm test commands, which is expected for a TDD skill but should remain user-controlled.

1
Files scanned
365
Lines analyzed
2
Review items
1
False positives ignored
Capability review items (1)

These are real local capabilities that may be expected for this skill, so they require review but are not counted as confirmed malicious behavior.

Low
Local Test Command Guidance
The skill instructs the agent to run local npm test commands to verify red and green TDD states. This is a legitimate development workflow, but it is still external command execution and should use the project test runner only.
The referenced lines contain explicit local npm test commands and expected output. They do not include shell interpolation, downloads, credential access, or arbitrary command construction.
Static false positives ignored (1)

These static matches were dismissed by semantic review or matched schema-only tokens, so they are shown for transparency but do not drive the quality score.

Low
Static Analyzer False Positives in Markdown Examples
The reported weak cryptography and network reconnaissance hits do not correspond to actionable code in SKILL.md. They appear in plain documentation, Markdown fences, tables, or test examples with no executable malicious behavior.
Manual review of these lines found instructional prose and test-quality examples, not cryptographic code or network scanning. No executable network or crypto implementation is present in the skill file.
Audited by: codex

Jan 17, 2026, 05:04 AM

This is a documentation-only skill containing Test-Driven Development guidelines. No executable code, network calls, file system access, or external commands. Pure educational content. All 53 static findings are false positives from the scanner misinterpreting markdown code block delimiters and JSON metadata as executable code.

2
Files scanned
365
Lines analyzed
1
Review items
0
False positives ignored
Audited by: claude

Jan 17, 2026, 05:04 AM

This is a documentation-only skill containing Test-Driven Development guidelines. No executable code, network calls, file system access, or external commands. Pure educational content. All 53 static findings are false positives from the scanner misinterpreting markdown code block delimiters and JSON metadata as executable code.

2
Files scanned
365
Lines analyzed
1
Review items
0
False positives ignored
Audited by: claude

Jan 10, 2026, 01:32 PM

This is a documentation-only skill containing Test-Driven Development guidelines. No executable code, network calls, file system access, or external commands. Pure educational content.

2
Files scanned
365
Lines analyzed
0
Review items
0
False positives ignored
No confirmed security findings were recorded for this completed audit.
Audited by: claude

Jan 10, 2026, 01:32 PM

This is a documentation-only skill containing Test-Driven Development guidelines. No executable code, network calls, file system access, or external commands. Pure educational content.

2
Files scanned
365
Lines analyzed
0
Review items
0
False positives ignored
No confirmed security findings were recorded for this completed audit.
Audited by: claude

Jan 10, 2026, 01:32 PM

This is a documentation-only skill containing Test-Driven Development guidelines. No executable code, network calls, file system access, or external commands. Pure educational content.

2
Files scanned
365
Lines analyzed
0
Review items
0
False positives ignored
No confirmed security findings were recorded for this completed audit.
Audited by: claude