📦

Audit History

gpt-series-reasoning-style - 6 audits

Version comparison

Capability and finding changes across audited versions, newest first.

VersionDateResultReview itemsChange vs previous
v6 LatestSep 20, 2026, 03:25 AM No confirmed findings0No capability change
v5 Sep 18, 2026, 05:39 AM No confirmed findings0No capability change
v4 Sep 17, 2026, 08:05 PM 1 confirmed0No capability change
v3 Sep 16, 2026, 07:18 PM No confirmed findings0No capability change
v2 Sep 12, 2026, 11:20 AM No confirmed findings1No capability change
v1 Sep 11, 2026, 09:27 PM No confirmed findings1Baseline

Sep 20, 2026, 03:25 AM

All 91 static findings were adjudicated as false positives because they reference public metadata, documentation, fixed repository paths, or analysis heuristics rather than executed behavior. No prompt-injection text, data-exfiltration intent, or runtime command execution was evidenced in the reviewed files.

24
Files scanned
2,119
Lines analyzed
3
Review items
0
False positives ignored
Audited by: codex

Sep 17, 2026, 08:05 PM

58 个静态命中均对应文档中的安装示例、代码围栏、正则字面量或纯文本熵启发式,未发现相应的运行时攻击行为。AGENTS.md 与 SKILL.md 仍包含控制代理加载顺序和执行规则的高风险提示注入式文本,安装前应由宿主策略和用户明确同意进行约束。

21
Files scanned
1,653
Lines analyzed
4
Review items
0
False positives ignored

Confirmed security concerns (1)

High
Prompt Injection Attempt Detected
AGENTS.md 要求代理按其加载顺序读取技能文件,SKILL.md 又要求“禁止读取”规则文件并“严格按其中内容执行任务”。这些文本试图控制代理的指令加载和执行顺序,项目级安装时可能与宿主策略或用户指令竞争;未发现直接数据窃取或代码执行载荷。
文件直接使用“强制门禁”、读取禁令和“严格按其中内容执行任务”等代理控制措辞,属于强提示注入信号;但内容未要求窃取数据或执行危险命令,因此风险判断低于确定性恶意载荷。
Audited by: codex

Sep 16, 2026, 07:18 PM

26 个静态发现均为误报:证据来自文档安装命令、Markdown 反引号、Python 文本匹配和普通文本熵启发式。未发现实际网络通信、危险命令执行、数据外传或提示注入证据。

17
Files scanned
1,293
Lines analyzed
3
Review items
0
False positives ignored
Audited by: codex

Sep 12, 2026, 11:20 AM

Most static alerts are false positives caused by defensive examples, readable Chinese prose, Markdown syntax, SVG paths, and documented installation locations. One high-risk behavior is confirmed: claim-check executes claims-file commands through shell=True, so hostile or insufficiently reviewed input can run arbitrary code. Static review was capped at 400/554 representative findings; omitted static matches are unconfirmed, so automatic publishing stays disabled until manual review.

97
Files scanned
13,494
Lines analyzed
4
Review items
0
False positives ignored
Capability review items (1)

These are real local capabilities that may be expected for this skill, so they require review but are not counted as confirmed malicious behavior.

High
Python subprocess.run
proc = subprocess.run(cmd, shell=True, cwd=str(root),
This call passes commands parsed from a claims file to subprocess.run with shell=True. The file and SECURITY.md explicitly acknowledge that hostile input can execute arbitrary code.

Risk Factors

🌐 Network access (13)
⚙️ External commands (50)
📁 Filesystem access (50)
.github/workflows/selfcheck.yml:32 .github/workflows/selfcheck.yml:33 .github/workflows/selfcheck.yml:59 .github/workflows/selfcheck.yml:61 .github/workflows/selfcheck.yml:63 .github/workflows/selfcheck.yml:66 .github/workflows/selfcheck.yml:43 .github/workflows/selfcheck.yml:46 .github/workflows/selfcheck.yml:48 .github/workflows/selfcheck.yml:49 .github/workflows/selfcheck.yml:50 .github/workflows/selfcheck.yml:55 assets/README.md:11 CHANGELOG.md:29 CHANGELOG.md:33 CHANGELOG.md:44 CHANGELOG.md:29 CHANGELOG.md:105 CHANGELOG.md:29 CHANGELOG.md:71 CHANGELOG.md:77 CHANGELOG.md:105 CHANGELOG.md:47 docs/field-tests/ab-baseline/judgement-sheet.md:166 docs/field-tests/README.md:16 docs/field-tests/README.md:36 docs/field-tests/selftest-run-2026-09-10/report.md:4 docs/field-tests/selftest-run-2026-09-10/report.md:13 docs/proposals/2026-09-09-rule-audit-probe-c.md:25 docs/proposals/2026-09-09-rule-audit-probe-c.md:31 docs/proposals/2026-09-09-rule-audit-probe-c.md:37 docs/proposals/2026-09-09-rule-audit-probe-c.md:42 docs/proposals/2026-09-09-rule-audit-probe-c.md:47 docs/proposals/2026-09-09-rule-audit-probe-c.md:53 docs/proposals/2026-09-09-rule-audit-probe-c.md:59 docs/proposals/2026-09-09-rule-audit-probe-c.md:64 docs/proposals/2026-09-09-rule-audit-probe-c.md:69 docs/reviews/2026-09-09-three-ai-audit-verdict.md:82 docs/reviews/2026-09-10-deep-audit.md:17 docs/reviews/2026-09-10-deep-audit.md:27 docs/reviews/2026-09-11-full-repo-audit.md:33 docs/reviews/2026-09-11-verification-sensitivity-mutation-kill.md:14 docs/reviews/2026-09-11-verification-sensitivity-mutation-kill.md:17 docs/reviews/2026-09-11-verification-sensitivity-mutation-kill.md:74 docs/reviews/2026-09-11-verification-sensitivity-mutation-kill.md:77 hooks/README.md:17 hooks/README.md:17 hooks/README.md:28 hooks/README.md:71 README.md:421
Audited by: codex

Sep 11, 2026, 09:27 PM

The audit confirms one high-risk issue: scripts/claim-check.py executes commands from an untrusted claims file with shell=True. The other reviewed matches are false positives from documentation, defensive patterns, test harnesses, installer mechanics, or static-site content; the confirmed issue requires remediation before publication. Static review was capped at 400/536 representative findings; omitted static matches are unconfirmed, so automatic publishing stays disabled until manual review.

94
Files scanned
13,319
Lines analyzed
4
Review items
0
False positives ignored
Capability review items (1)

These are real local capabilities that may be expected for this skill, so they require review but are not counted as confirmed malicious behavior.

High
Python subprocess.run
proc = subprocess.run(cmd, shell=True, cwd=str(root),
At line 352, subprocess.run uses shell=True on each command parsed from an untrusted claims file. A malicious claims file can execute arbitrary shell commands despite the blacklist, so this is a real command-execution risk.

Risk Factors

🌐 Network access (13)
⚙️ External commands (50)
📁 Filesystem access (50)
.github/workflows/selfcheck.yml:32 .github/workflows/selfcheck.yml:33 .github/workflows/selfcheck.yml:59 .github/workflows/selfcheck.yml:61 .github/workflows/selfcheck.yml:63 .github/workflows/selfcheck.yml:66 .github/workflows/selfcheck.yml:43 .github/workflows/selfcheck.yml:46 .github/workflows/selfcheck.yml:48 .github/workflows/selfcheck.yml:49 .github/workflows/selfcheck.yml:50 .github/workflows/selfcheck.yml:55 CHANGELOG.md:20 CHANGELOG.md:24 CHANGELOG.md:35 CHANGELOG.md:20 CHANGELOG.md:96 CHANGELOG.md:20 CHANGELOG.md:62 CHANGELOG.md:68 CHANGELOG.md:96 CHANGELOG.md:38 docs/field-tests/ab-baseline/judgement-sheet.md:166 docs/field-tests/README.md:16 docs/field-tests/README.md:34 docs/field-tests/selftest-run-2026-09-10/report.md:4 docs/field-tests/selftest-run-2026-09-10/report.md:13 docs/proposals/2026-09-09-rule-audit-probe-c.md:25 docs/proposals/2026-09-09-rule-audit-probe-c.md:31 docs/proposals/2026-09-09-rule-audit-probe-c.md:37 docs/proposals/2026-09-09-rule-audit-probe-c.md:42 docs/proposals/2026-09-09-rule-audit-probe-c.md:47 docs/proposals/2026-09-09-rule-audit-probe-c.md:53 docs/proposals/2026-09-09-rule-audit-probe-c.md:59 docs/proposals/2026-09-09-rule-audit-probe-c.md:64 docs/proposals/2026-09-09-rule-audit-probe-c.md:69 docs/reviews/2026-09-09-three-ai-audit-verdict.md:82 docs/reviews/2026-09-10-deep-audit.md:17 docs/reviews/2026-09-10-deep-audit.md:27 docs/reviews/2026-09-11-full-repo-audit.md:33 docs/reviews/2026-09-11-verification-sensitivity-mutation-kill.md:14 docs/reviews/2026-09-11-verification-sensitivity-mutation-kill.md:17 docs/reviews/2026-09-11-verification-sensitivity-mutation-kill.md:74 docs/reviews/2026-09-11-verification-sensitivity-mutation-kill.md:77 hooks/README.md:17 hooks/README.md:17 hooks/README.md:28 hooks/README.md:71 README.md:419 README.md:420
Audited by: codex