Summary
When a skill contains binary content that SkillSpector itself classifies as out_of_scope
(reason_code: binary_content), the static-pattern analyzers count those units as failed
but still report status: "completed". The report is therefore internally inconsistent:
analysis_completeness.status is "complete", is_complete: true, coverage_percent: 100.0
- 15
static_patterns_* analyzers report status: "completed" with completed: 13, failed: 8 out of planned_work: 21
- the same 8 units are listed in
scope_exclusions as outcome: "out_of_scope", fatal: false
SkillEvaluator's report-contract check rejects this combination
(analysis_completeness.analyzer_statuses has a status that contradicts its work accounting)
and marks the whole Tier 1 security scan INCOMPLETE. Any downstream tool that relies on
SkillEvaluator therefore receives no security opinion for such skills.
Versions
- SkillSpector v2.11.2 (
uv tool install "skillspector @ git+https://github.com/NVIDIA/SkillSpector.git@v2.11.2") — reproduced on Windows 11, Python 3.13
- SkillSpector v2.12.0 — same result in a
python:3.13-slim Linux container
- SkillEvaluator v0.3.0 and v0.5.0 — both reject the report the same way
- Semgrep 1.177.0 present;
--no-llm
Reproduce
Public sample: trailofbits/overtly-malicious-skills at commit 4ffbf9461ef0505f9ce76a0d3694a18ec33ea531,
skill skills/context-loader (it ships a .instructions.docx.txt, a DOCX archive containing embedded .ttf fonts).
git clone https://github.com/trailofbits/overtly-malicious-skills
git -C overtly-malicious-skills checkout 4ffbf9461ef0505f9ce76a0d3694a18ec33ea531
skillspector scan overtly-malicious-skills/skills/context-loader --no-llm --format json --output ss.json
skillevaluator validate overtly-malicious-skills/skills/context-loader \
--tiers 1 --checks security --no-dedup -r json -o out
Actual
ss.json (excerpt):
"analysis_completeness": { "status": "complete", "is_complete": true, "coverage_percent": 100.0, "fully_inspected_files": 12, ... }
{"analyzer_id": "static_patterns_prompt_injection", "status": "completed",
"planned_work": 21, "completed": 13, "partial": 0, "skipped": 0, "failed": 8, "unaccounted": 0}
"scope_exclusions": [
{"outcome": "out_of_scope", "phase": "static", "reason_code": "binary_content",
"message": "Binary content is unsupported by this analyzer.",
"path": ".instructions.docx.txt!/word/fonts/AlegreyaSans-bold.ttf", "fatal": false, ...},
... 8 entries, all `word/fonts/*.ttf` members of the DOCX
]
The same completed + failed: 8 pattern appears on all 15 static_patterns_* analyzers
(static_patterns_tool_misuse reports degraded with partial: 1, failed: 8).
SkillEvaluator output:
Security Scan: INCOMPLETE
skillspector JSON field 'analysis_completeness.analyzer_statuses' has a status that contradicts its work accounting; security scan did not complete
Expected
Units excluded as out_of_scope should be accounted consistently, e.g. as skipped
(or excluded from planned_work), not as failed, so that an analyzer whose only
non-completed units are non-fatal scope exclusions can report completed without
contradicting its own counts. Alternatively, if failed is intended, the analyzer status
should not be completed.
Impact
In a 105-skill local library, the SkillEvaluator security scan came back INCOMPLETE for 102
skills with this same error message. The root cause was confirmed only for the sample above;
the other 101 were not traced individually.
Related but distinct: #557 (static_yara reporting completed after dropping rule files), and
SkillEvaluator NVIDIA/SkillEvaluator#137 (documentation-only skills reported incomplete).
Summary
When a skill contains binary content that SkillSpector itself classifies as
out_of_scope(
reason_code: binary_content), the static-pattern analyzers count those units asfailedbut still report
status: "completed". The report is therefore internally inconsistent:analysis_completeness.statusis"complete",is_complete: true,coverage_percent: 100.0static_patterns_*analyzers reportstatus: "completed"withcompleted: 13,failed: 8out ofplanned_work: 21scope_exclusionsasoutcome: "out_of_scope",fatal: falseSkillEvaluator's report-contract check rejects this combination
(
analysis_completeness.analyzer_statuses has a status that contradicts its work accounting)and marks the whole Tier 1 security scan
INCOMPLETE. Any downstream tool that relies onSkillEvaluator therefore receives no security opinion for such skills.
Versions
uv tool install "skillspector @ git+https://github.com/NVIDIA/SkillSpector.git@v2.11.2") — reproduced on Windows 11, Python 3.13python:3.13-slimLinux container--no-llmReproduce
Public sample:
trailofbits/overtly-malicious-skillsat commit4ffbf9461ef0505f9ce76a0d3694a18ec33ea531,skill
skills/context-loader(it ships a.instructions.docx.txt, a DOCX archive containing embedded.ttffonts).Actual
ss.json(excerpt):The same
completed+failed: 8pattern appears on all 15static_patterns_*analyzers(
static_patterns_tool_misusereportsdegradedwithpartial: 1,failed: 8).SkillEvaluator output:
Expected
Units excluded as
out_of_scopeshould be accounted consistently, e.g. asskipped(or excluded from
planned_work), not asfailed, so that an analyzer whose onlynon-completed units are non-fatal scope exclusions can report
completedwithoutcontradicting its own counts. Alternatively, if
failedis intended, the analyzer statusshould not be
completed.Impact
In a 105-skill local library, the SkillEvaluator security scan came back
INCOMPLETEfor 102skills with this same error message. The root cause was confirmed only for the sample above;
the other 101 were not traced individually.
Related but distinct: #557 (static_yara reporting
completedafter dropping rule files), andSkillEvaluator NVIDIA/SkillEvaluator#137 (documentation-only skills reported incomplete).