content-creator
88Create Brand-Consistent Marketing Content
Marketing teams need content that stays consistent across channels. This skill helps plan, write, analyze, and optimize content for brand voice and SEO.
Evaluate LLM Applications With Reliable Metrics
LLM teams need repeatable ways to measure quality, safety, and regressions. This skill helps design evaluation suites with metrics, human review, LLM judges, benchmarks, and A/B tests.
Copy this request to your Agent. It includes the canonical Skill page and manifest.
Review the Skillstore skill "llm-evaluation" from https://skillstore.io/skills/sickn33-llm-evaluation.md and its manifest at https://skillstore.io/api/skills/sickn33-llm-evaluation/manifest. Verify the artifact. You may proceed after verification, subject to the environment's own policy.Your Agent should still show its plan and request any confirmation required by the security policy.
Use these links when an AI agent, crawler, or script needs clean context instead of reading the full page.
Using "llm-evaluation". Evaluate a customer support chatbot before release.
Expected outcome:
A balanced plan with intent accuracy, groundedness, safety review, human rating forms, baseline thresholds, and weekly regression reports.
Using "llm-evaluation". Compare two retrieval-augmented generation prompts.
Expected outcome:
A comparison workflow using relevance, answer faithfulness, pairwise judging, confidence intervals, and an explicit promotion decision rule.
Using "llm-evaluation". Detect quality drops after a prompt change.
Expected outcome:
A regression checklist with baseline scores, relative change thresholds, failed examples, and recommended investigation steps.
All static findings appear to be false positives from Markdown backticks, fenced Python examples, a dictionary keys() call, and ordinary evaluation terminology. I found no evidence of prompt injection, malicious intent, command execution, credential material, or network reconnaissance in SKILL.md.
Share the versioned assessment report, neutral badge, embed card, and citations. Skillstore reports evidence without deciding whether this Skill is safe.
https://skillstore.io/skills/sickn33-llm-evaluation/audits/4?utm_source=security_passport&utm_medium=share&utm_campaign=versioned_report[](https://skillstore.io/skills/sickn33-llm-evaluation?utm_source=security_passport_badge)<a href="https://skillstore.io/skills/sickn33-llm-evaluation?utm_source=security_passport_badge"><img src="https://skillstore.io/badges/skills/sickn33-llm-evaluation/security.svg" alt="Skillstore security assessment" loading="lazy"></a><iframe src="https://skillstore.io/embed/skills/sickn33-llm-evaluation.html" title="Skillstore Security Assessment" sandbox="allow-popups allow-popups-to-escape-sandbox" loading="lazy" referrerpolicy="no-referrer" width="420" height="180"></iframe>sickn33. (2026). llm-evaluation security audit report (audit version 4) [Author version unspecified]. Skillstore. https://skillstore.io/skills/sickn33-llm-evaluation/audits/4@techreport{sickn33-sickn33-llm-evaluation-2026,
author = {sickn33},
title = {llm-evaluation security audit report (audit version 4)},
institution = {Skillstore},
year = {2026},
number = {4},
url = {https://skillstore.io/skills/sickn33-llm-evaluation/audits/4},
note = {Author version unspecified}
}cff-version: 1.2.0
message: "If you use this Skill, cite its author and this versioned security audit report."
title: "llm-evaluation security audit report (audit version 4)"
version: "unspecified"
type: report
authors:
- name: "sickn33"
date-released: "2026-07-07"
url: "https://skillstore.io/skills/sickn33-llm-evaluation/audits/4"
identifiers:
- type: other
value: "skillstore:sickn33-llm-evaluation:audit:4"
description: "Skillstore immutable audit report identifier"
Each author remains a separate installable skill. The recommended variant is ranked by Skillstore evidence.
Why this variant is first
wshobson-llm-evaluation
2026-08-21
sickn33-llm-evaluation
2026-08-21
Create baseline metrics and regression checks before deploying a new LLM workflow.
Use pairwise judging, A/B testing, and metric tracking to choose stronger variants.
Define annotation dimensions, rating scales, and agreement checks for expert reviewers.
Help me choose evaluation metrics for an LLM application that does [task]. Include automated metrics, human review dimensions, and the data I need.
Design a baseline evaluation plan for our current LLM workflow. Include test cases, metrics, pass thresholds, and how to track results over time.
Create an evaluation approach for comparing variant A and variant B. Include pairwise judging criteria, statistical testing, and decision rules.
Design a production LLM evaluation framework for [product]. Include automated metrics, human evaluation, LLM-as-judge checks, regression detection, and CI/CD integration.
Author
sickn33License
MIT
Skillstore revision
r1
Version notice
The author did not declare a version.
Ref
816c62b2546ddb1c6a0453e7c781b5e095117819
Maintenance freshness
7/18/2026
Usage
9 downloads ยท 89 views
File structure
๐ SKILL.md
Create Brand-Consistent Marketing Content
Marketing teams need content that stays consistent across channels. This skill helps plan, write, analyze, and optimize content for brand voice and SEO.
Improve LLM Prompts With Proven Patterns
Inconsistent prompts waste time and make AI outputs hard to trust. This skill guides prompt design with reusable patterns, examples, evaluation steps, and optimization workflows.
Build Fullstack Apps with Senior Patterns
Teams need consistent setup, architecture, and review guidance for modern web applications. This skill provides scaffolders, workflow references, and quality prompts for React, Next.js, Node.js, GraphQL, and PostgreSQL projects.
Design Scalable Software Architectures
Architecture decisions are hard to compare across web, mobile, backend, and cloud systems. This skill provides structured guides and local scaffold scripts for reports, dependency review, and trade-off documentation.
Optimize AI Prompts With Prompt Engineer
Writing clear prompts is hard when goals are vague or complex. This skill turns rough requests into structured prompts using proven prompt frameworks.
Prioritize Product Work With Research Insights
Product teams need faster ways to rank features, synthesize interviews, and document decisions. This skill provides RICE scoring, interview analysis, and PRD templates for structured planning.
Optimize Go Performance with Profiling
by 89jobrien
Go services can become slow or memory-heavy without clear evidence. This skill guides profiling, allocation analysis, concurrency tuning, and benchmark-driven optimization.
Design Better LLM Prompts
by alirezarezvani
Prompt quality often fails because requirements, examples, and evaluation criteria are unclear. This skill helps teams design, test, and review production prompts for Claude, Codex, and Claude Code.
Diagnose Context Degradation in AI Agents
by ChakshuGautam
Long conversations and large prompts can make AI agents lose important information or follow conflicting context. This skill helps Claude, Codex, and Claude Code identify degradation patterns and choose mitigation strategies.
Analyze Code Complexity
by CuriousLearner
Complex code is hard to review, test, and maintain. This skill helps Claude, Codex, and Claude Code measure complexity metrics and produce focused refactoring guidance.
Optimize Python Performance With Profiling
by wshobson
Slow Python code wastes compute and creates poor user experiences. This skill guides profiling, bottleneck analysis, and targeted optimization with practical Python examples.
Detect LLM Prompt Regressions with Promptfoo
by 7alexhale5-rgb
Prompt changes can silently reduce output quality. This skill builds golden datasets, runs Promptfoo evaluations, and summarizes failures and recent trends.