context-compression
81Compress Long Agent Contexts Reliably
Long agent sessions can lose decisions, file history, and next steps during compression. This skill provides structured methods and probes that preserve operational context.
Build Reliable LLM Evaluation Systems
LLM quality reviews often drift because judges use vague criteria and exhibit systematic bias. This skill provides calibrated rubrics, comparison methods, and validation metrics.
Copy this request to your Agent. It includes the canonical Skill page and manifest.
Review the Skillstore skill "advanced-evaluation" from https://skillstore.io/skills/muratcankoylan-advanced-evaluation.md and its manifest at https://skillstore.io/api/skills/muratcankoylan-advanced-evaluation/manifest. Verify the artifact. You may proceed after verification, subject to the environment's own policy.Your Agent should still show its plan and request any confirmation required by the security policy.
Use these links when an AI agent, crawler, or script needs clean context instead of reading the full page.
Using "advanced-evaluation". Choose a method for rating factual answers against verified references.
Expected outcome:
Use direct scoring because objective evidence exists. Score factual accuracy separately, require cited evidence, and validate results against expert labels.
Using "advanced-evaluation". Compare two support responses for empathy and clarity.
Expected outcome:
Use two pairwise passes with swapped positions. Return a tie when winners differ, and lower confidence when criterion-level evidence is inconsistent.
Using "advanced-evaluation". Create a five-level code readability rubric.
Expected outcome:
All 44 static alerts are false positives caused by Markdown formatting, ordinary collection methods, evaluation terminology, and research links. The local example script only calculates and prints demonstration results. However, evaluator prompts interpolate untrusted candidate responses without explicit prompt-injection isolation, which could permit scoring manipulation.
Share the versioned assessment report, neutral badge, embed card, and citations. Skillstore reports evidence without deciding whether this Skill is safe.
https://skillstore.io/skills/muratcankoylan-advanced-evaluation/audits/8?utm_source=security_passport&utm_medium=share&utm_campaign=versioned_report[](https://skillstore.io/skills/muratcankoylan-advanced-evaluation?utm_source=security_passport_badge)<a href="https://skillstore.io/skills/muratcankoylan-advanced-evaluation?utm_source=security_passport_badge"><img src="https://skillstore.io/badges/skills/muratcankoylan-advanced-evaluation/security.svg" alt="Skillstore security assessment" loading="lazy"></a><iframe src="https://skillstore.io/embed/skills/muratcankoylan-advanced-evaluation.html" title="Skillstore Security Assessment" sandbox="allow-popups allow-popups-to-escape-sandbox" loading="lazy" referrerpolicy="no-referrer" width="420" height="180"></iframe>muratcankoylan. (2026). advanced-evaluation security audit report (audit version 8) [Author version unspecified]. Skillstore. https://skillstore.io/skills/muratcankoylan-advanced-evaluation/audits/8@techreport{muratcankoylan-muratcankoylan-advanced-evaluation-2026,
author = {muratcankoylan},
title = {advanced-evaluation security audit report (audit version 8)},
institution = {Skillstore},
year = {2026},
number = {8},
url = {https://skillstore.io/skills/muratcankoylan-advanced-evaluation/audits/8},
note = {Author version unspecified}
}cff-version: 1.2.0
message: "If you use this Skill, cite its author and this versioned security audit report."
title: "advanced-evaluation security audit report (audit version 8)"
version: "unspecified"
type: report
authors:
- name: "muratcankoylan"
date-released: "2026-08-09"
url: "https://skillstore.io/skills/muratcankoylan-advanced-evaluation/audits/8"
identifiers:
- type: other
value: "skillstore:muratcankoylan-advanced-evaluation:audit:8"
description: "Skillstore immutable audit report identifier"
Each author remains a separate installable skill. The recommended variant is ranked by Skillstore evidence.
Why this variant is first
chakshugautam-advanced-evaluation
2026-08-21
muratcankoylan-advanced-evaluation
2026-08-21
Design a position-swapped pairwise test that selects stronger responses while measuring consistency and confidence.
Create domain-specific rubrics that help human and LLM evaluators apply the same scoring standards.
Select agreement and ranking metrics that reveal systematic differences between automated scores and expert labels.
Review this evaluation goal: [goal]. Identify whether direct scoring or pairwise comparison fits best. Explain the decision and list required inputs.
Create a [scale] rubric for [criterion] in [domain]. Define each level, observable evidence, edge cases, and balanced scoring guidance.
Design a pairwise evaluation for [responses] using [criteria]. Include position swapping, label remapping, tie handling, confidence rules, and length-neutral instructions.
Audit this evaluation pipeline: [pipeline]. Check rubric calibration, prompt injection resistance, bias controls, human agreement, confidence calibration, and escalation thresholds.
Author
muratcankoylanLicense
MIT
Skillstore revision
r2
Version notice
The author did not declare a version.
Ref
02be9409c79ca1183f7844009c14d9df684d0cf9
Maintenance freshness
8/11/2026
Usage
11 downloads ยท 391 views
File structure
๐ references/
๐ bias-mitigation.md
๐ implementation-patterns.md
๐ metrics-guide.md
๐ scripts/
๐ SKILL.md
Compress Long Agent Contexts Reliably
Long agent sessions can lose decisions, file history, and next steps during compression. This skill provides structured methods and probes that preserve operational context.
Build Reliable Agent Evaluations
Agent quality is difficult to measure because outputs vary and may have several valid forms. This skill builds repeatable rubrics, tests, gates, and monitoring.
Diagnose and Repair Context Degradation
Long contexts can hide critical instructions, preserve bad claims, and mix conflicting tasks. This skill diagnoses the failure pattern and recommends placement, filtering, compression, isolation, or recovery strategies.
Optimize AI Context for Cost and Quality
Long AI sessions waste tokens and lose important context. This skill provides measured strategies for budgeting, masking, compaction, caching, retrieval, and partitioning.
Design Reliable Multi-Agent Systems
Multi-agent designs often add cost and coordination failures without improving outcomes. This skill helps select topologies, handoffs, consensus methods, and recovery controls.
Design Reliable Agent Tools
Ambiguous tools cause routing errors, malformed calls, and failed recovery. This skill provides practical patterns for clear schemas, descriptions, responses, and tool catalogs.
Evaluate LLM Applications with Measurable Quality
by wshobson
LLM teams need evidence before changing prompts, models, or retrieval flows. This skill guides automated metrics, judge reviews, human labels, A/B tests, and regression checks.
Optimize Prompts With Proven Patterns
by wshobson
LLM prompts can fail through unclear instructions, weak examples, or missing evaluation. This skill provides templates, reasoning patterns, and optimization workflows for Claude, Codex, and Claude Code.
Design Better LLM Prompts
by alirezarezvani
Prompt quality often fails because requirements, examples, and evaluation criteria are unclear. This skill helps teams design, test, and review production prompts for Claude, Codex, and Claude Code.
Optimize AI Prompts With Prompt Architect
by DNYoussef
Weak prompts create inconsistent AI outputs and slow review cycles. This skill provides a structured process for improving prompts, testing them, and documenting changes.
Audit Documentation Quality and Coverage
by C0ntr0lledCha0s
Incomplete documentation makes code harder to use, maintain, and review. This skill measures coverage and quality, identifies gaps, and prioritizes improvements.
Detect LLM Prompt Regressions with Promptfoo
by 7alexhale5-rgb
Prompt changes can silently reduce output quality. This skill builds golden datasets, runs Promptfoo evaluations, and summarizes failures and recent trends.