benchmark-kernel
Benchmark FlashInfer Kernels Accurately
GPU kernel comparisons are unreliable when timing methods and parameters vary. This skill guides reproducible FlashInfer benchmarks using CUPTI or CUDA events.
Install with my Agent
Copy this request to your Agent. It includes the canonical Skill page and manifest.
Review the Skillstore skill "benchmark-kernel" from https://skillstore.io/skills/flashinfer-ai-benchmark-kernel.md and its manifest at https://skillstore.io/api/skills/flashinfer-ai-benchmark-kernel/manifest. Verify the artifact. You may proceed after verification, subject to the environment's own policy.Your Agent should still show its plan and request any confirmation required by the security policy.
Agent-readable resources
Use these links when an AI agent, crawler, or script needs clean context instead of reading the full page.
Test it
Using "benchmark-kernel". Compare FA2 and cuDNN for one decode-attention shape on an H100.
Expected outcome:
- Benchmark plan: compare both backends on identical shapes with five warmups and thirty measured iterations.
- Validation: enable reference checking and record median latency, variation, throughput, and reproducer details.
Using "benchmark-kernel". My kernel latency changes significantly between runs.
Expected outcome:
- Diagnosis: review warmup count, GPU load, clock behavior, cache state, and timing method.
- Next steps: increase warmups and iterations, isolate competing workloads, then compare median latency and variation.
Using "benchmark-kernel". I need a custom timing study for a new FlashInfer kernel.
Expected outcome:
- Design: wrap the kernel with stable inputs and measure it through bench_gpu_time.
- Report: include timing mode, cache policy, iteration counts, latency statistics, operation count, and calculated throughput.
Security Audit
SafeEighty-four Ruby or shell backtick findings and the network-reconnaissance finding are false positives caused by Markdown formatting and benchmark arguments. The sudo nvidia-smi example is a real privileged system change that can alter GPU clocks for all workloads, although it contains no injection path. No malicious intent, prompt injection, or data exfiltration was found.
Capability review items (1)
These are real local capabilities that may be expected for this skill, so they require review but are not counted as confirmed malicious behavior.
Risk Factors
โ๏ธ External commands (50)
Share & cite this report
Share the versioned assessment report, neutral badge, embed card, and citations. Skillstore reports evidence without deciding whether this Skill is safe.
Copy report link
https://skillstore.io/skills/flashinfer-ai-benchmark-kernel/audits/12?utm_source=security_passport&utm_medium=share&utm_campaign=versioned_reportMarkdown badge
[](https://skillstore.io/skills/flashinfer-ai-benchmark-kernel?utm_source=security_passport_badge)HTML badge
<a href="https://skillstore.io/skills/flashinfer-ai-benchmark-kernel?utm_source=security_passport_badge"><img src="https://skillstore.io/badges/skills/flashinfer-ai-benchmark-kernel/security.svg" alt="Skillstore security assessment" loading="lazy"></a>Embed card
<iframe src="https://skillstore.io/embed/skills/flashinfer-ai-benchmark-kernel.html" title="Skillstore Security Assessment" sandbox="allow-popups allow-popups-to-escape-sandbox" loading="lazy" referrerpolicy="no-referrer" width="420" height="180"></iframe>Academic citations (APA ยท BibTeX ยท CFF)
APA citation
flashinfer-ai. (2026). benchmark-kernel security audit report (audit version 12) [Author version unspecified]. Skillstore. https://skillstore.io/skills/flashinfer-ai-benchmark-kernel/audits/12BibTeX citation
@techreport{flashinfer-ai-flashinfer-ai-benchmark-kernel-2026,
author = {flashinfer-ai},
title = {benchmark-kernel security audit report (audit version 12)},
institution = {Skillstore},
year = {2026},
number = {12},
url = {https://skillstore.io/skills/flashinfer-ai-benchmark-kernel/audits/12},
note = {Author version unspecified}
}CITATION.cff
cff-version: 1.2.0
message: "If you use this Skill, cite its author and this versioned security audit report."
title: "benchmark-kernel security audit report (audit version 12)"
version: "unspecified"
type: report
authors:
- name: "flashinfer-ai"
date-released: "2026-07-23"
url: "https://skillstore.io/skills/flashinfer-ai-benchmark-kernel/audits/12"
identifiers:
- type: other
value: "skillstore:flashinfer-ai-benchmark-kernel:audit:12"
description: "Skillstore immutable audit report identifier"
Skillstore Score
Why this score Evidence Confidence: HighWhat You Can Build
Compare Kernel Backends
Measure identical attention or GEMM shapes across supported backends with correctness checks and consistent timing.
Build Repeatable Benchmark Suites
Run parameterized test lists, preserve reproducer details, and save metrics for later analysis.
Design Custom Timing Studies
Use bench_gpu_time for controlled experiments with warmups, iteration counts, cache behavior, and timing selection.
Try These Prompts
Prepare a minimal FlashInfer decode-attention benchmark for my [GPU] using [backends]. Include prerequisites, reference checking, timing behavior, and result interpretation.
Build a fair comparison for [routine] across [backends] using [shape and data types]. Keep inputs, warmups, iterations, and validation identical.
Create a batch benchmark plan covering [workloads]. Define the test matrix, CSV fields, correctness checks, reproducer details, and comparison criteria.
Design a bench_gpu_time experiment for [kernel]. Address CUPTI fallback, cache state, statistical stability, operation counts, throughput calculations, and likely measurement bias.
Best Practices
- Enable reference checking before trusting performance comparisons.
- Record GPU model, software versions, parameters, warmups, and measurement iterations.
- Use identical inputs and timing methods when comparing backends.
Avoid
- Do not compare results collected with different shapes, data types, or timing methods.
- Do not treat one iteration as representative performance.
- Do not use output mismatches as valid performance evidence.
Frequently Asked Questions
Does this skill run benchmarks automatically?
Is CUPTI required?
Which kernel families are covered?
Can the skill compare multiple backends?
Which results can I collect?
How can I improve reproducibility?
Developer Details
Author
flashinfer-aiLicense
MIT
Skillstore revision
r2
Version notice
The author did not declare a version.
Ref
c43861a65bb95efcae259cd161c9d6f4dc7eec6f
Maintenance freshness
7/25/2026
Usage
3 downloads ยท 291 views
File structure
๐ SKILL.md