agent-ui
78Build React Agent Interfaces
Teams need agent chat interfaces without rebuilding streaming, approvals, and tool displays. This skill guides setup of the inference.sh Agent component in React and Next.js.
Transcribe Audio with Whisper Models
Audio and video recordings are hard to search or reuse without accurate transcripts. This skill uses inference.sh Whisper models to create transcripts, translations, and timestamped segments.
Do not auto-install this skill.
The canonical policy requires operator review before any installation action.
Copy this request to your Agent. It includes the canonical Skill page and manifest.
Review the Skillstore skill "speech-to-text" from https://skillstore.io/skills/inference-sh-9-speech-to-text.md and its manifest at https://skillstore.io/api/skills/inference-sh-9-speech-to-text/manifest. Verify the artifact. Do not auto-install. Inspect the skill and report your findings, then wait for an operator or manual installation decision.Your Agent should still show its plan and request any confirmation required by the security policy.
Use these links when an AI agent, crawler, or script needs clean context instead of reading the full page.
Using "speech-to-text". A meeting recording URL with English speech.
Expected outcome:
A full transcript with detected English language and paragraph breaks for review.
Using "speech-to-text". A podcast episode URL with timestamp request enabled.
Expected outcome:
A transcript split into timed segments that can guide editing and captions.
Using "speech-to-text". A French audio recording URL with translation requested.
Expected outcome:
An English transcript that preserves the main spoken content.
Most Ruby backtick detections are false positives caused by Markdown fences and inline code. The Quick Start contains a confirmed curl-to-shell installer, and its installer URL is part of the same critical remote code execution risk. No prompt injection text or hidden exfiltration intent was found.
These are real local capabilities that may be expected for this skill, so they require review but are not counted as confirmed malicious behavior.
Share the versioned assessment report, neutral badge, embed card, and citations. Skillstore reports evidence without deciding whether this Skill is safe.
https://skillstore.io/skills/inference-sh-9-speech-to-text/audits/4?utm_source=security_passport&utm_medium=share&utm_campaign=versioned_report[](https://skillstore.io/skills/inference-sh-9-speech-to-text?utm_source=security_passport_badge)<a href="https://skillstore.io/skills/inference-sh-9-speech-to-text?utm_source=security_passport_badge"><img src="https://skillstore.io/badges/skills/inference-sh-9-speech-to-text/security.svg" alt="Skillstore security assessment" loading="lazy"></a><iframe src="https://skillstore.io/embed/skills/inference-sh-9-speech-to-text.html" title="Skillstore Security Assessment" sandbox="allow-popups allow-popups-to-escape-sandbox" loading="lazy" referrerpolicy="no-referrer" width="420" height="180"></iframe>inference-sh-9. (2026). speech-to-text security audit report (audit version 4) [Author version unspecified]. Skillstore. https://skillstore.io/skills/inference-sh-9-speech-to-text/audits/4@techreport{inference-sh-9-inference-sh-9-speech-to-text-2026,
author = {inference-sh-9},
title = {speech-to-text security audit report (audit version 4)},
institution = {Skillstore},
year = {2026},
number = {4},
url = {https://skillstore.io/skills/inference-sh-9-speech-to-text/audits/4},
note = {Author version unspecified}
}cff-version: 1.2.0
message: "If you use this Skill, cite its author and this versioned security audit report."
title: "speech-to-text security audit report (audit version 4)"
version: "unspecified"
type: report
authors:
- name: "inference-sh-9"
date-released: "2026-07-06"
url: "https://skillstore.io/skills/inference-sh-9-speech-to-text/audits/4"
identifiers:
- type: other
value: "skillstore:inference-sh-9-speech-to-text:audit:4"
description: "Skillstore immutable audit report identifier"
Convert recorded meetings into readable text for notes, search, and follow-up tasks.
Turn podcast audio into text that can support editing, summaries, and publishing.
Generate timestamped transcript segments that can be used as caption inputs.
Use this skill to transcribe this meeting audio URL: [audio URL]. Return the full text and detected language.
Transcribe this podcast audio URL with timestamps: [audio URL]. Provide the transcript and segment timing summary.
Use Whisper V3 Large to translate this non-English audio URL to English: [audio URL]. Return clear English transcript text.
Extract audio from this video URL, transcribe it with timestamps, and prepare caption-ready text: [video URL]. Show each step.
Author
inference-sh-9License
MIT
Skillstore revision
r1
Version notice
The author did not declare a version.
Ref
a06681402992ceae98ba04d54cfd4ab004862696
Maintenance freshness
7/18/2026
Usage
11 downloads ยท 224 views
File structure
๐ SKILL.md
Build React Agent Interfaces
Teams need agent chat interfaces without rebuilding streaming, approvals, and tool displays. This skill guides setup of the inference.sh Agent component in React and Next.js.
Add Tool Lifecycle UI to Agent Apps
Agent applications need clear tool status, approval, and result displays. This skill provides React and Next.js component guidance for those workflows.
Render JSON Widgets in React
Agent output often needs interactive UI instead of plain text. This skill shows how to render structured widget definitions as React and Next.js components.
Build React Chat Interfaces
Teams need chat interfaces that match modern React apps. This skill provides shadcn-style components and patterns for messages, inputs, avatars, and streaming states.
Improve AI Prompts for Text, Images, and Video
AI results often fail when prompts are vague or poorly structured. This skill provides reusable patterns for LLM, image, and video prompts.
Search Web Sources with Tavily and Exa
Research tasks need current web evidence and clean page text. This skill runs inference.sh search and extraction apps for Tavily and Exa.
Transcribe Audio with verging.ai STT
by verging-ai
Audio recordings are hard to review, search, and share as text. This skill guides Claude, Codex, and Claude Code through verging.ai speech-to-text transcription using OpenAI Whisper.
Manage Lark Minutes and Meeting Artifacts
by larksuite
Meeting recordings and transcripts are difficult to search, retrieve, and update consistently. This skill guides Claude, Codex, and Claude Code through supported lark-cli workflows for Lark Minutes.
Summarize URLs, Files, and Media
by steipete
Long sources make key information difficult to review. This skill uses the Summarize CLI to extract or condense URLs, documents, transcripts, audio, and video.
Control Sonos Speakers from Your AI Assistant
by aiskillstore
Managing Sonos rooms manually slows routine playback and grouping tasks. This skill guides Claude, Codex, and Claude Code through focused Sonos CLI commands.
Create Fal.ai Audio Workflows
by sickn33
Audio tasks need clear steps for generation, transcription, and review. This skill helps Claude, Codex, and Claude Code plan fal.ai audio workflows.
Transcribe Audio with Azure AI and Python
by sickn33
Speech transcription projects require correct Azure client setup and workflow choices. This skill provides Python examples for batch and real-time transcription, authentication, diarization, and language selection.