Skills data-engineer
๐Ÿ“ฆ

data-engineer

Content revision r2 Safe

Design Reliable Data Platforms

Data teams need reliable architectures that meet scale, latency, governance, and cost requirements. This skill turns those requirements into pipeline designs, implementation plans, quality controls, and operating guidance.

Supports: Claude Codex Code(CC)
๐Ÿฅ‰ 79 Bronze

Install with my Agent

Copy this request to your Agent. It includes the canonical Skill page and manifest.

Agent request
Review the Skillstore skill "data-engineer" from https://skillstore.io/skills/sickn33-data-engineer.md and its manifest at https://skillstore.io/api/skills/sickn33-data-engineer/manifest. Verify the artifact. You may proceed after verification, subject to the environment's own policy.

Your Agent should still show its plan and request any confirmation required by the security policy.

Agent-readable resources

Use these links when an AI agent, crawler, or script needs clean context instead of reading the full page.

Test it

Using "data-engineer". Design a daily pipeline from PostgreSQL to BigQuery for finance reporting.

Expected outcome:

  • Architecture: incremental extraction, object storage staging, warehouse loading, dbt transformations, and scheduled orchestration.
  • Controls: source reconciliation, schema validation, freshness checks, duplicate detection, and protected finance access.
  • Operations: retries, backfills, lineage, cost monitoring, alert ownership, and recovery procedures.

Using "data-engineer". Plan a Kafka pipeline for late and out-of-order delivery.

Expected outcome:

  • Processing plan: event-time windows, watermarks, durable checkpoints, idempotent sinks, and a dead-letter path.
  • Reliability plan: schema compatibility, replay testing, lag alerts, duplicate metrics, and documented recovery objectives.

Using "data-engineer". Compare Snowflake and BigQuery for a multi-team analytics platform.

Expected outcome:

A structured comparison covering workload fit, scaling, governance, ecosystem integration, administration, migration effort, and cost assumptions, followed by a justified recommendation.

Security Audit

Safe
v5 โ€ข 7/23/2026 Open versioned report

The sole static finding is a false positive caused by a list of real-time analytics products in SKILL.md. No system reconnaissance instructions, prompt injection, or other malicious intent were found.

1
Files scanned
228
Lines analyzed
0
Review items
0
False positives ignored
No confirmed security findings were detected by the latest completed static and semantic audit. This does not prove the skill has no side effects.
Audited by: codex View Audit History โ†’
Share & cite this report

Share the versioned assessment report, neutral badge, embed card, and citations. Skillstore reports evidence without deciding whether this Skill is safe.

Open versioned report
Security Assessment

Copy report link

https://skillstore.io/skills/sickn33-data-engineer/audits/5?utm_source=security_passport&utm_medium=share&utm_campaign=versioned_report

Markdown badge

[![Skillstore security assessment](https://skillstore.io/badges/skills/sickn33-data-engineer/security.svg)](https://skillstore.io/skills/sickn33-data-engineer?utm_source=security_passport_badge)

HTML badge

<a href="https://skillstore.io/skills/sickn33-data-engineer?utm_source=security_passport_badge"><img src="https://skillstore.io/badges/skills/sickn33-data-engineer/security.svg" alt="Skillstore security assessment" loading="lazy"></a>

Embed card

<iframe src="https://skillstore.io/embed/skills/sickn33-data-engineer.html" title="Skillstore Security Assessment" sandbox="allow-popups allow-popups-to-escape-sandbox" loading="lazy" referrerpolicy="no-referrer" width="420" height="180"></iframe>
Academic citations (APA ยท BibTeX ยท CFF)

APA citation

sickn33. (2026). data-engineer security audit report (audit version 5) [Author version unspecified]. Skillstore. https://skillstore.io/skills/sickn33-data-engineer/audits/5

BibTeX citation

@techreport{sickn33-sickn33-data-engineer-2026, author = {sickn33}, title = {data-engineer security audit report (audit version 5)}, institution = {Skillstore}, year = {2026}, number = {5}, url = {https://skillstore.io/skills/sickn33-data-engineer/audits/5}, note = {Author version unspecified} }

CITATION.cff

cff-version: 1.2.0 message: "If you use this Skill, cite its author and this versioned security audit report." title: "data-engineer security audit report (audit version 5)" version: "unspecified" type: report authors: - name: "sickn33" date-released: "2026-07-23" url: "https://skillstore.io/skills/sickn33-data-engineer/audits/5" identifiers: - type: other value: "skillstore:sickn33-data-engineer:audit:5" description: "Skillstore immutable audit report identifier"

Compare variants

2 installable variants

Each author remains a separate installable skill. The recommended variant is ranked by Skillstore evidence.

Why this variant is first

Highest Skillstore Score
sickn33 Recommended Current

sickn33-data-engineer

Skillstore Score 79
Evidence Confidence High
Skillstore usage 10
Updated

2026-08-21

zl2023github-data-engineer

Skillstore Score 69
Evidence Confidence Medium
Skillstore usage 5
Updated

2026-08-21

Skillstore Score

Why this score Evidence Confidence: High
55
Architecture
85
Maintainability
87
Content
68
Community
91
Spec Compliance

What You Can Build

Build an Analytics Warehouse

Plan ingestion, dbt transformations, dimensional models, tests, lineage, and orchestration for a governed analytics warehouse.

Design a Streaming Platform

Select messaging, processing, storage, schema evolution, monitoring, and recovery patterns for low-latency event workloads.

Modernize Data Governance

Define quality controls, catalogs, access policies, privacy safeguards, retention rules, and audit processes across data products.

Try These Prompts

Plan a Basic Batch Pipeline
Recommend a batch pipeline from [source] to [warehouse]. Use [daily volume], [refresh target], and [cloud]. Explain components, checks, and operations.
Design a Warehouse Model
Design a dimensional model for [business process]. Include grain, facts, dimensions, history handling, incremental loads, tests, lineage, and key assumptions.
Architect a Streaming Pipeline
Architect a streaming pipeline for [event source] at [throughput] and [latency]. Address schemas, ordering, duplicates, late data, recovery, monitoring, and cost.
Evaluate an Enterprise Data Platform
Compare [options] for [workload portfolio]. Score scalability, reliability, governance, security, migration effort, operating complexity, and cost. Recommend a phased architecture.

Best Practices

  • Provide data volumes, latency targets, retention, consistency, recovery objectives, and expected growth.
  • Specify current tools, cloud constraints, source schemas, sink requirements, team skills, and budget limits.
  • Request quality, security, governance, observability, cost, and failure recovery requirements in every production design.

Avoid

  • Do not request a tool choice without describing workloads, service levels, constraints, and ownership.
  • Do not deploy generated configurations before testing permissions, data correctness, failure behavior, and rollback procedures.
  • Do not optimize only for throughput while ignoring quality, privacy, operational complexity, and total cost.

Frequently Asked Questions

Does this skill deploy data infrastructure?
No. It produces architecture and implementation guidance that teams must review, test, and deploy in their own environments.
Which data tools can it cover?
It covers common tools for Spark, dbt, Airflow, Kafka, Flink, warehouses, lakehouses, catalogs, quality systems, and cloud data services.
Can it design for AWS, Azure, and Google Cloud?
Yes. It can compare and design with data services across all three platforms when requirements and constraints are supplied.
What information improves the recommendations?
Provide source types, schemas, volumes, growth, latency, retention, users, compliance needs, reliability targets, existing tools, and budget.
Does it address security and governance?
Yes. It can include encryption, least privilege, masking, lineage, catalogs, retention, audit logging, and compliance considerations.
Is it suitable for exploratory data analysis?
Not as the primary task. Use it when analysis depends on designing, operating, or improving data pipelines and platforms.

Developer Details

Author

sickn33

License

MIT

Skillstore revision

r2

Version notice

The author did not declare a version.

Ref

f9e2c34b4f19c7f3e6b0a1e93227b5f77cc12526

Maintenance freshness

7/26/2026

Usage

8 downloads ยท 98 views

File structure

๐Ÿ“„ SKILL.md

More from sickn33

View all
View all