# Build Reliable DevOps and SRE Workflows

Complex delivery and reliability work often spans disconnected tools and operational practices. This skill provides structured workflows for implementation, diagnosis, automation, and review.

## Install

```bash
npx skillstore add zl2023github/devops-sre-engineer
```

## Metadata

- Status: approved
- Slug: zl2023github-devops-sre-engineer
- Skillstore revision: r2
- Version status: missing
- Tree hash: 22b0c5053b472fdd39770716f309a07794476a856c78526728397b31cda0ce7d
- Author: zl2023github
- GitHub username: zl2023github
- License: MIT
- Repository: https://github.com/zl2023github/software-engineer-skills/tree/main/software-engineering/devops-sre-engineer
- Ref: 88a8e9a07f4c54ab105c1c41b6267c287146b07b
- Supported tools: Claude, Codex, Claude Code
- Audit status: complete
- Agent install advisory: blocked
- Manual install advisory: allowed\_with\_warning
- Artifact signature: available
- Audit attestation: unavailable
- Human verification: not\_verified
- Risk factors: external\_commands, network, filesystem
- Quality score: 38
- Quality tier: warning
- Public page: https://skillstore.pages.dev/skills/zl2023github-devops-sre-engineer
- Manifest: https://skillstore.pages.dev/api/skills/zl2023github-devops-sre-engineer/manifest

## Capabilities

- Designs CI/CD workflows for GitHub Actions, GitLab CI, Jenkins, ArgoCD, and Tekton.
- Plans Docker, Kubernetes, Helm, Kustomize, service mesh, and GitOps configurations.
- Develops Terraform, OpenTofu, Pulumi, Ansible, and Packer workflows.
- Defines observability, SLI, SLO, error budget, alerting, and incident response practices.
- Guides capacity testing, performance analysis, chaos experiments, backups, security scanning, and cost optimization.
- Produces operational checklists, troubleshooting sequences, postmortems, configuration plans, and automation guidance.

## Use Cases

- Create a Delivery Pipeline: Design staged build, test, scan, approval, deployment, health check, and rollback workflows for an application.
- Investigate a Production Incident: Organize impact assessment, recent-change review, metrics, logs, traces, dependency checks, recovery, and postmortem actions.
- Establish Reliability Controls: Define service indicators, objectives, error budgets, dashboards, alerts, capacity assumptions, and release policies.

## Prompt Templates

### Review My Deployment

```
Review my deployment setup for [service] in [environment]. Identify missing checks, risks, and the next three safe improvements.
```

### Diagnose an Incident

```
Diagnose [symptom] using these metrics, logs, events, and recent changes: [evidence]. Separate observations, hypotheses, tests, mitigation, and follow-up.
```

### Design a Delivery System

```
Design a CI/CD and GitOps workflow for [stack], [repository], and [platform]. Include approvals, security scans, promotion, rollback, and verification.
```

### Plan Reliability Engineering

```
Create a reliability plan for [service]. Define SLIs, SLOs, error budgets, capacity tests, observability, failure experiments, governance, and implementation phases.
```

## Limitations

- Results depend on accurate repository, environment, architecture, and incident context.
- Operational commands can affect live systems and require human approval.
- The skill cannot verify permissions, credentials, backups, or rollback readiness without environment access.
- Security, compliance, and production changes still require organizational review.

## Best Practices

- Provide architecture, environment, constraints, recent changes, and available evidence before requesting operational guidance.
- Review plans and previews before changes, then verify health indicators and rollback readiness.
- Run performance, chaos, security, and recovery tests only in approved scopes with monitoring and stop conditions.

## Anti Patterns

- Do not execute destructive commands from a generic example against an unverified target.
- Do not place secrets, tokens, private keys, or sensitive production data in prompts.
- Do not treat generated guidance as approval for production, security, or compliance changes.

## Security Audit

- Audited at: 2026-07-23T23:32:07.738\+00:00
- Summary: Most alerts are false positives caused by Markdown backticks, placeholders, standard device handling, readable Chinese text, and routine local diagnostics. The Docker Bench example grants a remote image Docker socket and host namespace access, creating critical host compromise risk. Unpinned cluster installation, ungated destructive commands, and unrestricted load testing require correction before publication.

## Stats

- Views: 1
- Downloads: 6
- Favorites: 0
- Popularity score: 0
