# Design Reliable Production MLOps Systems

Production machine learning requires coordinated pipelines, infrastructure, governance, deployment, and monitoring. This skill provides structured MLOps architecture guidance across major clouds and toolchains.

## Install

```bash
npx skillstore add sickn33/mlops-engineer
```

## Metadata

- Status: approved
- Slug: sickn33-mlops-engineer
- Skillstore revision: r2
- Version status: missing
- Tree hash: e0efe425dd64dd7233d830a02c8cad2cf869103c1d9bd2998043f69aeefc4768
- Author: sickn33
- GitHub username: sickn33
- License: MIT
- Repository: https://github.com/sickn33/antigravity-awesome-skills/tree/main/skills/mlops-engineer
- Ref: 81e05e636292629114b76cbb3922fbe57672fc02
- Supported tools: Claude, Codex, Claude Code
- Audit status: complete
- Agent install advisory: allowed
- Manual install advisory: allowed
- Artifact signature: available
- Audit attestation: unavailable
- Human verification: not\_verified
- Risk factors: external\_commands
- Quality score: 78
- Quality tier: bronze
- Public page: https://skillstore.pages.dev/skills/sickn33-mlops-engineer
- Manifest: https://skillstore.pages.dev/api/skills/sickn33-mlops-engineer/manifest

## Capabilities

- Design orchestration workflows using Kubeflow, Airflow, Prefect, Dagster, Argo, or managed cloud pipelines.
- Recommend experiment tracking, artifact versioning, model registry, and promotion workflows.
- Plan AWS, Azure, or GCP infrastructure for training, batch processing, and online inference.
- Define ML CI/CD strategies with testing, approval gates, canary releases, rollback, and GitOps.
- Develop monitoring plans for model performance, drift, data quality, infrastructure, and cost.
- Address security, compliance, scalability, disaster recovery, and operational documentation requirements.

## Use Cases

- Launch a Team MLOps Platform: Create a practical platform architecture for repeatable training, registry management, deployment, monitoring, and team ownership.
- Productionize a Research Model: Turn an experimental model workflow into a tested, versioned, observable, and recoverable production process.
- Govern Regulated ML Workloads: Plan controls, audit trails, approval gates, access management, and monitoring for regulated machine learning systems.

## Prompt Templates

### Outline a Starter MLOps Platform

```
Design a starter MLOps platform for [team and use case] on [cloud]. Include core components, workflow stages, ownership, and initial milestones.
```

### Select Tracking and Registry Tools

```
Compare [candidate tools] for experiment tracking and model registration. Evaluate integration, governance, scaling, operating effort, cost, and migration risks.
```

### Plan a Safe Model Release

```
Create a release plan for [model and service]. Include validation gates, canary deployment, monitoring thresholds, rollback criteria, approvals, and an operational runbook.
```

### Architect Regulated Multi-Cloud MLOps

```
Design a multi-cloud MLOps architecture for [workload] under [regulations]. Address identity, data boundaries, lineage, resilience, observability, cost, and disaster recovery.
```

## Limitations

- Provides architecture guidance and plans but does not directly provision infrastructure or deploy models.
- Requires environment details, workload targets, budget, and compliance constraints for specific recommendations.
- Cloud service availability, pricing, quotas, and product behavior require current provider documentation.
- The referenced implementation playbook is not included in the reported skill package.

## Best Practices

- Provide workload scale, latency, data sensitivity, compliance, budget, and recovery objectives.
- Define measurable validation, release, rollback, and monitoring criteria before choosing tools.
- Verify recommendations against current provider documentation and test them in a non-production environment.

## Anti Patterns

- Do not select a complex platform before defining operating requirements and team ownership.
- Do not promote models using only training metrics without data, service, and business validation.
- Do not place credentials in pipeline definitions, images, logs, prompts, or model artifacts.

## Security Audit

- Audited at: 2026-08-04T14:30:24.109\+00:00
- Summary: Both static findings are false positives caused by ordinary Markdown prose: a relative document reference and an Azure Event Grid mention. No executable command, reconnaissance instruction, prompt injection, or other intent-level security concern was found.

## Stats

- Views: 116
- Downloads: 11
- Favorites: 0
- Popularity score: 0
