agentsclimarketplace

Lineage pii and governance

Skill vaquarkhan/data-engineering-agent-skills/skills/lineage-pii-and-governance

Production-grade Agent Skills for data engineering AI agents: 73 workflows, platform presets, safe backfill/replay, Kafka & Spark reliability, MCP observability, and VS Code/JetBrains installers.

Install
npx -y skills add vaquarkhan/data-engineering-agent-skills --skill lineage-pii-and-governance

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 21 stars21 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Applies governance, lineage, ownership, and sensitive-data controls to data changes. Use when a pipeline touches published datasets, regulated information, or shared business metrics.

SKILL.md

2.4 KB, as published. Nobody here has run it

Lineage, PII, And Governance

Overview

Governance is an engineering concern, not a cleanup exercise. This skill ensures that every meaningful data change accounts for ownership, lineage, access, and sensitive-data handling before release.

When to Use

  • publishing a new table, model, or stream
  • changing business-critical metrics
  • handling personal, financial, health, or otherwise sensitive data
  • modifying access controls or data-sharing patterns
  • changing upstream or downstream lineage

Workflow

  1. Identify ownership and consumers. Every published dataset should have:

    • an owner
    • intended consumers
    • known downstream dependencies
  2. Classify the data. Determine whether fields are:

    • public
    • internal
    • confidential
    • regulated or sensitive
  3. Define required controls. Controls may include:

    • masking
    • tokenization
    • row-level restrictions
    • column-level restrictions
    • encryption requirements
    • retention or deletion rules
  4. Update lineage and documentation. Record how the data flows from source to publish layer, including major transformations.

  5. Verify policy enforcement in implementation. Do not stop at documentation. Check that access and masking rules are actually reflected in code or platform configuration.

Common Rationalizations

RationalizationReality
"It is only internal data."Internal datasets still create exposure, misuse, and compliance risk.
"We will document lineage later."Lineage that is not updated during change work becomes stale immediately.
"Security will handle masking downstream."Sensitive data should be controlled as close to production as possible.

Red Flags

  • published data has no named owner
  • sensitive fields are copied without classification
  • lineage updates are missing for a shared metric
  • access rules are assumed but not enforced

Verification

  • Dataset ownership and consumers are identified
  • Sensitive fields are classified
  • Required controls are implemented or explicitly planned
  • Lineage and documentation reflect the change
  • Governance checks are based on real enforcement, not comments alone

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.