agentsclimarketplace

Ceres publish dataset

Skill AndreaBozzo/Ceres-Claude-Skill/ceres-publish-dataset

Claude Code Skill for Ceres

Install
npx -y skills add AndreaBozzo/Ceres-Claude-Skill --skill ceres-publish-dataset

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when publishing a HuggingFace dataset export for the Ceres Open Data Index. Covers the full workflow — data quality checks, Parquet export, snapshot manifest, coverage/quality reports, snapshot changelog, README dataset card update, versioning, and push to HuggingFace.

SKILL.md

9.0 KB, as published. Nobody here has run it

Ceres Publish Dataset — HuggingFace Export

Recurring workflow to export the Ceres open data index to HuggingFace as a curated Parquet dataset. Ceres v0.5.0 is the baseline for the reproducible snapshot contract: manifest, integrity checks, coverage/quality reports, alias-aware duplicate provenance, identity index, and snapshot changelogs.

Tracking issue: https://github.com/AndreaBozzo/Ceres/issues/89

Checklist

  1. Run data quality checks (noise, duplicates, new portals)
  2. Export: ceres export --format parquet --output ~/ceres-open-data-index
  3. Remove any stale per-portal parquet files from previous exports (e.g. renamed portals)
  4. Update README.md dataset card (counts, portal table, snapshot date, bias notes)
  5. Version the snapshot (commit + tag)
  6. Push to HuggingFace
  7. (Optional) Regenerate HuggingFace Space visualization

Export Command

ceres export --format parquet --output ~/ceres-open-data-index

The export streams all non-stale datasets from PostgreSQL, applies curation, and writes:

  • all.parquet — Canonical complete flattened dataset
  • data/<portal-name>.parquet — Per-portal subsets (repeat rows from all.parquet; never sum into the canonical total)
  • identity.parquet — Slim per-record fingerprint (source_portal, original_id, content_hash) used to diff snapshots; one row per all.parquet row
  • metadata.json — Versioned snapshot manifest: stable snapshot_id, UTC generated_at, Ceres version/commit, portal-config checksum, duplicate_detection provenance (method/version/alias_groups), curation row counts, per-portal inclusion status, and SHA-256 checksums for every file
  • reports.json — Machine-readable coverage and quality report: coverage by portal/type/profile/language, field-completeness rates (description, license, organization, tags, modification date), and curation outcomes (raw, exported, filtered, duplicate-flagged, duplicate-detection method/version, excluded portals)
  • report.md — Human-readable summary of reports.json for the dataset card / release notes
  • changelog.json / changelog.md — Snapshot-to-snapshot diff (added/changed/removed/unchanged, with per-portal summaries), written only when --previous <DIR> points at a prior snapshot; otherwise a zeroed baseline changelog with compared: false

Verify the SHA-256 checksums in metadata.json before publishing a copied or mirrored snapshot. reports.json is derived from the same export pass, so its figures agree with the manifest.

Diffing against the previous snapshot

To publish a changelog, point --previous at the prior published snapshot directory:

ceres export --format parquet --output ~/ceres-open-data-index --previous ~/ceres-open-data-index

The diff is keyed by the stable identity (source_portal + original_id) read from each snapshot's identity.parquet. "Changed" means the content_hash (SHA-256 of title+description) differs; source modification timestamps are not used because portal coverage of them is incomplete.

Portal names are resolved from ~/.config/ceres/portals.toml. Portals not in the config fall back to hostname-based naming (e.g. https://data.gov.ro becomes data-gov-ro).

The export takes ~30-40 minutes for 900k+ datasets due to JSONB flattening.

Export Schema (14 columns)

ColumnTypeDescription
original_idstringDataset ID from source portal
source_portalstringPortal base URL
portal_namestringHuman-readable portal name
urlstringDirect URL to dataset page
titlestringDataset title
descriptionstringDataset description (nullable)
tagsstringComma-separated tag names (nullable)
organizationstringPublishing organization (nullable)
licensestringLicense title or identifier (nullable)
metadata_createdstringOriginal creation date ISO 8601 (nullable)
metadata_modifiedstringLast modification date ISO 8601 (nullable)
first_seen_atstringWhen Ceres first indexed this dataset (RFC 3339)
languagestringPrimary language code (nullable)
is_duplicatebooleanHeuristic signal: same title (case-insensitive) appears on another portal. Not canonical deduplication

Curation Rules (applied automatically during export)

  • Noise filter: Removes datasets where title < 5 chars, description is empty, or title contains "test"/"prova"/"esempio" (case-insensitive substring match)
  • Duplicate flag (heuristic, not canonical dedup): Same title (case-insensitive) across different portals sets is_duplicate=true. Duplicates are kept, not removed. Portals may declare aliases = [...] in portals.toml; aliased/mirror URLs are folded onto their canonical portal first, so a mirror is not counted as an independent source. The matching rule and version are recorded in metadata.json under duplicate_detection. Core SQL: SELECT LOWER(title) FROM datasets GROUP BY LOWER(title) HAVING COUNT(DISTINCT <canonicalized source_portal>) > 1
  • Metadata flattening: Tags from metadata.tags[].name, organization from metadata.organization.title, license from metadata.license_title

HuggingFace Dataset Repository

  • Location: ~/ceres-open-data-index
  • Remote: https://huggingface.co/datasets/AndreaBozzo/ceres-open-data-index
  • Branch: main
  • LFS: *.parquet tracked via Git-LFS (configured in .gitattributes)

Commit and push workflow

cd ~/ceres-open-data-index
git add all.parquet identity.parquet data/ metadata.json reports.json report.md changelog.json changelog.md README.md
git commit -m "vN: Month Year export (X portals, Yk datasets)"
git tag vN
git push origin main
git push origin vN

README Dataset Card — What to Update

The README at ~/ceres-open-data-index/README.md uses HuggingFace dataset card format with YAML frontmatter. Update these sections:

  1. YAML frontmatter: language list (add new language codes), size_categories, tags
  2. Opening line: Total datasets count, portal count, country count
  3. Files section: Row count, file count
  4. Portal table (### Splits by portal): Add new portals, update all counts, sort by count descending
  5. Curation section: Noise filtering count, duplicate flagging count
  6. Update frequency: Snapshot date
  7. Known biases: Geographic skew percentages, language distribution, portal selection notes
  8. Citation block: Snapshot date, portal count

Curation counts and per-portal totals come from metadata.json; coverage breakdowns (by type/profile/language) and field-completeness rates come from reports.json (or the rendered report.md). Both are generated by the export.

Export History

VersionDateExportedFilteredDuplicatesPortalsCountries
v12026-02-12230,3154,69358,364238 + intl
v22026-02-25349,8366,53259,062259 + intl
v32026-03-30890,14310,33175,5273213 + intl

Key Files

FilePurpose
~/.config/ceres/portals.tomlPortal config — controls name resolution during export
~/ceres-open-data-index/README.mdHuggingFace dataset card
~/ceres-open-data-index/metadata.jsonVersioned snapshot manifest (generated, do not edit)
~/ceres-open-data-index/reports.jsonCoverage and quality report (generated, do not edit)
~/ceres-open-data-index/report.mdHuman-readable coverage/quality summary (generated)
crates/ceres-core/src/parquet_export.rsExport logic (noise filter, duplicate flagging, flattening, manifest, reports)
crates/ceres-db/src/repository.rsSQL queries (duplicate detection, dataset streaming)

HuggingFace Space (Optional)

The visualization dashboard at https://huggingface.co/spaces/AndreaBozzo/Ceres is a static Plotly site generated from the exported data.

  • Location: ~/Documenti/Ceres(huggingfacespace)
  • Generator: ~/Documenti/open-data-galaxy/galaxy_constellation.py
  • Regenerate after pushing the dataset, then push the Space repo separately

Pre-export Data Quality Checks

# Overall stats
ceres stats

# Check noise candidates in DB (optional)
docker exec ceres_db psql -U ceres_user -d ceres_db -c "
  SELECT COUNT(*) FILTER (WHERE LENGTH(title) < 5) as tiny_titles,
         COUNT(*) FILTER (WHERE description IS NULL OR TRIM(description) = '') as empty_desc,
         COUNT(*) FILTER (WHERE LOWER(title) LIKE '%test%' OR LOWER(title) LIKE '%prova%' OR LOWER(title) LIKE '%esempio%') as noise_titles
  FROM datasets WHERE NOT is_stale;
"

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.