agentsclimarketplace

Ceres publish dataset

Skill AndreaBozzo/Ceres-Claude-Skill/ceres-publish-dataset

Use when publishing a HuggingFace dataset export for the Ceres Open Data Index. Covers the full workflow — data quality checks, Parquet export, snapshot manifest, coverage/quality reports, snapshot changelog, README dataset card update, versioning, and push to HuggingFace.From its SKILL.md

Install
npx -y skills add AndreaBozzo/Ceres-Claude-Skill --skill ceres-publish-dataset

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

9.0 KB, ~2.2k tokens by cl100k_base, as published. Nobody here has run it

Ceres Publish Dataset — HuggingFace Export

Recurring workflow to export the Ceres open data index to HuggingFace as a curated Parquet dataset. Ceres v0.5.0 is the baseline for the reproducible snapshot contract: manifest, integrity checks, coverage/quality reports, alias-aware duplicate provenance, identity index, and snapshot changelogs.

Tracking issue: https://github.com/AndreaBozzo/Ceres/issues/89

Checklist

  1. Run data quality checks (noise, duplicates, new portals)
  2. Export: ceres export --format parquet --output ~/ceres-open-data-index
  3. Remove any stale per-portal parquet files from previous exports (e.g. renamed portals)
  4. Update README.md dataset card (counts, portal table, snapshot date, bias notes)
  5. Version the snapshot (commit + tag)
  6. Push to HuggingFace
  7. (Optional) Regenerate HuggingFace Space visualization

Export Command

ceres export --format parquet --output ~/ceres-open-data-index

The export streams all non-stale datasets from PostgreSQL, applies curation, and writes:

  • all.parquet — Canonical complete flattened dataset
  • data/<portal-name>.parquet — Per-portal subsets (repeat rows from all.parquet; never sum into the canonical total)
  • identity.parquet — Slim per-record fingerprint (source_portal, original_id, content_hash) used to diff snapshots; one row per all.parquet row
  • metadata.json — Versioned snapshot manifest: stable snapshot_id, UTC generated_at, Ceres version/commit, portal-config checksum, duplicate_detection provenance (method/version/alias_groups), curation row counts, per-portal inclusion status, and SHA-256 checksums for every file
  • reports.json — Machine-readable coverage and quality report: coverage by portal/type/profile/language, field-completeness rates (description, license, organization, tags, modification date), and curation outcomes (raw, exported, filtered, duplicate-flagged, duplicate-detection method/version, excluded portals)
  • report.md — Human-readable summary of reports.json for the dataset card / release notes
  • changelog.json / changelog.md — Snapshot-to-snapshot diff (added/changed/removed/unchanged, with per-portal summaries), written only when --previous <DIR> points at a prior snapshot; otherwise a zeroed baseline changelog with compared: false

Verify the SHA-256 checksums in metadata.json before publishing a copied or mirrored snapshot. reports.json is derived from the same export pass, so its figures agree with the manifest.

Diffing against the previous snapshot

To publish a changelog, point --previous at the prior published snapshot directory:

ceres export --format parquet --output ~/ceres-open-data-index --previous ~/ceres-open-data-index

The diff is keyed by the stable identity (source_portal + original_id) read from each snapshot's identity.parquet. "Changed" means the content_hash (SHA-256 of title+description) differs; source modification timestamps are not used because portal coverage of them is incomplete.

Portal names are resolved from ~/.config/ceres/portals.toml. Portals not in the config fall back to hostname-based naming (e.g. https://data.gov.ro becomes data-gov-ro).

The export takes ~30-40 minutes for 900k+ datasets due to JSONB flattening.

Export Schema (14 columns)

ColumnTypeDescription
original_idstringDataset ID from source portal
source_portalstringPortal base URL
portal_namestringHuman-readable portal name
urlstringDirect URL to dataset page
titlestringDataset title
descriptionstringDataset description (nullable)
tagsstringComma-separated tag names (nullable)
organizationstringPublishing organization (nullable)
licensestringLicense title or identifier (nullable)
metadata_createdstringOriginal creation date ISO 8601 (nullable)
metadata_modifiedstringLast modification date ISO 8601 (nullable)
first_seen_atstringWhen Ceres first indexed this dataset (RFC 3339)
languagestringPrimary language code (nullable)
is_duplicatebooleanHeuristic signal: same title (case-insensitive) appears on another portal. Not canonical deduplication

Curation Rules (applied automatically during export)

  • Noise filter: Removes datasets where title < 5 chars, description is empty, or title contains "test"/"prova"/"esempio" (case-insensitive substring match)
  • Duplicate flag (heuristic, not canonical dedup): Same title (case-insensitive) across different portals sets is_duplicate=true. Duplicates are kept, not removed. Portals may declare aliases = [...] in portals.toml; aliased/mirror URLs are folded onto their canonical portal first, so a mirror is not counted as an independent source. The matching rule and version are recorded in metadata.json under duplicate_detection. Core SQL: SELECT LOWER(title) FROM datasets GROUP BY LOWER(title) HAVING COUNT(DISTINCT <canonicalized source_portal>) > 1
  • Metadata flattening: Tags from metadata.tags[].name, organization from metadata.organization.title, license from metadata.license_title

HuggingFace Dataset Repository

  • Location: ~/ceres-open-data-index
  • Remote: https://huggingface.co/datasets/AndreaBozzo/ceres-open-data-index
  • Branch: main
  • LFS: *.parquet tracked via Git-LFS (configured in .gitattributes)

Commit and push workflow

cd ~/ceres-open-data-index
git add all.parquet identity.parquet data/ metadata.json reports.json report.md changelog.json changelog.md README.md
git commit -m "vN: Month Year export (X portals, Yk datasets)"
git tag vN
git push origin main
git push origin vN

README Dataset Card — What to Update

The README at ~/ceres-open-data-index/README.md uses HuggingFace dataset card format with YAML frontmatter. Update these sections:

  1. YAML frontmatter: language list (add new language codes), size_categories, tags
  2. Opening line: Total datasets count, portal count, country count
  3. Files section: Row count, file count
  4. Portal table (### Splits by portal): Add new portals, update all counts, sort by count descending
  5. Curation section: Noise filtering count, duplicate flagging count
  6. Update frequency: Snapshot date
  7. Known biases: Geographic skew percentages, language distribution, portal selection notes
  8. Citation block: Snapshot date, portal count

Curation counts and per-portal totals come from metadata.json; coverage breakdowns (by type/profile/language) and field-completeness rates come from reports.json (or the rendered report.md). Both are generated by the export.

Export History

VersionDateExportedFilteredDuplicatesPortalsCountries
v12026-02-12230,3154,69358,364238 + intl
v22026-02-25349,8366,53259,062259 + intl
v32026-03-30890,14310,33175,5273213 + intl

Key Files

FilePurpose
~/.config/ceres/portals.tomlPortal config — controls name resolution during export
~/ceres-open-data-index/README.mdHuggingFace dataset card
~/ceres-open-data-index/metadata.jsonVersioned snapshot manifest (generated, do not edit)
~/ceres-open-data-index/reports.jsonCoverage and quality report (generated, do not edit)
~/ceres-open-data-index/report.mdHuman-readable coverage/quality summary (generated)
crates/ceres-core/src/parquet_export.rsExport logic (noise filter, duplicate flagging, flattening, manifest, reports)
crates/ceres-db/src/repository.rsSQL queries (duplicate detection, dataset streaming)

HuggingFace Space (Optional)

The visualization dashboard at https://huggingface.co/spaces/AndreaBozzo/Ceres is a static Plotly site generated from the exported data.

  • Location: ~/Documenti/Ceres(huggingfacespace)
  • Generator: ~/Documenti/open-data-galaxy/galaxy_constellation.py
  • Regenerate after pushing the dataset, then push the Space repo separately

Pre-export Data Quality Checks

# Overall stats
ceres stats

# Check noise candidates in DB (optional)
docker exec ceres_db psql -U ceres_user -d ceres_db -c "
  SELECT COUNT(*) FILTER (WHERE LENGTH(title) < 5) as tiny_titles,
         COUNT(*) FILTER (WHERE description IS NULL OR TRIM(description) = '') as empty_desc,
         COUNT(*) FILTER (WHERE LOWER(title) LIKE '%test%' OR LOWER(title) LIKE '%prova%' OR LOWER(title) LIKE '%esempio%') as noise_titles
  FROM datasets WHERE NOT is_stale;
"

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.