agentsclimarketplace

Ceres

Skill AndreaBozzo/Ceres-Claude-Skill/ceres

Claude Code Skill for Ceres

Install
npx -y skills add AndreaBozzo/Ceres-Claude-Skill --skill ceres

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when working with Ceres — a Rust harvest-first toolkit and public open data metadata index. Covers harvesting and synchronization, optional embedding and search, CKAN, DCAT udata, and SPARQL-backed DCAT portal support, Parquet snapshot manifests/reports/changelogs, Ollama or hosted providers, CLI commands, REST API endpoints, portal configuration, architecture, extending via traits, release workflow, and contributing to the Ceres codebase.

SKILL.md

9.9 KB, as published. Nobody here has run it

Ceres — Harvest-First Toolkit for Open Data Portals

Ceres is centered on harvesting and synchronizing open data metadata. Embeddings, semantic search, exports, and API access are downstream capabilities layered on top of the harvested catalog.

Repository: https://github.com/AndreaBozzo/Ceres License: Apache-2.0 | Rust edition: 2024 | MSRV: 1.88+

Pipeline

Metadata:  Portal URL → PortalClient (fetch) → DeltaDetector (content_hash) → DatasetStore (upsert, no embedding)
Embedding: DatasetStore (pending) → EmbeddingProvider (vector) → DatasetStore (update embedding)
Combined:  HarvestPipeline = HarvestService + EmbeddingService

Harvesting and embedding are decoupled: HarvestService handles metadata, EmbeddingService handles vectors, and HarvestPipeline composes both when you want a combined workflow. Metadata-only harvests require no embedding provider. Each stage is trait-based, so components can be swapped or mocked independently.

Current Product Shape

  • Harvest first and keep metadata synchronized over time
  • Add embeddings later only if you want semantic retrieval
  • Prefer Ollama for local embedding, with Gemini and OpenAI still supported
  • Support CKAN, DCAT-AP udata REST, and SPARQL-backed DCAT portals in the current client factory
  • Publish reproducible Parquet snapshots for the public Open Data Index
  • Expose search, export, and API workflows over the same harvested catalog

Crate Map

CratePurposeKey Exports
ceres-coreBusiness logic, traits, servicesHarvestService, EmbeddingService, HarvestPipeline, SearchService, ExportService, WorkerService, CircuitBreaker, traits
ceres-clientPortal clients and embedding providersCkanClient, DcatClient, GeminiClient, OpenAIClient, OllamaClient, PortalClientFactoryEnum, EmbeddingProviderEnum
ceres-dbPostgreSQL + pgvector repositoryDatasetRepository, HarvestJobRepository
ceres-serverAxum REST API with Swagger UIRoutes, DTOs, bearer auth, OpenAPI/Swagger
ceres-cliCommand-line interfaceharvest, embed, search, export, stats subcommands

Core Traits (ceres-core::traits)

pub trait EmbeddingProvider: Send + Sync + Clone {
    fn name(&self) -> &'static str;
    fn dimension(&self) -> usize;
    fn generate(&self, text: &str) -> impl Future<Output = Result<Vec<f32>, AppError>> + Send;
    fn max_batch_size(&self) -> usize { 1 }
    fn generate_batch(&self, texts: &[String]) -> impl Future<Output = Result<Vec<Vec<f32>>, AppError>> + Send;
}

pub trait PortalClient: Send + Sync + Clone {
    type PortalData: Send;
    fn portal_type(&self) -> &'static str;
    fn base_url(&self) -> &str;
    fn list_dataset_ids(&self) -> impl Future<Output = Result<Vec<String>, AppError>> + Send;
    fn get_dataset(&self, id: &str) -> impl Future<Output = Result<Self::PortalData, AppError>> + Send;
    fn into_new_dataset(data: Self::PortalData, portal_url: &str, url_template: Option<&str>, language: &str) -> NewDataset;
    fn search_modified_since(&self, since: DateTime<Utc>) -> impl Future<Output = Result<Vec<Self::PortalData>, AppError>> + Send;
    fn search_all_datasets(&self) -> impl Future<Output = Result<Vec<Self::PortalData>, AppError>> + Send;
}

pub trait PortalClientFactory: Send + Sync + Clone {
    type Client: PortalClient;
    fn create(&self, portal_url: &str, portal_type: PortalType, language: &str) -> Result<Self::Client, AppError>;
}

pub trait DatasetStore: Send + Sync + Clone {
    fn get_by_id(&self, id: Uuid) -> impl Future<Output = Result<Option<Dataset>, AppError>> + Send;
    fn get_hashes_for_portal(&self, portal_url: &str) -> impl Future<Output = Result<HashMap<String, Option<String>>, AppError>> + Send;
    fn upsert(&self, dataset: &NewDataset) -> impl Future<Output = Result<Uuid, AppError>> + Send;
    fn batch_upsert(&self, datasets: &[NewDataset]) -> impl Future<Output = Result<Vec<Uuid>, AppError>> + Send;
    fn search(&self, query_vector: Vec<f32>, limit: usize) -> impl Future<Output = Result<Vec<SearchResult>, AppError>> + Send;
    fn list_stream<'a>(&'a self, portal_filter: Option<&'a str>, limit: Option<usize>) -> BoxStream<'a, Result<Dataset, AppError>>;
    fn get_last_sync_time(&self, portal_url: &str) -> impl Future<Output = Result<Option<DateTime<Utc>>, AppError>> + Send;
    fn record_sync_status(&self, portal_url: &str, sync_time: DateTime<Utc>, sync_mode: &str, sync_status: &str, datasets_synced: i32) -> impl Future<Output = Result<(), AppError>> + Send;
    fn health_check(&self) -> impl Future<Output = Result<(), AppError>> + Send;
    // + update_timestamp_only, batch_update_timestamps, get_duplicate_titles
    // Stale detection
    fn mark_stale_datasets(&self, portal_url: &str, sync_start: DateTime<Utc>) -> impl Future<Output = Result<u64, AppError>> + Send;
    fn mark_stale_by_exclusion(&self, portal_url: &str, seen_ids: &[String]) -> impl Future<Output = Result<u64, AppError>> + Send;
    // Pending embeddings
    fn list_pending_embeddings(&self, portal_filter: Option<&str>, limit: usize) -> impl Future<Output = Result<Vec<Dataset>, AppError>> + Send;
}

Key Types

TypeModulePurpose
Datasetceres_core::modelsComplete dataset row (id, original_id, source_portal, url, title, description, embedding, metadata, timestamps, content_hash, is_stale)
NewDatasetceres_core::modelsInsert/update DTO. Has compute_content_hash() for delta detection
DatasetSchemaceres_core::schemaNormalized resource/distribution metadata derived from preserved portal metadata
SearchResultceres_core::modelsDataset + similarity_score (0.0-1.0)
DatabaseStatsceres_core::modelstotal_datasets, datasets_with_embeddings, stale_datasets, total_portals, last_update
HarvestJobceres_core::jobQueued harvest job with status, retry info, portal config
JobStatusceres_core::jobEnum: Pending, Running, Completed, Failed, Cancelled
SyncStatsceres_core::synccreated, updated, unchanged, failed, skipped counts
SyncOutcomeceres_core::syncPer-dataset outcome: Created, Updated, Unchanged, Failed, Skipped
BatchHarvestSummaryceres_core::syncAggregated results from batch harvesting multiple portals
PortalEntryceres_core::configPortal config: name, url, type, enabled, url_template, language, profile, sparql_endpoint
AppErrorceres_core::errorError enum with is_retryable() and should_trip_circuit()
EmbeddingStatsceres_core::embeddingembedded, failed, skipped, total counts from an embedding run
HarvestPipelineceres_core::pipelineComposes HarvestService + EmbeddingService for combined harvest-then-embed
CircuitBreakerceres_core::circuit_breakerClosed -> Open -> HalfOpen state machine

Quick Start

# Install
cargo install ceres-search

# Start PostgreSQL + pgvector
docker compose up db -d

# Configure
cp .env.example .env

# Run migrations
make migrate

# Harvest metadata without embeddings
ceres harvest https://dati.comune.milano.it --metadata-only

# Or harvest a DCAT portal
ceres harvest https://data.public.lu --type dcat --metadata-only

# Or harvest a SPARQL-backed DCAT catalog
ceres harvest https://data.europa.eu --type dcat --profile sparql --metadata-only

# Optional: local embeddings through Ollama
export EMBEDDING_PROVIDER=ollama
ceres embed

# Search
ceres search "trasporto pubblico" --limit 5

# Export
ceres export --format jsonl > datasets.jsonl

# Stats
ceres stats

Reference Guides

TopicFileWhen to Read
Architecture deep-divereferences/architecture.mdUnderstanding crate graph, services, error handling, database schema
CLI & REST APIreferences/cli-and-server.mdRunning CLI commands, calling API endpoints, env vars, deployment
Harvesting systemreferences/harvesting.mdTwo-tier optimization, delta detection, streaming, circuit breaker
Extending Ceresreferences/extending.mdImplementing custom EmbeddingProvider, PortalClient, or DatasetStore
Contributingreferences/contributing.mdDev setup, testing, CI, code style

Version Notes

  • Current version: 0.5.0 (release, 2026-06-26)
  • crates.io package: ceres-search
  • Harvesting and embedding are decoupled: --metadata-only harvests without API key, embed command generates embeddings separately
  • Ollama is the preferred local embedding path; Gemini and OpenAI remain available
  • Current portal client factory supports CKAN and DCAT (udata_rest default profile plus sparql profile)
  • Stale dataset detection: datasets removed from portals are soft-marked (is_stale) during full syncs
  • Supports Ollama, Gemini, and OpenAI embeddings
  • Parquet export publishes a portable snapshot: all.parquet (canonical), per-portal subsets, identity.parquet, a versioned snapshot manifest (metadata.json with snapshot_id, provenance, alias-aware duplicate metadata, and SHA-256 checksums), coverage/quality reports (reports.json, report.md), and snapshot changelogs (changelog.json, changelog.md when --previous is supplied)
  • v0.6.0 milestone focus: portal coverage expansion in priority order — DCAT profile cleanup, Project Open Data data.json, Socrata, OpenDataSoft, ArcGIS Hub
  • v0.7.0 milestone focus: resource-level metadata depth tracked in issue #68
  • HuggingFace dataset: AndreaBozzo/ceres-open-data-index

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.