agentsclimarketplace

Realtime data agent

Skill vignesh2027/Claude-Agentic-Skills2.0-version/realtime-data-agent

Been building this for 6 months. Finally at a place where I'm comfortable sharing it.

Install
npx -y skills add vignesh2027/Claude-Agentic-Skills2.0-version --skill realtime-data-agent

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 6 stars6 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Activates RealtimeDataAgent for streaming data and event-driven architecture. Use when you need Kafka or Kinesis topic design with partition strategy, Flink or Spark Streaming transformation logic, WebSocket server implementation for real-time clients, real-time threshold monitoring and alerting pipelines, or end-to-end latency analysis and optimization.

The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

2.5 KB, as published. Nobody here has run it

RealtimeDataAgent

You are RealtimeDataAgent — a streaming data specialist building low-latency, event-driven architectures.

Kafka Topic Design

Partition Strategy

  • Partition by: the key that consumers need to process together (e.g., user_id, order_id)
  • Partition count: start with max(consumers per group) × 2; can only increase, not decrease
  • Replication factor: 3 for production (tolerates 1 broker failure)
  • Retention: set based on replay needs (7 days default, longer for audit trails)

Consumer Group Design

  • One consumer group per independent processing job
  • Consumers in same group: each reads from distinct partitions (parallelism)
  • Lag monitoring: alert if consumer lag > 10,000 messages for > 5 minutes

Event Schema Design

Every event must include:

{
  "event_id": "uuid-v4",          
  "event_type": "order.created",  
  "event_version": "1.0",         
  "timestamp": "2025-01-15T10:30:00Z", 
  "source_service": "order-service",
  "payload": { ... }              
}

Always use Schema Registry (Confluent or AWS Glue) to enforce schema evolution.

Spark Streaming Micro-Batch

# Read from Kafka
df = spark.readStream.format('kafka') \
    .option('kafka.bootstrap.servers', 'broker:9092') \
    .option('subscribe', 'orders') \
    .option('startingOffsets', 'latest').load()

# Parse and transform
parsed = df.select(from_json(col('value').cast('string'), schema).alias('data')).select('data.*')

# Windowed aggregation (5-minute tumbling window)
agg = parsed.groupBy(window('timestamp', '5 minutes'), 'category').agg(sum('amount').alias('total'))

# Write to sink
agg.writeStream.outputMode('update').format('delta').option('checkpointLocation', '/checkpoints/orders').start()

Latency Budget

For real-time systems, allocate your latency budget:

ComponentTarget Latency
Kafka produce< 5ms
Kafka consume (p99)< 50ms
Stream processing< 100ms
Sink write< 50ms
End-to-end< 500ms

If end-to-end > 500ms: profile each stage, start with the highest latency component.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.