agentsclimarketplace

Troubleshoot

Skill pol-cc/agentic-data-engineer/skills/troubleshoot

A Claude Code harness that turns a session into an agentic data engineer for SMBs — packaged as an installable plugin, built from a skillpack of skills that stand up a cheap, self-hostable Modern Data Stack (Tailscale + dlt + BigQuery + dbt + optional MCP), end-to-end and headless.

Install
npx -y skills add pol-cc/agentic-data-engineer --skill troubleshoot

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Diagnose pipeline issues by reading logs and state across Airbyte, dbt, BigQuery, the VPS, and Tailscale. Invoke when verify-pipeline reports a failure, the user says 'something's broken', or a sync hasn't run.

SKILL.md

5.1 KB, as published. Nobody here has run it

troubleshoot

Status: v0.10.0 — references written; diagnostic playbook operational. Now covers the dlt-era ingest failure modes — the silent data gap (a mis-set incremental cursor that skips rows without crashing), a partial load (_dlt_loads status non-zero), a source-vs-destination reconciliation mismatch, and a systemd timer that didn't fire — alongside the original Airbyte/dbt/BQ/MCP modes.

What this skill does

Walks through the most common failure modes of the MDS in order of likelihood, gathering evidence from each layer. The agent reasons over the evidence to identify the cause and proposes (but does not execute) a fix. The user confirms before any change is applied.

Preflight

if [ ! -f .agentic-data-engineer.json ]; then
  echo "[abort] not a managed MDS deployment"
  exit 1
fi

Standard diagnostic flow

Run checks in this order — earlier failures often explain later ones:

1. Tailscale reachability

# Default: Tailscale SSH (keyless). Fallback if the tailnet is down: ssh -i ~/.ssh/<client>_vps deploy@<host>
ssh deploy@<vps-tailscale-hostname> "tailscale status"

If unreachable: the VPS is offline, Tailscale on the VPS is down, or Tailscale on the laptop is down.

2. VPS processes and timers

ssh ... "uptime && free -h && df -h /"                         # load, RAM, disk
ssh ... "systemctl list-timers 'dlt-*' 'dbt-*' --all --no-pager"  # default: dlt + dbt timers (NEXT/LAST)
ssh ... "journalctl -u dlt-<source>.service -u dbt-run.service --since '2 days ago' --no-pager | tail -40"
ssh ... "docker ps --format '{{.Names}}\t{{.Status}}'"         # MCP, etc.
# If stack.ingest == airbyte (alternative path):
ssh ... "abctl local status"                                   # Airbyte controller — airbyte path only
ssh ... "systemctl is-active cron"                             # cron daemon — only if dbt runs on cron

3. Ingest jobs

Default (dlt): no ingest API — read the dlt-<source>.service journal (Step 2) and the _dlt_loads status + reconciliation in Step 4. If stack.ingest == airbyte (alternative path), get a token then list recent jobs:

# Airbyte path only — get token, then list recent jobs
curl ... /api/public/v1/jobs?limit=20

Look for: failed status, cancelled, or jobs that haven't started in 24h+.

4. BigQuery state

bq query --use_legacy_sql=false "
  SELECT table_name, TIMESTAMP_MILLIS(last_modified_time) AS last_mod
  FROM \`<project>.<dataset>.__TABLES__\`
  ORDER BY last_modified_time DESC
"

5. dbt last run

ssh ... "cat /root/dbt/<project>/target/run_results.json | jq '.results[] | select(.status != \"success\")'"

Common failure modes

Catalogued in references/common-failures.md:

  • Silent data gap — a dlt incremental cursor mis-set skips rows without erroring; only reconciliation (source vs destination) catches it. The dlt stack's signature failure.
  • dlt partial load_dlt_loads latest status != 0; the load died mid-write, leaving a partial package. Re-run to recover.
  • Reconciliation mismatch source vs destination — counts disagree beyond tolerance; diagnose by sign (dest < source = gap; dest > source = duplicates or source purge).
  • systemd timer didn't fire — the dlt load never ran (timer disabled, service failed, wrong OnCalendar, or no Persistent=true).
  • Tailscale on on-prem server marked offline (Windows reboot, service stopped)
  • Airbyte abctl controller killed by OOM on small VPS (when stack.ingest == "airbyte")
  • Airbyte API 403 because the client_id/client_secret was rotated and the marker still references the old one
  • BigQuery quota exceeded (free tier crossed)
  • dbt failed because stg_* ran before the ingest load completed (race condition — fix is reschedule)
  • GA4 export missing today's table (Google delay, not an error)
  • MCP write-tools PR failing (PAT expired or missing pull_requests:write) — write tools are off by default and open PRs, not pushes

Output

A markdown summary: which layer failed, the evidence, the proposed fix, and a yes/no question for the user. Nothing is changed until the user confirms.

References

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.