Capture mongodb create
Agent skills for setting up and operating Estuary data pipelines through your AI assistant.
npx -y skills add estuary/agent-skills --skill capture-mongodb-createAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 6 stars6 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Create a MongoDB CDC capture using flowctl. Use when setting up real-time streaming from MongoDB Atlas, DocumentDB, or self-hosted MongoDB. Use when user says "capture MongoDB", "stream from Mongo", "MongoDB CDC", or "connect MongoDB to Estuary".
SKILL.md
7.9 KB, as published. Nobody here has run it
Create MongoDB Capture
Create a MongoDB capture using flowctl to stream data from MongoDB collections into Estuary collections using Change Data Capture (CDC).
Applies to: source-mongodb, source-amazon-documentdb, source-azure-cosmos-db
Step 0: Load Connector Documentation
Before proceeding, fetch the official connector docs for prerequisites, config reference, and deployment-specific setup.
Always load the main page: https://docs.estuary.dev/reference/Connectors/capture-connectors/MongoDB/
Then load the variant subpage based on the user's deployment:
| Deployment | Docs URL |
|---|---|
| MongoDB Atlas | Main page covers this |
| Self-hosted MongoDB | Main page covers this |
| Amazon DocumentDB | https://docs.estuary.dev/reference/Connectors/capture-connectors/MongoDB/amazon-documentdb/ |
| Azure Cosmos DB | https://docs.estuary.dev/reference/Connectors/capture-connectors/MongoDB/azure-cosmosdb/ |
Use WebFetch to load these pages. Together they cover:
- Prerequisites (replica set requirement, user permissions)
- Full config property reference
- Capture modes (Change Stream Incremental, Batch Snapshot, Batch Incremental)
- SSH tunnel configuration
- Network access / IP allowlisting
This skill provides the flowctl workflow and troubleshooting that docs don't cover.
Step 1: Gather Requirements
Before writing any YAML, ask the user:
- Deployment type? — MongoDB Atlas, self-hosted replica set, Amazon DocumentDB, or Azure Cosmos DB
- Network path? — Direct connection (Atlas/cloud with IP allowlist), SSH tunnel (private network), Private Link (AWS/Azure/GCP), or ngrok (local dev)
- Non-default data plane? — Most users use the default. Ask if they need a non-default data plane.
- Database and collections? — Which database, all collections or specific subset
- Capture mode? — Change Stream Incremental (default CDC), Batch Snapshot, or Batch Incremental
Critical check: MongoDB CDC requires a replica set. Atlas and DocumentDB always are. Self-hosted standalone will NOT work — must be converted to replica set first.
Step 2: Find the Correct Connector Version
Always use the latest numbered version tag. Query the connector registry to find it:
flowctl raw get --table connector_tags \
--query 'documentation_url=ilike.*source-mongodb*' \
--query 'select=image_tag,documentation_url' \
--output yaml
Choose the connector image:
| Deployment | Connector Image |
|---|---|
| MongoDB Atlas / Self-hosted | ghcr.io/estuary/source-mongodb |
| Amazon DocumentDB | ghcr.io/estuary/source-mongodb (same connector, different config) |
| Azure Cosmos DB | ghcr.io/estuary/source-mongodb (same connector, different config) |
Step 3: Help User Complete Prerequisites
Walk the user through prerequisites from the docs loaded in Step 0:
- Replica set —
rs.status()should return replica set info, not an error - User permissions — needs
readrole on target database andreadonlocal(for oplog) - Oplog retention — at least 24 hours recommended to avoid forced re-backfills
- Network access — Estuary IPs allowlisted, or SSH tunnel configured
Step 4: Create the Capture Spec File
Build flow.yaml using the config reference from the docs. Minimal required config:
captures:
<tenant>/<path>/source-mongodb:
endpoint:
connector:
image: ghcr.io/estuary/source-mongodb:<version>
config:
address: "<connection_string>"
database: "<database_name>"
user: "<username>"
password: "<password>"
bindings: []
Important: The user and password fields are required even if MongoDB auth is disabled.
For SSH tunnel, add networkTunnel.sshForwarding block — see docs for full config.
Connection String Formats
These are critical and easy to get wrong — not fully covered in docs:
# MongoDB Atlas (SRV)
mongodb+srv://cluster0.xxxxx.mongodb.net/?authSource=admin
# MongoDB Atlas (standard)
mongodb://shard-00-00.xxxxx.mongodb.net:27017,.../?ssl=true&replicaSet=atlas-xxxxx&authSource=admin
# Self-hosted (single node replica set)
mongodb://hostname:27017/?authSource=admin&directConnection=true
# Amazon DocumentDB
mongodb://docdb-cluster.xxxxx.us-east-1.docdb.amazonaws.com:27017/?ssl=true&replicaSet=rs0&retryWrites=false
# Via ngrok (local dev)
mongodb://0.tcp.ngrok.io:12345/?authSource=admin&directConnection=true
Key parameters:
authSource=admin— required when user is defined in admin dbdirectConnection=true— use for single-node connectionsssl=true— required for Atlas and DocumentDB
Step 5: Discover and Publish
# Discover collections
flowctl discover --source flow.yaml
# Review the generated bindings
cat flow.yaml
# Publish the capture
flowctl catalog publish --source flow.yaml --auto-approve
Step 6: Verify
# Check status (expect PENDING → BACKFILLING → OK: Streaming Change Events)
flowctl catalog status <tenant>/<path>/source-mongodb
# View recent logs
flowctl logs --task <tenant>/<path>/source-mongodb --since 5m | jq -c '{ts, message}'
# Read captured data
flowctl collections read --collection <tenant>/<path>/<database>/<collection> --uncommitted | head -10
Status progression:
PENDING— normal for ~30 seconds during shard assignmentBACKFILLING— initial snapshot of collectionsOK: Streaming Change Events— CDC running normally
Troubleshooting
"not a replica set" or "change stream not supported"
Cause: MongoDB is standalone, not a replica set
Fix: Atlas/DocumentDB are always replica sets. For self-hosted:
# Add to mongod.conf, restart, then:
mongo --eval "rs.initiate()"
"not authorized" or "Authentication failed"
Cause: Invalid credentials or missing permissions
Fix:
- Verify username/password
- Check
authSourceparameter (usuallyadmin) - Grant required roles:
db.grantRolesToUser("flow_capture", [
{ role: "read", db: "target_database" },
{ role: "read", db: "local" }
])
Missing authSource=admin in connection string
Cause: User authenticates against admin db but authSource not specified
Fix: Add ?authSource=admin to connection string.
"server selection error" or "no reachable servers"
Cause: Incorrect connection string or network issues
Fix:
- Verify connection string format matches deployment type
- For Atlas SRV records, ensure DNS resolution works
- Check if SSL/TLS is required (
ssl=true)
"resume token not found" or forced re-backfill
Cause: Oplog rolled over while connector was paused/stopped
Impact: Connector must re-snapshot all data (happens automatically)
Prevention: Increase oplog size, keep retention at least 24 hours, don't pause captures for extended periods.
"user and password are required"
Cause: Config missing user/password fields
Fix: Always include user and password even if MongoDB auth is disabled — use placeholder values.
Understanding MongoDB "update" semantics
In MongoDB, deleting a field from a document appears as an "update" event, not a delete. The captured document reflects the new state without that field.
Capture stuck in PENDING
Wait 30-60 seconds — this is normal during shard assignment. If still stuck:
flowctl logs --task <tenant>/<path>/source-mongodb --since 5m | jq 'select(.level == "error" or .level == "warn")'
Related Skills
connector-disable-enable— Pause/restart existing capturesconnector-delete-recreate— Nuclear option for stuck capturesestuary-logs— Deep log analysisestuary-catalog-status— Status checking