Okra curl
Agent skills for OkraPDF — PDF extraction, document chat, structured data extraction
npx -y skills add okrapdf/skills --skill okra-curlAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
HTTP/curl cookbook for the OkraPDF REST API — workflow-bound document uploads, passive file assets, chat completions, structured extraction, collections, and exports.
SKILL.md
11.2 KB, ~3.2k tokens by cl100k_base, as published. Nobody here has run it
OkraPDF HTTP API
Direct HTTP access to OkraPDF. OpenAI-compatible chat completions, REST endpoints for everything.
Auth
All authenticated endpoints require:
Authorization: Bearer YOUR_API_KEY
Base URL: https://api.okrapdf.com
Choose the right upload surface
POST /v1/documents— upload and immediately start Okra processing (OCR, chat, exports, entities)POST /v1/files— store a PDF as a passive asset without binding it to a workflow
Use documents when you want the PDF parsed right away. Use files when you want Cloudinary-style storage first and will decide later what to do with the PDF.
Upload and process a document
From URL
curl -X POST https://api.okrapdf.com/v1/documents \
-H "Authorization: Bearer $OKRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://arxiv.org/pdf/2307.09288"}'
From file (multipart)
curl -X POST https://api.okrapdf.com/v1/documents \
-H "Authorization: Bearer $OKRA_API_KEY" \
-F "[email protected]" \
-F "page_images=cover"
Upload options (form fields or JSON):
page_images—none,cover(default),lazyprocessor—textlayer,llamaparse,unstructured,azure-distrategy— processing strategy override
Response:
{
"document_id": "doc-abc123",
"phase": "extracting",
"pages_total": 42
}
Passive file assets
Store the original PDF without starting OCR or document lifecycle workflows.
Upload from file (multipart)
curl -X POST https://api.okrapdf.com/v1/files \
-H "Authorization: Bearer $OKRA_API_KEY" \
-F "[email protected];type=application/pdf"
Response:
{
"id": "doc-file123",
"object": "file",
"name": "report.pdf",
"mime": "application/pdf",
"size": 1048576,
"upload_mode": "multipart",
"workflow_bound": false,
"urls": {
"bytes": "https://api.okrapdf.com/v1/files/doc-file123/bytes"
}
}
List files
curl "https://api.okrapdf.com/v1/files?limit=20" \
-H "Authorization: Bearer $OKRA_API_KEY"
Download original PDF bytes
curl "https://api.okrapdf.com/v1/files/doc-file123/bytes" \
-H "Authorization: Bearer $OKRA_API_KEY" \
-o report.pdf
Delete a stored file
curl -X DELETE "https://api.okrapdf.com/v1/files/doc-file123" \
-H "Authorization: Bearer $OKRA_API_KEY"
Large PDFs: direct-to-R2 upload
Use this when you want the browser or client to upload bytes straight to R2 instead of proxying the whole PDF through Okra's Worker.
FILE=report.pdf
SHA256=$(shasum -a 256 "$FILE" | awk '{print $1}')
FILE_SIZE=$(stat -f '%z' "$FILE")
PRESIGN=$(curl -s -X POST https://api.okrapdf.com/v1/files/presign \
-H "Authorization: Bearer $OKRA_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"fileName\":\"report.pdf\",\"fileSize\":$FILE_SIZE,\"sha256\":\"$SHA256\"}")
FILE_ID=$(echo "$PRESIGN" | jq -r '.id')
UPLOAD_URL=$(echo "$PRESIGN" | jq -r '.upload_url')
CONTENT_TYPE=$(echo "$PRESIGN" | jq -r '.upload_headers["Content-Type"]')
curl -X PUT "$UPLOAD_URL" \
-H "Content-Type: $CONTENT_TYPE" \
--data-binary "@$FILE"
curl -X POST https://api.okrapdf.com/v1/files/finalize \
-H "Authorization: Bearer $OKRA_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"id\":\"$FILE_ID\",\"fileName\":\"report.pdf\",\"sha256\":\"$SHA256\"}"
Notes:
- Use exactly the returned
upload_headers. Today that is justContent-Type. POST /v1/files/presignplusPUTplusPOST /v1/files/finalizeis the raw HTTP flow. The SDK hides this behind onefiles.upload()call.workflow_bound: falsemeans the file is stored but not processed yet.
Check Status
curl https://api.okrapdf.com/v1/documents/doc-abc123/status \
-H "Authorization: Bearer $OKRA_API_KEY"
Response (key fields, abbreviated — full payload includes spec, capabilities, cache, etc.):
{"documentId":"doc-abc123","phase":"complete","pagesCompleted":42,"pagesTotal":42,"totalNodes":318,"verifiedNodes":318,"fileName":"report.pdf","pdfSha256":"..."}
Read Content
Full document (markdown)
curl https://api.okrapdf.com/v1/documents/doc-abc123/full.md \
-H "Authorization: Bearer $OKRA_API_KEY"
Specific pages
curl "https://api.okrapdf.com/v1/documents/doc-abc123/pages/3" \
-H "Authorization: Bearer $OKRA_API_KEY"
All pages as JSON
curl https://api.okrapdf.com/v1/documents/doc-abc123/pages \
-H "Authorization: Bearer $OKRA_API_KEY"
Chat Completions (OpenAI-compatible)
Ask a question about a document
curl -X POST https://api.okrapdf.com/document/doc-abc123/chat/completions \
-H "Authorization: Bearer $OKRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "What was total revenue in 2024?"}],
"stream": false
}'
Structured extraction (JSON schema)
curl -X POST https://api.okrapdf.com/document/doc-abc123/chat/completions \
-H "Authorization: Bearer $OKRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Extract all line items from this invoice"}],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "invoice",
"schema": {
"type": "object",
"properties": {
"line_items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"description": {"type": "string"},
"amount": {"type": "number"}
}
}
},
"total": {"type": "number"}
}
},
"strict": true
}
},
"stream": false
}'
Streaming
curl -X POST https://api.okrapdf.com/document/doc-abc123/chat/completions \
-H "Authorization: Bearer $OKRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Summarize this document"}],
"stream": true
}'
Exports
# Markdown
curl https://api.okrapdf.com/exports/doc-abc123/markdown \
-H "Authorization: Bearer $OKRA_API_KEY"
# Excel
curl -o report.xlsx https://api.okrapdf.com/exports/doc-abc123/excel \
-H "Authorization: Bearer $OKRA_API_KEY"
# DOCX
curl -o report.docx https://api.okrapdf.com/exports/doc-abc123/docx \
-H "Authorization: Bearer $OKRA_API_KEY"
# JSON snapshot
curl https://api.okrapdf.com/exports/doc-abc123/snapshot \
-H "Authorization: Bearer $OKRA_API_KEY"
Page Images
Deterministic URLs, CDN-cached:
# Page 1 image
curl -o page1.png https://res.okrapdf.com/v1/documents/doc-abc123/pg_1.png
# With resize transform
curl -o thumb.png "https://res.okrapdf.com/v1/documents/doc-abc123/w_400,h_300/pg_1.png"
Tables and Entities
# List tables
curl https://api.okrapdf.com/v1/documents/doc-abc123/entities/tables \
-H "Authorization: Bearer $OKRA_API_KEY"
# List all entities
curl https://api.okrapdf.com/v1/documents/doc-abc123/entities \
-H "Authorization: Bearer $OKRA_API_KEY"
Collections
Group documents and query across all of them at once.
Create (with optional seed documents)
curl -X POST https://api.okrapdf.com/v1/collections \
-H "Authorization: Bearer $OKRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "Q4 Earnings",
"description": "Quarterly earnings reports",
"document_ids": ["doc-abc123", "doc-def456"]
}'
Add documents
curl -X POST https://api.okrapdf.com/v1/collections/col-xxx/documents \
-H "Authorization: Bearer $OKRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"document_ids": ["doc-ghi789"]}'
Query across documents
Two query modes:
| Mode | Behavior | Best for |
|---|---|---|
fanout (default) | Separate completion per document, NDJSON stream | Per-document answers, structured extraction |
sandbox | Single LLM with grep/Python over all docs | Cross-document search, comparisons, aggregation |
Fanout (default):
curl -X POST https://api.okrapdf.com/v1/collections/col-xxx/query \
-H "Authorization: Bearer $OKRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"query": "What was total revenue in Q4?"}'
Fanout NDJSON response events:
{"type":"start","query_id":"...","doc_count":7}
{"type":"result","doc_id":"doc-xxx","answer":"Apple reported revenue of..."}
{"type":"done","completed":7,"failed":0}
Sandbox mode:
curl -X POST https://api.okrapdf.com/v1/collections/col-xxx/query \
-H "Authorization: Bearer $OKRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"prompt": "Compare R&D spending across all companies. Show a table.",
"mode": "sandbox"
}'
Sandbox returns a single JSON response:
{
"answer": "Based on the 10-K filings...",
"model": "moonshotai/kimi-k2.5",
"mode": "sandbox",
"usage": {"inputTokens": 19364, "outputTokens": 1804, "costUsd": 0.005},
"tool_calls": 5,
"duration_ms": 78166
}
Get collection
curl https://api.okrapdf.com/v1/collections/col-xxx \
-H "Authorization: Bearer $OKRA_API_KEY"
List collections
curl https://api.okrapdf.com/v1/collections \
-H "Authorization: Bearer $OKRA_API_KEY"
Remove documents
curl -X DELETE https://api.okrapdf.com/v1/collections/col-xxx/documents \
-H "Authorization: Bearer $OKRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"document_ids": ["doc-abc123"]}'
Export collection
# NDJSON stream (default)
curl -N "https://api.okrapdf.com/v1/collections/col-xxx/export?format=markdown" \
-H "Authorization: Bearer $OKRA_API_KEY"
# Zip archive
curl -L "https://api.okrapdf.com/v1/collections/col-xxx/export?format=zip" \
-H "Authorization: Bearer $OKRA_API_KEY" -o collection.zip
Delete collection (preserves documents)
curl -X DELETE https://api.okrapdf.com/v1/collections/col-xxx \
-H "Authorization: Bearer $OKRA_API_KEY"
Fan-out pattern (scripting)
For maximum control, query individual documents in parallel:
for doc_id in doc-abc123 doc-def456 doc-ghi789; do
curl -s -X POST "https://api.okrapdf.com/document/$doc_id/chat/completions" \
-H "Authorization: Bearer $OKRA_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"messages\": [{\"role\": \"user\", \"content\": \"What was total revenue?\"}], \"stream\": false}" &
done
wait
List Documents
curl "https://api.okrapdf.com/v1/documents?limit=20" \
-H "Authorization: Bearer $OKRA_API_KEY"
Pagination: ?limit=20&after=cursor_token
Available Vendors
curl https://api.okrapdf.com/v1/vendors
Returns list of available OCR processors with capabilities and pricing tier.
Error Handling
| Status | Meaning |
|---|---|
| 202 | Accepted, processing async |
| 400 | Bad request (check body) |
| 401 | Missing or invalid API key |
| 404 | Document not found |
| 409 | Conflict (document already exists) |
| 429 | Rate limited |
| 500 | Server error |