Glean core workflow b
'Execute Glean secondary workflow: bulk document indexing, custom datasource connectors,From its SKILL.md
npx -y skills add jeremylongshore/claude-code-plugins-plus-skills --skill glean-core-workflow-bAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
What its file declares
Copied from the file, not written here
The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
3.5 KB, 773 tokens by cl100k_base, as published. Nobody here has run it
Glean Core Workflow B: Indexing & Connectors
Overview
Build custom Glean connectors: set up datasources, bulk index documents, manage content lifecycle, and configure permissions.
Instructions
Step 1: Create Custom Datasource
await fetch(`${GLEAN}/index/v1/adddatasource`, {
method: 'POST', headers: idxHeaders,
body: JSON.stringify({
name: 'internal_docs',
displayName: 'Internal Documentation',
datasourceCategory: 'PUBLISHED_CONTENT',
urlRegex: 'https://docs.internal.company.com/.*',
isOnPrem: false,
}),
});
Step 2: Bulk Index Documents
// Bulk indexing replaces ALL documents in the datasource
const uploadId = `upload-${Date.now()}`;
// Send documents in batches of 100
for (let i = 0; i < allDocs.length; i += 100) {
const batch = allDocs.slice(i, i + 100);
const isFirst = i === 0;
const isLast = i + 100 >= allDocs.length;
await fetch(`${GLEAN}/index/v1/bulkindexdocuments`, {
method: 'POST', headers: idxHeaders,
body: JSON.stringify({
datasource: 'internal_docs',
uploadId,
isFirstPage: isFirst,
isLastPage: isLast,
documents: batch.map(doc => ({
id: doc.id,
title: doc.title,
url: doc.url,
body: { mimeType: 'text/html', textContent: doc.content },
author: { email: doc.authorEmail },
updatedAt: doc.updatedAt,
permissions: { allowAnonymousAccess: true },
})),
}),
});
console.log(`Indexed batch ${i/100 + 1} (${batch.length} docs)`);
}
Step 3: Set Document Permissions
// Control who can see documents in search results
await fetch(`${GLEAN}/index/v1/indexdocuments`, {
method: 'POST', headers: idxHeaders,
body: JSON.stringify({
datasource: 'internal_docs',
documents: [{
id: 'confidential-001',
title: 'Board Meeting Notes',
url: 'https://docs.internal.company.com/board/q1-2025',
body: { mimeType: 'text/plain', textContent: '...' },
permissions: {
allowedUsers: [{ email: '[email protected]' }, { email: '[email protected]' }],
},
}],
}),
});
Step 4: Delete Documents
// Remove specific documents from the index
await fetch(`${GLEAN}/index/v1/deletedocument`, {
method: 'POST', headers: idxHeaders,
body: JSON.stringify({
datasource: 'internal_docs',
objectType: 'Document',
id: 'doc-to-delete',
}),
});
Error Handling
| Error | Cause | Solution |
|---|---|---|
uploadId already used | Reusing bulk upload ID | Generate unique uploadId per run |
document too large | Content exceeds limit | Truncate body to ~100KB |
invalid permissions | Malformed user/group | Use valid email addresses |
Resources
Next Steps
For common errors, see glean-common-errors.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.