Python docx
Skill cxcscmu/SkillLearnBench/skills/b1-one-shot-claude-sonnet-4-6/offer-letter-generator/python-docx
Core python-docx usage for reading, modifying, and saving Word documents including paragraphs, runs, tables, headers, and footers.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill python-docxAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
2.7 KB, 608 tokens by cl100k_base, as published. Nobody here has run it
python-docx Core Usage
Installation
pip install python-docx
Basic Document Operations
from docx import Document
# Open existing document
doc = Document('template.docx')
# Access paragraphs
for para in doc.paragraphs:
print(para.text) # Full paragraph text
for run in para.runs: # Individual styled runs
print(run.text)
# Save
doc.save('output.docx')
Document Structure
A .docx file has:
doc.paragraphs— body paragraphs (does NOT include headers/footers)doc.tables— top-level tables in bodydoc.sections— page sections (each has.headerand.footer)
Paragraphs and Runs
A Paragraph is a block of text. It contains multiple Run objects, each with their own formatting (bold, italic, font size, etc.).
Critical: Word often splits a single logical text segment across multiple runs (due to spell-check marks, formatting, or XML internals). Always read para.text for the full text — never rely on individual runs.
para.text # Full concatenated text of all runs
para.runs # List of Run objects
run.text # Text of this run
run.bold # Bold formatting
run.italic # Italic formatting
run.font.size # Font size (in EMUs; divide by 12700 for pt)
run.font.name # Font name
Tables
for table in doc.tables:
for row in table.rows:
for cell in row.cells:
for para in cell.paragraphs:
print(para.text)
for nested_table in cell.tables: # recurse for nesting
pass
Headers and Footers
for section in doc.sections:
header = section.header
footer = section.footer
for para in header.paragraphs:
print(para.text)
for para in footer.paragraphs:
print(para.text)
Modifying Text While Preserving Formatting
When replacing text, preserve the first run's formatting and clear subsequent runs:
def set_paragraph_text(para, new_text):
"""Replace all run text in a paragraph, keeping first run's formatting."""
if not para.runs:
return
para.runs[0].text = new_text
for run in para.runs[1:]:
run.text = ''
Common Pitfalls
doc.paragraphsdoes NOT include paragraphs inside table cellsdoc.paragraphsdoes NOT include header/footer paragraphs- Split runs are the norm, not the exception — always work at
para.textlevel - Clearing all runs and rebuilding loses per-run formatting (bold/italic on parts of text)
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.