document
document
¶
Document and closely related objects.
Document
¶
Document(element: CT_Document, part: DocumentPart)
Bases: ElementProxy
WordprocessingML (WML) document.
Not intended to be constructed directly. Use docx.Document to open or create
a document.
Source code in src/docx/document.py
alt_chunks
property
¶
alt_chunks: List[AltChunk]
The AltChunk objects in the document body, in document order.
Only alt-chunks that are direct children of the body appear here; the schema also allows one inside a table cell or other block container.
comments
property
¶
comments: Comments
A Comments object providing access to comments added to the document.
custom_properties
property
¶
custom_properties: CustomProperties
A CustomProperties object providing the arbitrary named values attached to this document.
Behaves as a mutable mapping of name to value. The part holding them is created
the first time this is used, so a document that never touches it gains no
/docProps/custom.xml.
core_properties
property
¶
A CoreProperties object providing Dublin Core properties of document.
extended_properties
property
¶
An ExtendedProperties object providing the application-specific properties of the document, such as word count and producing application.
footnotes
property
¶
footnotes: Footnotes
A Footnotes object providing access to the footnotes of this document.
The footnotes part is created the first time this is used, so a document that
never touches it gains no /word/footnotes.xml.
custom_xml_parts
property
¶
custom_xml_parts: Tuple[CustomXmlPart, ...]
The custom XML data store items of this document, in relationship order.
Each part offers .item_id, .schema_refs, .element and .xml. The item
content is arbitrary caller-supplied XML, so .element is a plain parsed tree
with no element classes of its own.
has_macros
property
¶
True when this document carries a VBA project.
The cheap predicate; vba_project is what reads the bytes.
vba_project
deletable
property
writable
¶
The macro project of this document as bytes, or None when it has none.
A .docm or .dotm carries its macros in word/vbaProject.bin, an OLE
compound file. This library does not parse it, but it round-trips untouched, so
the two operations people actually want are expressible:
Strip the macros from a document received from elsewhere:
Transplant a project authored in Word into a generated document:
Assigning switches the main part to the macro-enabled content type, and removing switches it back. Word silently ignores macros in a document whose main part does not claim to be macro-enabled, and warns the user about macros in one that claims to be but is not, so the two are kept in step rather than left to the caller.
Note this sets the content type; it does not choose the file extension for you.
A macro-enabled document conventionally has a .docm extension.
endnotes
property
¶
endnotes: Endnotes
An Endnotes object providing access to the endnotes of this document.
The endnotes part is created the first time this is used, so a document that
never touches it gains no /word/endnotes.xml.
fields
property
¶
fields: List[Field]
A Field for each field in the document body, in document order.
Outermost first, so a PAGEREF nested in a table-of-contents entry follows the
TOC field containing it. Fields in a header or footer are not in the document
part and so are not included; reach those through the paragraphs of the header
or footer.
form_fields
property
¶
form_fields: List[FormField]
A FormField instance for each legacy form field in the document body.
Fields appear in document order, including those inside tables. Fields in a header or footer are not in the document part and so are not included; reach those through the paragraphs of the header or footer.
floating_shapes
property
¶
The FloatingShapes collection for this document.
A floating shape is anchored rather than inline: it is positioned against the page, the margin, the column or the paragraph, and text wraps around it. These do not appear in inline_shapes, whose position properties would be meaningless for them.
embedded_objects
property
¶
embedded_objects: List[EmbeddedObject]
The OLE objects embedded in the document body, in document order.
An embedded object is a whole file carried inside the document — a spreadsheet, a PDF, another document — which Word opens in its own application on double-click. Extracting them is the useful half:
for obj in document.embedded_objects:
if obj.blob is not None:
Path(obj.filename or "attachment").write_bytes(obj.blob)
Objects in a header, a footer or a footnote belong to those parts and are not included; reach them through the container concerned.
images
property
¶
images: Tuple[Image, ...]
The distinct images embedded in this document's body, in relationship order.
This is the package-level view, the counterpart of reaching an image through the shape that displays it. Several shapes can share one image part, so this is shorter than inline_shapes whenever a picture is used twice, and it includes images no shape displays — a picture left behind when its paragraph was deleted, for instance.
Only images related from the main document part appear here. A picture in a header, a footer or a comment belongs to that part's relationships instead.
A linked image is not included: its bytes are not in the package. Neither is a
relationship of image type whose target is not an image part, which does occur —
see the same guard in Package._gather_image_parts().
theme
property
¶
theme: Theme | None
The document's Theme, or None when it carries no theme part.
The theme is where a theme typeface token such as "minorHAnsi" becomes a
real font name, and where a theme colour becomes an RGB value:
For the large class of documents that set no explicit w:rFonts/@w:ascii
anywhere, this is the only place the typeface the text is actually rendered in
can be found; see also Font.theme_typeface.
inline_shapes
property
¶
The InlineShapes collection for this document.
An inline shape is a graphical object, such as a picture, contained in a run of text and behaving like a character glyph, being flowed like other text in a paragraph.
content_controls
property
¶
content_controls: List[ContentControl]
The structured document tags (content controls) in the document body.
In document order, outermost first. The content of a control appears in
.paragraphs, .tables and .iter_inner_content() as though the wrapper were
not there; this is how the wrapper itself is reached.
math
property
¶
math: List[Math]
The equations in the document body, in document order.
Equations in a header, a footer, a footnote or a comment are in those parts rather than the body and are not included; reach them through the container concerned. See Paragraph.math for why equation text is not part of Paragraph.text.
numbering
property
¶
numbering: Numbering
A Numbering object providing access to the list definitions of this document.
The numbering part is created the first time this is used, so a document that
never touches it gains no /word/numbering.xml.
list_numbers
property
¶
list_numbers: List[tuple[Paragraph, str]]
(paragraph, number) for each list paragraph in the body, in document order.
The number is what a reader sees — "1.", "a)", "iii." — which Word computes from
numbering.xml at display time rather than storing in the body:
Paragraphs inside tables are included, since they count towards the same lists. This walks the document once, which is why it exists alongside Paragraph.list_number: reading that for every paragraph is quadratic.
paragraphs
property
¶
paragraphs: List[Paragraph]
The Paragraph instances in the document, in document order.
A paragraph wrapped in a w:sdt (content control) appears in this list, in the
position of its wrapper.
A revision mark such as w:ins or w:del wraps runs rather than paragraphs, so
it does not affect which paragraphs appear here; it affects their text. See
Paragraph.text and Paragraph.original_text.
is_template
property
¶
True when this document is a Word template, a .dotx or .dotm.
A template holds the same markup as a document and differs only in the content type of its main part, which is what tells Word to start a new document from it rather than open it for editing.
revisions
property
¶
revisions: List[Revision]
A Revision for each tracked change in the document body, in document order.
Empty for a document that has not been through review. Revisions in a header, footer or footnote are not in the document part and so are not included; reach those through the paragraphs of the story concerned.
sections
property
¶
sections: Sections
Sections object providing access to each section in this document.
watermarks
property
¶
watermarks: List[Watermark]
Every watermark in the document, in section and header order.
Empty when the document has none. Ordinarily one per header rather than one per document, since a watermark is a shape in a header and each header carries its own.
settings
property
¶
settings: Settings
A Settings object providing access to the document-level settings.
tables
property
¶
tables: List[Table]
All Table instances in the document, in document order.
Note that only tables appearing at the top level of the document appear in this
list; a table nested inside a table cell does not appear. A table wrapped in a
w:sdt (content control) does appear. A row marked as inserted or deleted
appears as an ordinary row; see revisions.
add_alt_chunk
¶
add_alt_chunk(
chunk: bytes | str | PathLike[str] | IO[bytes],
content_type: str,
) -> AltChunk
Return an AltChunk newly added at the end of the document body.
chunk is the embedded document, given as bytes, as a path to a file (a string
or os.PathLike), or as a file-like object open for binary read.
content_type states its format, e.g. "text/html", "application/rtf" or
"application/vnd.openxmlformats-officedocument.wordprocessingml.document";
Word chooses an importer from it, so it must be right.
Word performs the import when it opens the document, which means the embedded
content is not visible to this library. Its paragraphs and tables do not appear
in Document.paragraphs, Document.tables or Document.iter_inner_content(),
and it contributes no styles, numbering or images to this document until Word
has rewritten the file.
Source code in src/docx/document.py
add_comment
¶
add_comment(
runs: Run | Sequence[Run],
text: str | None = "",
author: str = "",
initials: str | None = "",
) -> Comment
Add a comment to the document, anchored to the specified runs.
runs can be a single Run object or a non-empty sequence of Run objects. Only the
first and last run of a sequence are used, it's just more convenient to pass a whole
sequence when that's what you have handy, like paragraph.runs for example. When runs
contains a single Run object, that run serves as both the first and last run.
A comment can be anchored only on an even run boundary, meaning the text the comment "references" must be a non-zero integer number of consecutive runs. The runs need not be contiguous per se, like the first can be in one paragraph and the last in the next paragraph, but all runs between the first and the last will be included in the reference.
The comment reference range is delimited by placing a w:commentRangeStart element before
the first run and a w:commentRangeEnd element after the last run. This is why only the
first and last run are required and why a single run can serve as both first and last.
Word works out which text to highlight in the UI based on these range markers.
text allows the contents of a simple comment to be provided in the call, providing for
the common case where a comment is a single phrase or sentence without special formatting
such as bold or italics. More complex comments can be added using the returned Comment
object in much the same way as a Document or (table) Cell object, using methods like
.add_paragraph(), .add_run()`, etc.
The author and initials parameters allow that metadata to be set for the comment.
author is a required attribute on a comment and is the empty string by default.
initials is optional on a comment and may be omitted by passing None, but Word adds an
initials attribute by default and we follow that convention by using the empty string
when no initials argument is provided.
Source code in src/docx/document.py
add_caption
¶
add_caption(
label: str,
text: str = "",
*,
style: str | None = "Caption",
separator: str = " ",
restart_at_heading_level: int | None = None,
before: Paragraph | None = None,
) -> Caption
Add a numbered, cross-referenceable caption and return it.
Word numbers each label series independently and renumbers the whole series
when one is inserted, which is the point of using a SEQ field rather than a
typed number. The number is therefore not in the document until Word computes
it; set Settings.update_fields_on_open to have it do so on open.
The caption is bookmarked with a _Ref-prefixed name and the returned object
carries it, so a cross-reference is a one-liner:
caption = document.add_caption("Figure", "Cross-section of the assembly")
document.add_paragraph().add_field(
fields.cross_reference(caption.bookmark_name)
)
The _Ref naming is not decoration: Word's own cross-reference dialogue offers
only targets whose bookmark name follows it, so a caption bookmarked with an
arbitrary name is one the user cannot reference from the UI.
style is the paragraph style, "Caption" by default, which is what Word uses;
pass None to leave the paragraph unstyled. separator goes between the number
and text. restart_at_heading_level restarts the numbering at each heading of
that level, giving the "Figure 2-1" style. before places the caption
immediately before an existing paragraph, which is where a table caption goes.
Source code in src/docx/document.py
add_heading
¶
Return a heading paragraph newly added to the end of the document.
The heading paragraph will contain text and have its paragraph style
determined by level. If level is 0, the style is set to Title. If level
is 1 (or omitted), Heading 1 is used. Otherwise the style is set to Heading
{level}. Raises ValueError if level is outside the range 0-9.
Source code in src/docx/document.py
add_paragraph
¶
add_paragraph(
text: str = "",
style: str | ParagraphStyle | None = None,
) -> Paragraph
Return paragraph newly added to the end of the document.
The paragraph is populated with text and having paragraph style style.
text can contain tab (\t) characters, which are converted to the
appropriate XML form for a tab. text can also include newline (\n) or
carriage return (\r) characters, each of which is converted to a line
break.
Source code in src/docx/document.py
add_picture
¶
add_picture(
image_path_or_stream: str | PathLike[str] | IO[bytes],
width: int | Length | None = None,
height: int | Length | None = None,
description: str | None = None,
title: str | None = None,
svg_fallback: str
| PathLike[str]
| IO[bytes]
| None = None,
honor_exif_orientation: bool = True,
)
Return new picture shape added in its own paragraph at end of the document.
The picture contains the image at image_path_or_stream, scaled based on
width and height. If neither width nor height is specified, the picture
appears at its native size. If only one is specified, it is used to compute a
scaling factor that is then applied to the unspecified dimension, preserving the
aspect ratio of the image. The native size of the picture is calculated using
the dots-per-inch (dpi) value specified in the image file, defaulting to 72 dpi
if no value is specified, as is often the case.
description is the picture's alternative text, which is what a screen reader
announces and what an accessibility check looks for; title is the separate
caption-like field Word writes alongside it.
svg_fallback is the raster image shown in place of an SVG wherever the vector
source cannot be rendered, and honor_exif_orientation applies a photo's EXIF
Orientation as a rotation in the DrawingML; see Run.add_picture() for both.
Source code in src/docx/document.py
add_section
¶
Return a Section object newly added at the end of the document.
The optional start_type argument must be a member of the WdSectionStart
enumeration, and defaults to WD_SECTION.NEW_PAGE if not provided.
Source code in src/docx/document.py
add_table
¶
add_table(
rows: int,
cols: int,
style: str | _TableStyle | None = None,
*,
title: str | None = None,
description: str | None = None,
)
Add a table having row and column counts of rows and cols respectively.
style may be a table style object or a table style name. If style is None,
the table inherits the default table style of the document.
description is the table's alternative text, which is what a screen reader
announces and what an accessibility check looks for. title is the separate,
caption-like field Word writes alongside it. Both are omitted from the XML when
None.
Source code in src/docx/document.py
bookmarks
¶
bookmarks() -> Bookmarks
The Bookmarks in this document, in document order.
Bookmarks Word maintains for itself, such as _GoBack and the _Toc… anchors,
are left out of the collection; reach them through .iter_all().
Source code in src/docx/document.py
cleanup
¶
cleanup(
*,
styles: bool = True,
numbering: bool = True,
media: bool = True,
latent_styles: bool = False,
keep: Tuple[str, ...] = (),
) -> CleanupResult
Remove what this document carries that nothing points at; report what went.
A document created by this library defines 164 styles and references one, and carries numbering definitions for lists it does not have:
Three separate kinds of dead weight, each with its own flag: unused style
definitions, numbering definitions no content or style references, and image
parts nothing in the document part refers to — the last of which the .delete()
methods leave behind as a matter of course.
This is destructive. For styles, the reachability closure in
Styles.usage is the only thing standing between it and a document whose
formatting has quietly changed; keep names styles to preserve along with their
dependencies, for ones you plan to apply but have not yet.
latent_styles is off by default and separate from styles on purpose:
dropping a w:lsdException changes what a user sees in Word's style gallery
rather than how the document renders.
Source code in src/docx/document.py
add_custom_xml_part
¶
add_custom_xml_part(
xml: str | bytes,
schema_refs: Tuple[str, ...] = (),
*,
item_id: str | None = None,
) -> CustomXmlPart
Add an item to the custom XML data store and return its part.
The custom XML data store is where a document-generation pipeline keeps its
data: whole XML documents against a caller-supplied schema, which content
controls in the document bind to through w:dataBinding and Word keeps in step
with what it displays:
document.add_custom_xml_part(
"<invoice><total>42.00</total></invoice>",
schema_refs=("urn:example:invoice",),
)
This is a different thing from custom_properties, which is a flat list
of named scalars in docProps/custom.xml.
A customXml/itemN.xml part is created for xml, along with the
itemPropsN.xml sidecar Word identifies it by, carrying the namespaces named in
schema_refs and a GUID.
item_id is that GUID, in Word's "{XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX}"
shape. One is generated at random when it is omitted, which is what Word does —
but a random value is the one thing in this library's output that is not a
function of its input, so pass an item_id of your own when byte-reproducible
output matters. It only has to be unique within the document.
Source code in src/docx/document.py
remove_vba_project
¶
Remove this document's VBA project; return how many parts were removed.
The word/vbaData.xml sibling, which holds command-bar and macro-name
customisations, goes with it rather than being left orphaned. Zero for a
document that carries no project. Equivalent to del document.vba_project.
Source code in src/docx/document.py
iter_inner_content
¶
replace_text
¶
replace_text(
old: str,
new: str,
*,
count: int = -1,
regex: bool = False,
flags: int = 0,
tables: bool = True,
headers_footers: bool = False,
footnotes: bool = False,
) -> int
Replace occurrences of old with new in this document; return how many.
The match is made against each paragraph's text as a whole, so it succeeds whether or not Word split the text across runs; see Paragraph.replace_text for what happens to formatting.
What gets searched is explicit rather than incidental, because "replace it everywhere" means different things to different callers and getting it wrong is invisible until someone reads the header:
document.replace_text("{{name}}", "Ada") # body only
document.replace_text("{{name}}", "Ada", headers_footers=True) # and those
The document body, including tables unless tables is False, is always
searched. Headers and footers of every section — default, first-page and
even-page alike — are searched when headers_footers is True, and footnotes
and endnotes when footnotes is True. Comments are never searched: a comment
is somebody's remark about the document rather than part of it.
count of -1 replaces every match; any other value limits the total across
everything searched, in the order given above. regex and flags are as for
Paragraph.replace_text.
Source code in src/docx/document.py
save
¶
Save this document to path_or_stream.
path_or_stream can be either a path to a filesystem location (a string or
os.PathLike) or a file-like object.
as_template selects whether the result is a Word template (.dotx /
.dotm) or an ordinary document (.docx / .docm). The default of
None keeps whichever this document already is, so a template opened and saved
is still a template. Pass False to generate a document from a template, or
True to turn a document into one. Macro-enabled input stays macro-enabled
either way.
Note this sets the content type; it does not choose the file extension for you.
Source code in src/docx/document.py
accept_all_revisions
¶
Accept every tracked change in the document body; return how many.
Insertions become ordinary text, deletions go, formatting-change records are dropped leaving the current formatting, and a deleted paragraph mark merges its paragraph with the one after it. The result is the document as Paragraph.text already reads it.
Source code in src/docx/document.py
reject_all_revisions
¶
Reject every tracked change in the document body; return how many.
The reverse of accept_all_revisions: the result is the document as Paragraph.original_text reads it.
Source code in src/docx/document.py
add_text_watermark
¶
add_text_watermark(
text: str,
*,
font: str = "Calibri",
font_size: Length | int | None = None,
color: str = "C0C0C0",
opacity: float | None = None,
angle: float = 315,
width: Length | int = Pt(468),
height: Length | int = Pt(234),
bold: bool = False,
italic: bool = False,
) -> List[Watermark]
Add a text watermark to the whole document, returning the watermarks added.
The faint "DRAFT" or "CONFIDENTIAL" behind the content:
document.add_text_watermark("DRAFT")
document.add_text_watermark("CONFIDENTIAL", color="FF0000", angle=0)
Every section is covered, and within each the default, first-page and even-page headers alike, so the watermark does not disappear on a page that uses a different header. A header shared between sections is written to once.
The arguments are as for Section.add_text_watermark, which is also how a watermark is applied to one section rather than the whole document.
Source code in src/docx/document.py
add_image_watermark
¶
add_image_watermark(
image_path_or_stream: str | PathLike[str] | IO[bytes],
*,
width: Length | int | None = None,
height: Length | int | None = None,
washout: bool = True,
scale: float = 1.0,
) -> List[Watermark]
Add an image watermark to the whole document; see add_text_watermark.
washout applies Word's brightness-and-contrast correction, which is what makes
a logo read as a background rather than sitting opaquely over the text.
Source code in src/docx/document.py
remove_watermark
¶
Remove every watermark from the document, returning how many were removed.
Source code in src/docx/document.py
_Body
¶
_Body(body_elm: CT_Body, parent: ProvidesStoryPart)
Bases: BlockItemContainer
Proxy for <w:body> element in this document.
It's primary role is a container for document content.