Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a two-stage pipeline: retrieve the webpage HTML, then parse its readable content and write equivalent Word structures with Beautiful Soup and python-docx. This approach gives you control over headings, lists, tables, images, links, cleanup rules, and output streams, but it cannot automatically reproduce every CSS layout or JavaScript-rendered element.

What the conversion pipeline does

python-docx creates and updates Microsoft Word .docx files. Beautiful Soup turns HTML into a navigable tree of Python objects. Keep those responsibilities separate:

  1. Fetch HTML with an HTTP client, applying timeouts, authentication, robots rules and responsible rate limits.
  2. Remove boilerplate and select the article region with Beautiful Soup.
  3. Map HTML semantics to Word headings, paragraphs, list styles, tables, pictures and hyperlinks.
  4. Save a .docx file or write it to an in-memory stream.

The supported target is Office Open XML .docx. Legacy binary .doc files require a separate conversion step.

Install the Python dependencies

python -m pip install requests beautifulsoup4 python-docx lxml

lxml is optional when using Python’s built-in html.parser, but it is useful for more tolerant HTML parsing. Image downloads also need a readable local path or file-like object that python-docx can pass to add_picture.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Minimal HTML-to-DOCX conversion

For a saved HTML file, this is a reliable starting point. It removes common non-content elements, chooses an article when available, and maps headings, paragraphs and list items to Word styles.

from bs4 import BeautifulSoup
from docx import Document

html = open("page.html", encoding="utf-8").read()
soup = BeautifulSoup(html, "html.parser")

for node in soup.select("script, style, template, nav, footer, aside"):
    node.decompose()

article = soup.select_one("article") or soup.body or soup

doc = Document()
for element in article.find_all(["h1", "h2", "h3", "p", "li"]):
    text = element.get_text(" ", strip=True)
    if not text:
        continue
    if element.name == "h1":
        doc.add_heading(text, level=0)
    elif element.name in {"h2", "h3"}:
        doc.add_heading(text, level=int(element.name[1]))
    elif element.name == "li":
        doc.add_paragraph(text, style="List Bullet")
    else:
        doc.add_paragraph(text)

doc.save("webpage.docx")

This deliberately favors readable text over visual fidelity. The article selector is only a starting point: each site needs inspection because publishers use different containers and navigation markup.

Fetch a live webpage safely

Separate retrieval from parsing so network failures do not become parsing bugs. Set a timeout, identify your client, check the status code and retain the response encoding chosen by the server.

from pathlib import Path
import requests
from bs4 import BeautifulSoup
from docx import Document

url = "https://example.com/article"
r = requests.get(
    url,
    headers={"User-Agent": "WebpageToDocx/1.0 (+contact@example.com)"},
    timeout=30,
)
r.raise_for_status()
r.encoding = r.encoding or "utf-8"
soup = BeautifulSoup(r.text, "html.parser")

for node in soup.select("script, style, template, nav, footer, aside"):
    node.decompose()
article = soup.select_one("article") or soup.body or soup

doc = Document()
for element in article.find_all(["h1", "h2", "h3", "p", "li"]):
    text = element.get_text(" ", strip=True)
    if not text:
        continue
    if element.name == "h1":
        doc.add_heading(text, level=0)
    elif element.name in ("h2", "h3"):
        doc.add_heading(text, level=int(element.name[1]))
    elif element.name == "li":
        style = "List Number" if element.parent and element.parent.name == "ol" else "List Bullet"
        doc.add_paragraph(text, style=style)
    else:
        doc.add_paragraph(text)

doc.save("article.docx")

Do not bypass access controls or ignore a site’s terms. Authenticated pages may require a session, cookies or an authorization header; rate-limit repeated requests and handle 429 responses with backoff.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Preserve headings, lists, links, images and tables

Headings and paragraphs

Use Word heading styles rather than manually inserting bold text. Word can then build a navigation pane and table of contents. Add one paragraph per readable HTML block and collapse incidental whitespace without joining separate sentences.

Ordered and unordered lists

Use List Bullet and List Number. Nested lists need recursive handling if their hierarchy matters; a simple find_all loop can flatten them.

Tables

Create a Word table for each HTML table, copy rows and cells, and decide how to treat colspan and rowspan. A basic rectangular table conversion is:

from docx import Document

def add_html_table(doc, table_node):
    rows = table_node.find_all("tr")
    matrix = []
    for row in rows:
        matrix.append([cell.get_text(" ", strip=True)
                       for cell in row.find_all(["th", "td"], recursive=False)])
    width = max((len(row) for row in matrix), default=0)
    if not width:
        return
    table = doc.add_table(rows=len(matrix), cols=width)
    table.style = "Table Grid"
    for r, row in enumerate(matrix):
        for c, value in enumerate(row):
            table.cell(r, c).text = value

For production data, add explicit handling for merged cells, header rows and cells containing links or images.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Images

Resolve relative URLs against the page URL, download only permitted resources, check content type and size, then pass a local filename or file-like object to doc.add_picture. Set a width such as Inches(6) to prevent oversized images. A failed image should be logged and skipped rather than aborting the whole document.

Links

get_text() preserves visible link text but not clickability. To create clickable hyperlinks, add an external hyperlink relationship to the document part and insert a run containing the link text. If clickable links are not required, include the URL in parentheses after the visible text so the destination remains recoverable.

A fuller converter with content selection

The following program handles paragraphs, headings, lists, tables and local image downloads while retaining a clean separation between fetching and document generation.

from io import BytesIO
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
from docx import Document
from docx.shared import Inches

URL = "https://example.com/article"
HEADERS = {"User-Agent": "WebpageToDocx/1.0"}

response = requests.get(URL, headers=HEADERS, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
for node in soup.select("script, style, template, nav, footer, aside, form"):
    node.decompose()
root = soup.select_one("article") or soup.body or soup

doc = Document()

def add_image(src):
    image_url = urljoin(URL, src)
    try:
        image = requests.get(image_url, headers=HEADERS, timeout=20)
        image.raise_for_status()
        if image.headers.get("Content-Type", "").startswith("image/"):
            doc.add_picture(BytesIO(image.content), width=Inches(6))
    except requests.RequestException:
        pass

for node in root.find_all(["h1", "h2", "h3", "p", "li", "table", "img"]):
    if node.name == "img" and node.get("src"):
        add_image(node["src"])
        continue
    if node.name == "table":
        rows = node.find_all("tr")
        values = [[c.get_text(" ", strip=True)
                   for c in row.find_all(["th", "td"], recursive=False)]
                  for row in rows]
        width = max((len(row) for row in values), default=0)
        if width:
            table = doc.add_table(rows=len(values), cols=width)
            table.style = "Table Grid"
            for r, row in enumerate(values):
                for c, value in enumerate(row):
                    table.cell(r, c).text = value
        continue
    text = node.get_text(" ", strip=True)
    if not text:
        continue
    if node.name == "h1":
        doc.add_heading(text, level=0)
    elif node.name in {"h2", "h3"}:
        doc.add_heading(text, level=int(node.name[1]))
    elif node.name == "li":
        parent = node.find_parent(["ol", "ul"])
        style = "List Number" if parent and parent.name == "ol" else "List Bullet"
        doc.add_paragraph(text, style=style)
    else:
        doc.add_paragraph(text)

doc.save("article.docx")

This example can duplicate text inside complex tables or nested lists because it walks all descendants. A production converter should process direct children recursively, mark handled subtrees, and implement the site’s actual content selectors.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In-memory output for a web service

python-docx accepts file-like inputs and outputs. Save to BytesIO and return the bytes with the DOCX content type instead of writing a temporary file.

from io import BytesIO

def document_bytes(doc):
    stream = BytesIO()
    doc.save(stream)
    stream.seek(0)
    return stream.getvalue()

# Flask example:
# return Response(document_bytes(doc),
#                 mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document",
#                 headers={"Content-Disposition": "attachment; filename=article.docx"})

JavaScript-rendered pages and fidelity limits

An HTTP request receives the server’s HTML, not necessarily the DOM produced after JavaScript runs. If the article appears only after client-side rendering, use an approved browser automation layer to render it first, then pass the resulting HTML to the parser. A full browser or document-conversion engine can preserve CSS layout more faithfully, but adds browser binaries, startup time, sandboxing and operational complexity. Beautiful Soup plus python-docx offers finer Python control and lower moving parts, not guaranteed visual parity with Word or every webpage.

Or skip the browser setup

If you need a clean capture of a rendered page before further processing, ScreenshotNeo provides a website screenshot API and MCP server. A single request can return PNG, JPEG, WebP or PDF. For example:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for request options. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP tools—take_screenshot, get_page_info and capture_pdf—let Claude, Cursor and other MCP clients request captures.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Free plan includes 1,000 screenshots per month without a card. Paid plans start at $5 for 3,000 shots; every plan includes every feature. Sign up free for ScreenshotNeo.

Performance, reliability and cost considerations

  • Reuse an HTTP session for multiple pages so connections can be reused.
  • Set separate, finite timeouts for HTML and image downloads; never allow an unbounded request to hold a worker.
  • Cache fetched HTML and images when permitted, but honor freshness, authentication and robots requirements.
  • Limit image dimensions and total bytes to avoid memory spikes in batch jobs.
  • Process independent URLs with bounded concurrency, then retry transient network failures with exponential backoff.
  • Write intermediate logs containing URL, selector chosen, item counts and skipped resources, not page secrets.
  • Validate that the output is a readable ZIP-based .docx before delivering it.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting

The document is empty

The page may have no article element, may require JavaScript, or may return an access-denied page. Save and inspect response.text, choose the site’s real content selector, or render the page with a browser before parsing.

Navigation and cookie text appears

Expand the removal selector list with the site’s classes or IDs. Generic selectors cannot identify every sidebar, consent dialog or newsletter component.

Accented characters are corrupted

Use response.apparent_encoding only when the server declaration is wrong, and pass the correct encoding to Beautiful Soup. For local files, always open with the page’s actual encoding, commonly UTF-8.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Lists lose nesting or numbering

Do not flatten every li in one pass. Walk direct children recursively and assign indentation or a matching Word list level.

Images fail

Relative URLs, hotlink protection, authentication, unsupported formats and oversized files are common causes. Resolve with urljoin, send appropriate headers, enforce a size limit, and log skipped images.

Tables are malformed

Handle colspan and rowspan explicitly, and ignore layout tables when they are not content. A rectangular table loop cannot represent merged cells correctly.

Word will not open the file

Ensure the stream is rewound before returning it, do not append text after saving, and use a .docx extension with the Office Open XML content type. Legacy .doc is not an output format supported by this API.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The page returns 403 or 429

Respect the site’s access policy, identify your client, slow requests, authenticate only when authorized, and retry 429 responses after the server’s specified delay. Do not attempt to defeat bot protections.

FAQ

Can I convert an existing HTML string instead of a URL?

Yes. Pass the string directly to BeautifulSoup; only the retrieval stage changes.

Can the script update an existing Word template?

Yes. Construct Document("template.docx"), then add or replace content before saving a new file.

Will the output preserve the webpage’s CSS exactly?

No. The workflow maps semantic content to Word structures; exact browser layout requires a rendering or document-conversion engine.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I convert an existing HTML string instead of a URL?

Yes. Pass the string directly to BeautifulSoup and use the same parsing and DOCX generation stages.

Can the script update an existing Word template?

Yes. Open a template with Document(“template.docx”), add the converted content, and save a new .docx file.

Will the output preserve the webpage’s CSS exactly?

No. Semantic mapping preserves content structure, while exact visual rendering requires a browser or specialized conversion engine.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.