Extract Text from XML

Turn any XML document into clean, readable plain text instantly.

XML to Text Converter

Paste or upload your XML and clickExtract Text to get plain text.
Live mode is off by default; tick it if you prefer instant results.


Load an example

Settings

Readable name: value pairs, the classic tag-and-data text, or values only.

Placed between each name and its value in Tag + data Pair mode.

Joins the records (the root element and each direct child). Paragraph is the default.

Strip leading and trailing spaces from text and attributes. Turn it off to keep the original whitespace.

Paste some XML and click Extract Text to get started.


Overview

XML (Extensible Markup Language) is a widely used format for storing and transporting structured data. It wraps content inside tags like <title> and <author>, often with attributes like id="1". While machines read XML easily, the tags, brackets and attributes make it hard for humans to skim the actual content.

Our Extract Text from XML tool removes all XML markup and leaves only the meaningful text: tag names that show structure, the text content between tags, and any attributes that describe elements. It is built entirely with the browser's native DOMParser, so your XML never leaves your device and no server is involved.


What This Tool Does

Strips XML Syntax: The tool parses your XML with the browser's native parser and walks every element in document order to build a flat text string. Tag names, attributes and text content all contribute to the result, while namespace declarations such as xmlns and xmlns:soap are ignored so they never clutter the output.

Example: <catalog><book isbn="978-1-2345-6789-0"><title>The Glass Lighthouse</title></book></catalog> produces two records — catalog and book isbn: 978-1-2345-6789-0; title: The Glass Lighthouse in the default Tag + data Pair mode.

Line Breaks Between Siblings: The root element is written on its own record, and each direct child element of the root starts a new record, so multiple top-level items are easy to scan. Deeper nested elements remain on the same record as their parent to keep context intact. Records are separated by a blank line (Paragraph) by default.

Choose Your Output Mode: The Output mode menu offers three layouts. Tag + data Pair (the default) renders readable name: value pairs joined with semicolons. Tag + data returns the classic flat form: tag names and content separated by spaces, with attributes placed right after the tag name. Data only keeps every value (text content and attribute values) but drops all tag and attribute names.

Example: for <catalog><book isbn="978-1-2345-6789-0"><title>The Glass Lighthouse</title></book></catalog> the book record is book isbn: 978-1-2345-6789-0; title: The Glass Lighthouse in Tag + data Pair, book isbn 978-1-2345-6789-0 title The Glass Lighthouse in Tag + data, and 978-1-2345-6789-0 The Glass Lighthouse in Data only.

Two Separators, Clear Roles: The Tag separator dropdown belongs to Tag + data Pair mode and controls the text placed between each name and its value; it is Colon by default, so book isbn: 978-1-2345-6789-0; title: The Glass Lighthouse reads cleanly. Choose Space for isbn 978-1-2345-6789-0, or type any character with Custom.... The separate Record separator joins the records themselves (the root plus one record per direct child of the root) in every mode; Paragraph (a blank line) is the default, while New line gives scannable rows, Comma gives CSV-style lists, Space flattens everything into a single line, Tab suits spreadsheet imports and Semicolon gives semicolon-delimited output.

Handles Namespaces Cleanly: Documents that declare namespaces (SOAP envelopes, RSS feeds, sitemaps) stay readable. The xmlns and xmlns:prefix declarations are skipped, and tag and attribute names use their local name, so <soap:Envelope> contributes envelope instead of a namespace URL.

Trim Whitespace: The Trim whitespace toggle is on by default. It strips leading and trailing spaces, tabs and newlines from text content and attribute values, and skips indentation-only nodes, so <a>  hi  </a> becomes a: hi in Tag + data Pair mode. Turn it off to keep the original whitespace exactly as written — the same element then becomes a:   hi  .

Works Offline in Your Browser: All processing happens client-side using JavaScript and the browser's built-in XML parser. Nothing you paste or upload is stored or transmitted anywhere, which makes it safe for sensitive documents.

Clear Error Handling: If the XML is malformed — missing closing tags, mismatched quotes, or illegal characters — the tool reports the parser's error message and clears the output so you can fix the input and try again.

Live Mode (Off by Default): Live Mode is off by default so the page does not reprocess large XML on every keystroke. Tick Live Mode (Update as you type) if you want the output to refresh instantly as you type, upload a file, or change an option. Conversions still run entirely in your browser, so live updates stay private to your device.


How It Works

1. Paste or Upload XML: Type or paste your XML into the Input box, click Load an example to see a pre-filled sample, or upload a .xml / .txt file from your device.

2. Choose Your Output Format: Pick the Output mode: Tag + data Pair for readable name: value records (the default), Tag + data for the classic tag-and-data text, or Data only for values with no tag or attribute names. In Tag + data Pair mode the Tag separator (Colon by default) sits between each name and its value. The Record separator (Paragraph by default) decides how records are joined in every mode — blank lines, new lines, commas, spaces, tabs or semicolons. Toggle Trim whitespace to remove indentation leakage from text content and attributes.

3. Click Extract Text (or use Live Mode): Press the Extract Text button, or tick Live Mode to have the output refresh automatically as you type, upload a file, or change an option. The browser parses the XML, walks every element and text node in document order, and builds the flat text output using your chosen settings.

4. Read the Result: The result appears in the Output box, with the root element and each direct child separated by your chosen record delimiter. Tag + data Pair shows readable name: value pairs, Tag + data shows the classic space-separated form with tag names, content and attributes, and Data only shows just the values.

5. Save or Copy: Use the Copy icon to put the result on your clipboard, or Download to save it as a .txt file. Press Clear All to start over.


Examples

Click Try it on any example below to load it straight into the tool. The outputs shown use the default Tag + data Pair mode.

Example 1: Library book catalog
Input (XML)
<library name="Central City Library" established="1998">
  <book isbn="978-1-2345-6789-0" category="Fiction">
    <title>The Glass Lighthouse</title>
    <author>Elena Voss</author>
    <year>2023</year>
    <available>true</available>
  </book>
  <book isbn="978-0-9876-5432-1" category="Science">
    <title>Rocks from Mars</title>
    <author>Dr. Amir Patel</author>
    <year>2020</year>
    <available>false</available>
  </book>
</library>
Output (plain text)
library name: Central City Library; established: 1998

book isbn: 978-1-2345-6789-0; category: Fiction; title: The Glass Lighthouse; author: Elena Voss; year: 2023; available: true

book isbn: 978-0-9876-5432-1; category: Science; title: Rocks from Mars; author: Dr. Amir Patel; year: 2020; available: false

The root <library> gets its own record with its name and established attributes. Each <book> child becomes its own record, with attributes and child elements written as name: value pairs joined by semicolons. A blank line separates the records (the default Paragraph separator).

Example 2: Conference talk schedule
Input (XML)
<schedule date="2025-03-14">
  <session room="Hall A">
    <time>09:00</time>
    <speaker>Maya Lin</speaker>
    <topic>Designing APIs People Love</topic>
  </session>
  <session room="Hall B">
    <time>11:30</time>
    <speaker>Omar Faruk</speaker>
    <topic>Debugging Distributed Systems</topic>
  </session>
</schedule>
Output (plain text)
schedule date: 2025-03-14

session room: Hall A; time: 09:00; speaker: Maya Lin; topic: Designing APIs People Love

session room: Hall B; time: 11:30; speaker: Omar Faruk; topic: Debugging Distributed Systems

Nested elements like <time> and <speaker> stay on the same record as their parent <session> and become name: value pairs. The root <schedule> attribute (date) is written the same way.

Example 3: RSS news feed
Input (XML)
<rss version="2.0">
  <channel>
    <title>Weekly Tech Digest</title>
    <link>https://example.com/tech</link>
    <description>Curated stories for developers.</description>
    <item>
      <title>Understanding WebAssembly</title>
      <author>Liam Chen</author>
      <pubDate>2025-04-01</pubDate>
    </item>
    <item>
      <title>10 CSS Tricks You Missed</title>
      <author>Sara Gomez</author>
      <pubDate>2025-04-08</pubDate>
    </item>
  </channel>
</rss>
Output (plain text)
rss version: 2.0

channel title: Weekly Tech Digest; link: https://example.com/tech; description: Curated stories for developers.; item title: Understanding WebAssembly; author: Liam Chen; pubdate: 2025-04-01; item title: 10 CSS Tricks You Missed; author: Sara Gomez; pubdate: 2025-04-08

Because both <item> elements are nested inside <channel>, they remain on the same record as their parent rather than splitting onto separate records. Only direct children of the root start new records.

Example 4: Contact list
Input (XML)
<contacts>
  <contact id="101" type="personal">
    <name>Nadia Okafor</name>
    <email>nadia@example.com</email>
    <phone>+1-555-0142</phone>
    <city>Lagos</city>
  </contact>
  <contact id="102" type="work">
    <name>Tomás Rivera</name>
    <email>tomas@example.com</email>
    <phone>+44-20-7946-0958</phone>
    <city>London</city>
  </contact>
</contacts>
Output (plain text)
contacts

contact id: 101; type: personal; name: Nadia Okafor; email: nadia@example.com; phone: +1-555-0142; city: Lagos

contact id: 102; type: work; name: Tomás Rivera; email: tomas@example.com; phone: +44-20-7946-0958; city: London

Multiple attributes per element (id and type) are written as name: value pairs, and each child element follows the same pattern. Each contact stays on its own record, making it easy to scan or paste into a spreadsheet.


When to Use This Tool

You can reach for this tool whenever you have XML data and want the human-readable content without the markup noise:

Skimming API Responses: Many web services return XML. Converting it to plain text lets you spot the data you care about at a glance, without opening a heavy XML viewer.

Feeding Other Text Tools: Once the XML is flattened to plain text, you can drop it straight into a word counter, sorter, or find-and-replace tool on this site.

Reviewing Configuration or Log Files: Application configs and logs are sometimes stored as XML. Extracting the text makes them searchable with any plain-text tool.

Quick Content Audits: Spot-check a batch of XML files by extracting the text and scanning the tag names and values for anomalies, missing fields or duplicates.


FAQs

FAQ 1: Is my XML sent to a server?
Answer: No. The XML you paste or upload is parsed and converted entirely in your browser with JavaScript. Nothing leaves your device.

FAQ 2: What is the difference between the three output modes?
Answer: Tag + data Pair (the default) renders readable name: value records. Tag + data returns the classic space-separated form with tag names, content and attributes. Data only drops all tag and attribute names and returns just the values (text content and attribute values).

FAQ 3: How are XML attributes shown in the output?
Answer: In Tag + data Pair mode, attributes are written as name: value pairs, so <book isbn="123"> contributes book isbn: 123. In Tag + data mode the attribute name and value sit right after the tag, giving book isbn 123. In Data only mode only the value remains: 123.

FAQ 4: How do I change what separates a name from its value?
Answer: Use the Tag separator dropdown in Tag + data Pair mode. It controls the text placed between each name and its value and is Colon by default. Choose Space, Comma, Semicolon, Tab or New line, or pick Custom... and type any separator you like. The control is hidden in Tag + data and Data only modes.

FAQ 5: How are records separated and how do I change the separator?
Answer: The root element is written first as its own record, then each direct child of the root becomes its own record, while deeper descendants stay on their parent's record. The Record separator joins these records: Paragraph (the default) leaves a blank line, New line gives one record per line, Comma gives CSV-style output, Tab suits spreadsheets, Space flattens everything into a single line, and Semicolon gives semicolon-delimited output.

FAQ 6: What does the Trim whitespace option do?
Answer: When on (the default), it removes leading and trailing spaces, tabs and newlines from every text node and attribute value, and skips indentation-only nodes. Turn it off to keep the original whitespace exactly as written, including the indentation and blank lines between elements.

FAQ 7: How are XML namespaces handled?
Answer: Namespace declarations (xmlns and xmlns:prefix) are ignored so they do not appear as data, and tag and attribute names are reduced to their local name. A SOAP document such as <soap:Envelope><soap:Body>...</soap:Body></soap:Envelope> therefore contributes envelope and body rather than namespace URLs.


Tools Worth Trying

TextToolz

The Ultimate Text Tools

TextToolz works seamlessly to let you convert and design your text. It is fast, reliable and secure. Trusted by thousands of users.