Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Convert CSV to XML in Java by parsing each CSV record with a dialect-aware parser, mapping its fields to the XML structure your application requires, and writing the result with StAX or Jackson. For large files or a custom XML contract, Apache Commons CSV with StAX offers explicit, row-at-a-time processing; Jackson is a good fit when the data already maps to Java objects.

Choose the CSV and XML libraries

CSV is not a single rigid format: files vary in delimiters, quoting, escaping, headers, and whitespace rules. Apache Commons CSV supports predefined formats such as RFC 4180 and Excel, plus custom configurations. Its project documentation describes it as a library for reading and writing variations of CSV: Apache Commons CSV.

For XML output, choose based on how much control you need over the document and whether you already have Java models:

Approach Memory approach XML shape control Best fit Main risk
Apache Commons CSV + StAX Can process one row at a time without retaining the full input. Explicit element and attribute writes. Large files and custom XML contracts. More mapping code to write and maintain.
Jackson CSV + Jackson XML Streaming APIs are available, but databinding can materialize Java objects. Defined through beans, annotations, or serializers. Data that already maps cleanly to Java models. Accidental object materialization or output that does not match the required XML contract.

Jackson’s project portal documents CSV and XML modules for both streaming and databinding: Jackson project. StAX is an iterative, event-based XML API, as described in Oracle’s StAX tutorial.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Decide the XML structure before converting

A converter cannot infer the XML vocabulary your application needs. For a flat table, a root element containing repeated row elements and fixed child elements is straightforward to stream. For example, a row with id and name columns could become <record> elements containing <id> and <name> children.

Before coding, specify:

  • The root and row element names, and which CSV columns map to each child or attribute.
  • Whether an empty field becomes an empty element, an omitted element, or an explicit nil value.
  • How missing required columns, extra columns, duplicate headers, and rows with an unexpected number of fields are handled.
  • Whether the target has an XSD or another downstream contract that the output must satisfy.

Use fixed XML names or an explicit mapping for columns. Do not copy arbitrary CSV headers directly into XML element names without validating or sanitizing them.

Convert CSV to XML with Commons CSV and StAX

This pattern reads UTF-8 CSV with its first record as headers, then writes a flat XML document. Adapt the column names, validation, and XML structure to your input and target contract.

try (Reader in = Files.newBufferedReader(csvPath, StandardCharsets.UTF_8);
     Writer out = Files.newBufferedWriter(xmlPath, StandardCharsets.UTF_8)) {
    CSVFormat format = CSVFormat.RFC4180.builder()
            .setHeader()
            .setSkipHeaderRecord(true)
            .build();

    XMLStreamWriter xw = XMLOutputFactory.newFactory()
            .createXMLStreamWriter(out);
    xw.writeStartDocument("UTF-8", "1.0");
    xw.writeStartElement("records");

    for (CSVRecord r : format.parse(in)) {
        xw.writeStartElement("record");
        xw.writeStartElement("id");
        xw.writeCharacters(r.get("id"));
        xw.writeEndElement();
        xw.writeStartElement("name");
        xw.writeCharacters(r.get("name"));
        xw.writeEndElement();
        xw.writeEndElement();
    }

    xw.writeEndElement();
    xw.writeEndDocument();
    xw.close();
}
  1. Open the files with explicit character encodings. UTF-8 is a common choice when it matches the CSV producer. If the input uses another encoding, select that encoding rather than assuming.
  2. Configure the CSV dialect. The example uses RFC 4180 and treats the first record as the header. Choose another predefined format or configure a custom delimiter, quote, escape, null string, whitespace, blank-line behavior, and header handling when the producer’s format requires it. See CSVFormat documentation.
  3. Validate the header and each row. Named access such as r.get("id") is clearer and less sensitive to column order than numeric indexes when headers are stable. Check required columns and row width, and define what happens when validation fails.
  4. Write the agreed XML mapping. StAX’s writeCharacters handles escaping text content; create element and attribute names from trusted, valid mappings.
  5. Close resources and handle failures. Use try-with-resources for the reader and writer, and close the XML writer after completing the document. When reporting malformed rows, include the row number and preserve the original exception cause.

Keep memory use predictable for large files

Process records as an iterable instead of loading the entire CSV into a collection. Commons CSV’s parser exposes record-wise iteration; its API documentation warns that getRecords() can consume significant resources because it loads the remaining input: CSVParser API.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The same principle applies to XML generation: write each row as it is parsed rather than building a complete XML document tree in memory. Jackson also has streaming APIs, but a databinding workflow may hold many mapped objects at once if the application collects them. There is no universal throughput or memory figure for these approaches; measure representative files with the actual dialect, mapping, and runtime environment.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Handle CSV edge cases and verify the output

  • Quoted commas and escaped quotes: Use a parser configured for the producer’s quoting and escape rules instead of splitting each line on commas.
  • Embedded newlines: A quoted CSV field may span lines, so treat parser records—not physical lines—as rows.
  • Alternate delimiters, blank lines, and spaces: Configure the dialect to match the file rather than relying on defaults.
  • BOMs and invalid character sequences: Confirm the input encoding and decide how a byte-order mark or malformed text should be handled.
  • Null and empty values: Define whether these are distinct in the input and how each is represented in XML.
  • Bad headers or row widths: Validate required columns and field counts, then either reject the file or report and skip invalid records according to an explicit policy.

Test with representative files that include these cases, then validate the generated XML against the target XSD or downstream contract when one exists.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.