7 min read
Editing a large XML file in Excel: the “sample first” method
By Quentin Delepierre, Salesforce Commerce Cloud consultant
A product export of several tens of MB is nothing unusual: a Salesforce Commerce catalog, a Google Shopping feed, an Akeneo PIM export. This guide explains how ExcelifyXML handles such files, and a simple workflow that lets you work on very large volumes while transforming only a small sample.
What ExcelifyXML accepts today
- Transformation (XML → CSV): up to 4 MB on the free plan, up to 30 MB on the VIP plan (secure direct upload, nothing retained). Before any credit is used, a free preview shows the detected row, the number of rows and columns; above 4 MB, the file is uploaded only once for the preview and the transformation.
- Reconstruction (CSV → XML): free and credit-less for everyone. It accepts a .excelify.zip bundle up to 4 MB — roughly 30 MB of XML once rebuilt, since the bundle is compressed.
- The rebuilt XML has no size ceiling: it is streamed back to you whatever its length.
The key idea: structure does not depend on volume
When ExcelifyXML transforms an XML file, it produces two files: data.csv (your data) and skeleton.json, the structure manifest — tag order, attributes, namespaces, how each row is rebuilt. That manifest describes the SHAPE of your data, not how much of it there is.
The practical consequence: a 200-product sample with exactly the same structure as your 50,000-product catalog yields exactly the same manifest. So you can transform the small file, then rebuild the big one.
The “sample first” workflow in 4 steps
- Extract a representative sample from your large XML: keep the header, the root tag, and a few dozen or hundred repeated elements (products, entries, URLs…) that cover every field in use. Any text editor will do — or ask your PIM/CMS for it.
- Drop that sample (a few KB to a few MB) on the Upload page. The free preview gives the number of columns: if it looks low, the sample is missing fields. Confirm (1 credit): you get the .excelify.zip bundle, to unzip (data.csv + skeleton.json).
- Open data.csv in Excel, Google Sheets or LibreOffice and replace its contents with the full dataset, keeping the same columns (same headers, same order). You can also generate those rows from a database or another spreadsheet: only the columns matter.
- On the Rebuild page, choose “Separate files” and drop your complete data.csv together with the original skeleton.json — no archive handling: the site compresses and sends both files for you. You get your complete XML, with the exact structure of the original. (If you prefer, you can also re-zip the two files yourself and drop the bundle.)
What you gain
- A single credit consumed, even for a 30 MB catalog: transforming the sample is enough, reconstruction is unlimited and free.
- A lightweight working file in Excel: you edit the volume in the spreadsheet, not in a 30 MB XML that makes your editor crawl.
- A repeatable flow: keep skeleton.json aside and regenerate your XML as often as you like from any CSV export with the same structure.
What a good sample contains
Keep everything around the products (declaration, root, header) and only cut between two complete products. Pick varied products: with and without a sale price, with one or several images, in every language.
<?xml version="1.0" encoding="UTF-8"?>
<catalog xmlns="…" catalog-id="master">
<header>…</header> ← kept as is
<product product-id="A1">…</product>
<product product-id="B7">…</product> ← a few dozen varied products
<product product-id="Z3">…</product>
</catalog> ← closing tag requiredPitfalls to avoid
- The sample must contain ALL the fields that exist in the large file: a field missing from the sample will be missing from the manifest, hence from the rebuilt XML.
- Keep the column headers intact (column names are the link between the CSV and the structure). Adding rows is free; renaming a column breaks the mapping.
- If your spreadsheet is in French, it saves CSV with semicolons: that is expected, ExcelifyXML detects the delimiter automatically, whatever it is.
- Save data.csv as "CSV UTF-8" if your data holds characters Excel's classic "CSV" format can't write (Greek, Cyrillic, Chinese…): it replaces them with "?". The reconstruction notices it and stops rather than produce a damaged XML.
- Paste and import codes as Text: opened by double-click, Excel turns EANs into "3.61235E+12", drops leading zeros and takes some values for dates. The reconstruction stops if a code turned into scientific notation or a date changed format.
Variant: the standalone Excel workbook
The same idea works with the standalone Excel workbook (2 credits): drop your sample on the Excel workbook page, paste all your rows into the workbook's "Data" sheet, below the two header rows and in the same column order (Paste Special › Values), then click "Generate XML".
The difference from online reconstruction: the XML is generated in Excel, on your computer, offline. So there's no size limit on our side, only Excel's limits apply. It needs desktop Excel, on Windows or Mac.
Privacy
Your files are never retained: processed in memory on European Union servers, then deleted immediately. Above 4 MB, the file goes through a private, encrypted transit space for the time of the preview and the transformation, then it's wiped (at the latest one hour after the upload). No metadata is kept.
Transform your sample and rebuild the full volume:
Transform a file