We will soon release expanded support for additional file types for when using Custom Extract Agents in Box Extract, enabling enterprises to extract structured data from a broader range of file types and more of your organization's content.
This will be a follow-up to additional file type support in Box Extract that we launched in February 2026, which introduced support for a defined set of file types. This release expands file type support by including all file types available via the Box Extract API introducing parity with what we already provide programmatically.
Box Extract will support extraction from the following files when using Custom Extract Agents:
- Images: bmp, cr2, crw, dcm, dicm, dicom, dng, gif, heic, jpeg, jpg, nef, png, raf, raw, svg, tga, tif, tiff, webp
- Design/CAD: ai, dwg, eps, idml, indd, indt, inx, ps, psd, xbd, xdw
- Documents: boxcanvas, boxnote, doc, docx, gdoc, msg, odt, pages, pdf, rtf, webdoc, wpd
- Presentations: gslide, gslides, key, odp, ppt, pptx
- Spreadsheets: csv, gsheet, numbers, ods, tsv, xls, xlsb, xlsm, xlsx
- Code/text: as, as3, asm, bat, c, cc, cmake, cpp, cs, css, cxx, diff, erb, groovy, h, haml, hh, htm, html, java, js, json, less, log, m, make, md, ml, mm, php, pl, plist, properties, py, rb, rst, sass, scala, scm, script, sh, sml, sql, txt, vi, vim, xhtml, xml, xsd, xsl, yaml
When a Custom Extract Agent is applied to a folder in Box, all the file types listed above will automatically have structured data extracted from them and applied as metadata alongside those files.
Custom Extract Agents applied to existing folders will continue extracting from the previously supported file types only. To enable support for the additional file types listed above, simply remove and re-assign the Custom Extract Agent to the folder.
Stay tuned to learn more about this release.