A developer chooses to use text pre-extraction to reindex Lucene indexes.
When is this recommended?
Text pre-extraction is recommended when an existing Lucene index with binary extraction enabled is being reindexed. It moves costly full-text extraction from binary content into a separate process, reducing the load of the reindexing operation. Image-heavy repositories do not receive the same benefit because images generally contain no extractable text. Adobe Experience Manager: Best Practices for Queries and Indexing
Community Discussion