Quote from: ZAP at Jan 22, 2007, 07:16 PM
I’ve used external search engines/indexers in the past, and generally they were a pain to reindex all the time (usually I set a crontab, but that just seemed clunky to me). So if MODx is going to maintain an index of page output in the future, then I would suggest that there be a simple trigger that can be called whenever page output is updated (by Jot, for example) and that it be capable of reindexing only the new content. This also raises the question of whether this data would have any relation to the cache files or not (since this would include much of the same info, but with tags and code stripped out).
This would be a custom search index tightly integrated with the core, which could be configurable to allow users to maintain search indexes of any kind of data, from any source, into custom indexes for specific purposes. Check out Lucene or Zend_Search_Lucene for examples of this kind of flexible search infrastructure.
Quote from: ZAP at Jan 22, 2007, 07:16 PM
Personally I use TVs for non-dynamic content regularly. For example, sidebar content or other info specific to that particular document. I would think that a standard routine (a la Ditto) to include those TVs when searching would be useful for many people. And I actually add randomized content to my pages a lot, and I wouldn’t want that content indexed for searching. So the future system that parses actual document output would need to also include something like the ability to set flags in comments for what not to search.
What you include in the indexes would be completely configurable. You could maintain indexes for private access, one for public access, and another with indexed PDF content linking to the PDF files it was indexed from. There would be a standard API for searching the index, creating and building indexes, and keeping them up to date.
Quote from: ZAP at Jan 22, 2007, 07:16 PM
In the meantime, sure I can hack the queries and add join statements to them, but not everyone is going to feel comfortable messing that deeply with the code.
Didn’t imply they were, just spelling out the technical challenges that must be overcome to provide a robust and generic solution to this problem of indexing/searching content that can be dynamic and/or static based on all kinds of variables, and would otherwise never have a chance to be searchable.