Hi,
Based on the following:
http://wiki.apache.org/solr/ExtractingRequestHandler
http://wiki.apache.org/solr/ContentStream
It looks like it's possible to ask Solr Cell to fetch a file from a URL, parse it using Tika and add it to the index, I think this would offload a lot of processing overhead to Solr server, and allow content outside Drupal to be indexed without adding it to Drupal.
I'd like to implement this, but I'm not sure what is the best way to do it. I believe I'll need to override SearchApiSolrService.indexItems and getFields at least, this seems to mean I need to subclass SearchApiSolrService. However, if I do subclass it, there's no way I can use this subclass within the search_api_solr module, so I'll have to create a new module for it and replicate all the stuff inside search_api_solr, which is a waste. As far as I can see, the easiest way to do this would be a patch for search_api_solr itself, any thoughts?
Thanks
Comments
Comment #0.0
jrao commentedfix url
Comment #1
OanaIlea commentedThis issue was closed due to lack of activity over a long period of time. If the issue is still acute for you, feel free to reopen it and describe the current state.