It would be great to be able to modify the item boost on a per-item basis. I.e., we should offer a processor that lets you pick a numeric field for each datasource/bundle and at indexing time multiplies the item boost with that field’s value. This way, you could easily configure individual items as being especially important – or, conversely, as not being very relevant as a search result.

If you would be interested in such a feature, please speak up here and, ideally, provide a patch. (A copy of \Drupal\search_api\Plugin\search_api\processor\TypeBoost would be a good place to start.)

Comments

mpp created an issue. See original summary.

mkalkbrenner’s picture

In general I'm not a friend of elevate.xml and prefer LTR (learn to rank).

In your case, why not just adding a keywords field to the entities and give them a high boost?

mpp’s picture

I like the idea of adding a keywords field with a high boost. Only issue here is that there's weight to distinguish between the keywords.

mkalkbrenner’s picture

Only issue here is that there's weight to distinguish between the keywords.

using LTR it's possible to weight terms individually.

Another approach are payloads as we already use for boosting entity types and HTML tags.

If you use (e)dismax you can also use "bq" and "bf".

And for sure you can use "eleavate".

From a module maintainers perspective, LTR is the way to go and the only approach worth implementing a UI for.
(Elevate is simply too limited and not worth the effort.)

mpp’s picture

Status: Active » Closed (won't fix)
mpp’s picture

@mkalkbrenner, let me know if you want me to open another feature request to implementing a UI for LTR to boost individual keywords.

mkalkbrenner’s picture

Title: Provide a UI to export keywords to elevate.xml » Custom boost factors
Status: Closed (won't fix) » Active

From a discussion on Slack (I hate it):

To recap: Someone wants to keep the entire boost calculation that exists already and wants to add another criteria based on a numeric field added to the entity.

I see two reliable ways to achieve such a feature

1. LTR
you can create a simple "model" that leverages the calculated relevance as Solr would normally so according to our Search API configuration as one "feature". Then you add a second "feature" based on your field's value.
LTR means learn to rank

2. payloads
Have a look at flattenKeysToPayloadScore().
You can add a numeric value as payload to your custom field when indexing. If I assume that your custom field is an integer or float you can index the same value again in it's payload. These payload values could be "added" to the relevancy calculation.
That's how I reimplemented the HTML tag based boosts.

mkalkbrenner’s picture

Component: User interface » Code

I think that it's not too hard to create a search_api_solr sub-module to offer such an entity field.
Or a Search API processor that acts on a numeric field of choice.

mkalkbrenner’s picture

Title: Custom boost factors » Use an entity field value as custom boost factor
Project: Search API Solr » Search API
Version: 8.x-3.x-dev » 8.x-1.x-dev
Component: Code » Plugins

Meanwhile there's another contrib module for LTR.

If someone is interested we should implement this simple approach:

Extend the TypeBoost Processor or derive from it. Offer a configuration form to select a numeric field for each datasource / bundle. This field's value has to be set as or add to the calculation of $item->setBoost(...) in preprocessIndexItems().

drunken monkey’s picture

If a few people would be interested in such a feature, sure, we could add such a processor. Would seem to be rather a niche functionality to me, though.
Also, in any case, I unfortunately don’t have the capacity to work on such feature requests in my spare time, so would wait for someone to provide a patch first. (Would provide details if someone is interested.)

dww’s picture

Would this be the right issue for adding a processor to search_api to boost results based on a date field?

https://www.drupal.org/project/issues/search_api?text=date+boost is turning up empty. ;)

I'm trying to improve the relevancy of search results on a site I'm building that involves events. The site owners would like upcoming events to be much more relevant than expired events.

I can't find any existing solution that lets me define a boost factor based on a single (date) field. I can boost all events relative to other node types, but I can't boost upcoming events and unboost (or whatever you call it when you reduce relevancy) to past events.

https://www.drupal.org/docs/8/modules/search-api-solr/search-api-solr-ho... is all about Solr, and sadly, I can't use Solr for this site due to hosting limitations.

https://www.drupal.org/project/search_api_boost_priority doesn't provide a processor like this.

Seems like a not-so-uncommon scenario. Folks also seem to love to be able to boost based on node creation date.

Is this the right issue to discuss/develop said feature, or should I open a new one?

Thanks!
-Derek

laura.gates’s picture

I would love a feature like this. Google Search Appliance used to have something like what was described in the description. With GSA hitting end of life a few years ago, UI solutions for things like this have been harder to find in search products. Modifying things like Best Bets in the UI would give clients a bit more freedom in their search results.

drunken monkey’s picture

Issue summary: View changes

Would this be the right issue for adding a processor to search_api to boost results based on a date field?

No, I think that’s a different feature, as it’s more backend-specific. (I.e., I don’t think this is currently even possible on the framework level with pure Search API.) This issue is about adding just a flat numeric boost on a per-item basis.
I see now, though, that the issue title and description are completely out of sync at this point, so I’d better bring them in line again …
For your suggestion, if there is really no issue for it yet, please open a new one. (Do check the Search API Solr issue queue first, though – as said, that feature would probably have to be backend-specific.)

@ Laura: With “a feature like this” do you mean the one from the (former) issue description or the one discussed in #9, and in the new issue description? I.e., what exactly do you want to be able to do?

laura.gates’s picture

@drunken monkey, I was referring to being able to edit the elevate.xml file in a Search API UI solution that is in the ticket description. I don't have easy access to the elevate.xml file settings with one of my projects.

drunken monkey’s picture

@drunken monkey, I was referring to being able to edit the elevate.xml file in a Search API UI solution that is in the ticket description. I don't have easy access to the elevate.xml file settings with one of my projects.

As that is Solr-specific, you should create an issue in the Search API Solr issue queue for that (if there isn’t one already).

mkalkbrenner’s picture

Assigned: Unassigned » mkalkbrenner
mkalkbrenner’s picture

StatusFileSize
new13.25 KB

here's a first draft.

mkalkbrenner’s picture

Status: Active » Needs review

Status: Needs review » Needs work

The last submitted patch, 17: 3055173.patch, failed testing. View results

mkalkbrenner’s picture

Status: Needs work » Needs review
StatusFileSize
new14.62 KB
new8.17 KB

Status: Needs review » Needs work

The last submitted patch, 20: 3055173_20.patch, failed testing. View results
- codesniffer_fixes.patch Interdiff of automated coding standards fixes only.

drunken monkey’s picture

Status: Needs work » Needs review
StatusFileSize
new15.78 KB
new15.11 KB

Thanks a lot for the patch, great work!

I’m not sure that much configuration is really needed, though. I would have just used a select box for a single field to use (always using 1.0 and max for boost/aggregation), or maybe checkboxes if we really want to allow multiple fields. It seems to me, though, that the use case would be a separate field, added specifically for this purpose, so there would always be just a single field with a single value. Or do you have a use case for actually re-using some other field(s), where this wouldn’t apply?
Anyways, if you do see a use case, I’m not that opposed to leaving the configuration in, but I usually aim to not add unnecessary configs, as the configuration for this module is far too complicated already anyways. (Partly because I haven’t always done a great job of this in the past.)

Speaking of this use case: I think in most cases, people will not actually want to index that boost field, just use its value for the item boost. So it would seem better to leave the user to choose any of the available (first-level) properties, not just the indexed fields.
However, I do see that that is a bit harder to implement – and to fit into the UI. So maybe just doing it like this is still the better solution – the overhead in the index should be pretty small anyways. (Or maybe we could even add an option that will remove the field from indexing after using its value? Would that make sense and be worthwhile, or just too complicated or too confusing?)

Furthermore, as discussed on Slack, item-level boosts for the DB backend unfortunately don’t work without keywords present. So you’ll either need to add a common keyword (like “node”) to all searches, or change the test to only check for correct computation of the boost values (like TypeBoostTest does).

Actually, I think the former makes the most sense, especially if we just move the query creation to a helper method. That way, you could also easily override that for the Solr tests. Implemented in the attached patch revision, along with some other changes (mostly nit-picks).

Some further remarks (plus one question):

  1. +++ b/config/schema/search_api.processor.schema.yml
    @@ -195,6 +195,25 @@ plugin.plugin_configuration.search_api_processor.type_boost:
    +        nullable: true
    

    Haven’t seen this anywhere yet, but good to know it exists!
    However, why is it needed here? Is it in case the list is empty? (Don’t think that made problems for us anywhere else before.)

  2. +++ b/src/Plugin/search_api/processor/NumberFieldBoost.php
    @@ -0,0 +1,151 @@
    +  protected static $boost_factors = [
    +    '0.0' => '0.0',
    

    I don’t think this needs to be a property. If we make this a variable in the form builder method, we can also just use a label for the “Don’t use” value, instead of a magic “0.0” value. (I used “Ignore” for the label, but open to input.)

  3. +++ b/tests/src/Kernel/Processor/NumberFieldBoostTest.php
    @@ -0,0 +1,257 @@
    +  public function setUp($processor = NULL): void {
    

    Can’t use that yet, as we still want to stay compatible to Drupal 8 and, thus, PHP 7.0 for a few more months.

  4. +++ b/tests/src/Kernel/Processor/NumberFieldBoostTest.php
    @@ -0,0 +1,257 @@
    +    // Add an integer field.
    +    $boostFieldStorage = FieldStorageConfig::create([
    +      'field_name' => 'boost',
    +      'entity_type' => 'node',
    +      'type' => 'integer',
    +      'cardinality' => 3,
    +    ]);
    +    $boostFieldStorage->save();
    

    Are you sure this worked for you? You’re using 'boost' as the field ID here, but 'field_boost' everywhere else. Had to change that to make the tests pass for me.

mkalkbrenner’s picture

Assigned: mkalkbrenner » Unassigned
Status: Needs review » Reviewed & tested by the community

It seems to me, though, that the use case would be a separate field, added specifically for this purpose, so there would always be just a single field with a single value. Or do you have a use case for actually re-using some other field(s), where this wouldn’t apply?

Scoring is always a difficult task with Solr. It might be that you perfectly adjusted your scoring but later you need to add a different scoring property like location or recent date and notice that the dimensions / relations aren't working. In this case it's easier to adjust the boost "factor" here instead of changing the field value of thousands of entities. That's why I introduced it.

I would have just used a select box for a single field to use (always using 1.0 and max for boost/aggregation)

The "first value" is always used for sorting. So this option should be available for boosting, too. But I agree that max is the most common option.

Speaking of this use case: I think in most cases, people will not actually want to index that boost field, just use its value for the item boost. So it would seem better to leave the user to choose any of the available (first-level) properties, not just the indexed fields.

I would suggest to keep the field in the index. But if you want to offer the optionally remove it, you can do it. I saw both use-cases.
In some cases the requirement was to boost on existing fields that are needed and present in the index anyway, for example stock of a product. Not the absolute value but an integer indicating "no stock", "low stock", "full stock".
In another case there's a special dedicated field for boosting which could be removed.

I general see your concerns and agree to keep things simple. But I already went through multiple iterations with this patch and already considered my proposal as the minimum set of features. So I agree with your patch in #22.

Thanks for your help!

drunken monkey’s picture

Status: Reviewed & tested by the community » Fixed

OK, all makes sense. Thanks for your detailed reply!
Committed. Thanks again!

Status: Fixed » Closed (fixed)

Automatically closed - issue fixed for 2 weeks with no activity.