Problem/Motivation

I need to boost content by its created and last changed date. Created is more important than last changed.
Sadly, I could neither find an option for this, nor documentation.

I looked into

I found

Processors: Boost more recent dates: A processor to boost more recent dates.

at https://www.drupal.org/docs/8/modules/search-api-solr/search-api-solr-ho... but it seems to be only existing for search_api_solr? For search_api_db I can't find something like that.

Furthermore I added the created and changed date as indexed field and though perhaps I could use Number field-based boosting for these fields, as they are timestamps. But these fields do not appear (otherwise that might be a nice solution). Just the uid and the "created by" field appears as boostable.

I also found https://drupal.stackexchange.com/questions/107023/how-to-boost-relevance... but think it shouldn't need custom implementations, as more complicated boostings are already possible in the UI.

So what am I missing here?
If this isn't available yet, could this please be changed into a Feature request?

Steps to reproduce

Proposed resolution

Remaining tasks

Issue fork search_api-3347901

Command icon Show commands

Start within a Git clone of the project using the version control instructions.

Or, if you do not have SSH keys set up on git.drupalcode.org:

Comments

Anybody created an issue. See original summary.

drunken monkey’s picture

Category: Support request » Feature request

No, this isn’t yet possible, but yes, we can change this into a feature request for it.
Unfortunately, I don’t have time to work on this myself at the moment (or in the near future), but I’d definitely help reviewing if someone else wants to give it a try.

anybody’s picture

Priority: Minor » Normal

Thank you very much for your reply @drunken monkey!

Could you give some feedback, if you think it would (not) make sense to simply expose the date fields to the existing number based boosting? At least if the values are timestamps (like created, changed, ...) that sounds relatively simple.

Some ideas, what you think otherwise could make sense, would be great. If you can say something about it without investing much time, of course! Just to ensure you like the way it goes...

ryan-l-robinson’s picture

+1 for this feature request

drunken monkey’s picture

Title: How to boost more recent content with database backend? » Allow dates to be used in number field-based boosting
Component: Database backend » Plugins
Status: Active » Needs review
StatusFileSize
new4.17 KB

You’re right, if that’s already a “good enough” solution for you, just treating dates as numbers would be pretty easy. It’s implemented in the attached patch – please test/review.
However, please note that this is a very rough mechanism, which will probably not work as expected for you: since dates are represented as the number of seconds since 1970, scores for dates in the same year will be very close to each other, so this will only have very limited effect. The proper way to boost recent content would be to add a dynamic boost based on not only the indexed date field value, but also the current date. This would, unfortunately, require support from the database backend.

anybody’s picture

Thanks @drunken monkey for your feedback AND work on this incl. tests. Whao! :)

I'll review this and mark it RTBC if it works as expected (I think it will - mitigated by the limitations you mentioned).

Shall I create a follow-up feature request for a full-featured integration?

drunken monkey’s picture

@Anybody: Any news on this? Would be great to get it committed.

Regarding the follow-up: I think we might as well leave this one open, right? But I’m fine either way. Very much depends on someone wanting to work on it, anyways – the issue will likely not be the main problem.

anybody’s picture

@drunken monkey sorry for the missing feedback! Too many projects in parallel, I sadly didn't have the time to try this yet. I'll reply as soon as possible. Sorry.

But I still like the plan and still think this makes a lot of sense.

anybody’s picture

Status: Needs review » Reviewed & tested by the community

Sorry for the delay @drunken monkey works fine for me so far, so RTBC! :) Thank you for the super helpful solution!

  • drunken monkey committed eb1431ae on 8.x-1.x
    Issue #3347901 by drunken monkey: Added possibility for dates to be used...
drunken monkey’s picture

Title: Allow dates to be used in number field-based boosting » Add a way to boost recent content
Status: Reviewed & tested by the community » Active

Good to hear, thanks for reporting back!
Merged.

As mentioned above, leaving this open in case someone wants to work on a proper solution.

herved’s picture

Hi everyone, I'm trying to rank newest documents higher according to a date field.
This patch is interesting but after trying it on that date field it doesn't seem to behave as expected.

This would be pretty easy with Solr but with this method, I wasn't able to apply a relevance boosting independently from keywords boosting. If the date field is present on all documents, and we search for some keywords, the boosting on the date doesn't seem to have any effect whatsoever since it seems to get merged with the keywords scoring/boosting.
And if we don't search for any keywords, this boosting has no effect at all and will be skipped in \Drupal\search_api_db\Plugin\search_api\backend\Database::createDbQuery

Am I missing something?

anybody’s picture

Priority: Normal » Major

Okay I think this is still a highly relevant feature for non-solr Search API configuration

It's absolutely typical to boost more recent content, but the workaround here isn't a proper fix. With the available number based boosting options you're able to boost the timestamp, but the resulting boost value is something like:

1708939471 * 1.0 = Boost: 1708939471
1708939471 * 1.1 = Boost: 18798334181
and even
1708939471 * 0.1 = Boost: 170893947.1

So I think we need a solution similar to what search_api_solr provides: https://www.drupal.org/docs/8/modules/search-api-solr/search-api-solr-ho... for the "created" and "updated" fields. Any ideas? I'm sadly not very experienced in this topic :/
Here's the solr processor implementation: https://git.drupalcode.org/project/search_api_solr/-/blob/4.x/src/Plugin...

Hope it's okay to set this feature request to major, as it's a really typical and relevant use-case in my eyes and not everyone has SOLR available.

anybody’s picture

Title: Add a way to boost recent content » Add a way to boost recent content (without SOLR)
Related issues: +#1694698: How to boost recently created content?, +#2335837: Boost Newer Content
herved’s picture

#13: this is what I attempted to explain in #12.
To me the patch from #5 now included is not really what most people would need because it is applied as a multiplier of the overall keyword score.
Therefore it is not independent from keywords, unlike in solr which gets added to the overall score, see https://git.drupalcode.org/project/search_api_solr/-/blob/4.x/src/Plugin...
Also, it has no effect if no keywords are given.

And if we don't search for any keywords, this boosting has no effect at all and will be skipped in \Drupal\search_api_db\Plugin\search_api\backend\Database::createDbQuery

scott_euser’s picture

Yes it looks like we need a separate boost plugin that works at the query stage rather than the process stage so it can be relative to the current date. Then whatever the current score is, add in the date factor to cause decay to the score perhaps:

So for simplicity if current is
SELECT item_id, score
FROM...

Then we could update this to
SELECT item_id, (score * 1 / (1 + TIMESTAMPDIFF(MONTH, FROM_UNIXTIME(value), CURDATE()) / 12)) AS score
FROM...

This would result in approximately:
Current year: Score x 1 - no change
1 year ago: Score x 0.5
2 years ago: Score x 0.333
3 years ago: Score x 0.250
4 years ago: Score x 0.200
10 years ago: Score x 0.100
15 years ago: Score x 0.067

Its not quite like the SOLR one but its a somewhat close approximation I think.

The 12 is the 'half-life' - ie, how long does the score take to divide in half. So the site builder could configure that, eg to make it divide in half in 60 months (5 years), or in 6 months, or whatever they want.

Any thoughts on that approach?

drunken monkey’s picture

Priority: Major » Normal

Yes, something like this seems like it would be the way to go. Instead of a processor, it could also be a Views query option (in the SearchApiQuery class). In any case, though, this will need to be declared as a feature so backends can declare whether they support it or not.

For backends supporting the feature, the processor (or the query plugin) could add an option to the search query containing both the used date field and the “half-life” value. The backend plugin would then need to check that option at search time and modify the result scores accordingly. (One additional hiccup being that, for the DB backend, we’d actually want to support all three DBMSs which I’d guess would all need a slightly different syntax for this.)

herved’s picture

Ideally it should also support future dates, resulting in a spike curve where the peak is NOW and decays in past or future.
In mysql at least, it looks like we can achieve the same formula as the one used in solr:
- solr: product(boost,recip(abs(ms(resolution,field_name)),m,a,b))
- mysql: boost * ( 1 / (m * ABS(TIMESTAMPDIFF(SECOND, FROM_UNIXTIME(field_name), NOW()) * 1000) + a) + b)

However I don't really know what would be the best to combine both numbers (keyword score that can be very high 1000-infinite), and this date boost. Maybe some kind normalization would be needed?
Something like: (score/(min_score+score)) where min_score = 1000
Then use a = 1, b= 0.1 producing a 0-1 date boost
and add both results?

I noticed that when keywords are given, we may need to join the main index table as search_api_db produces a group_by query which also makes things quite complicated.

hamilton1341’s picture

I found this snippet in src/Plugin/search_api/processor when I was looking into making my own patch:

        if ($value) {
              // Normalize values from dates (which are represented by UNIX
              // timestamps) to be not too large to store in the database.
              if ($field->getType() === 'date') {
                $value /= 1000000;
              }
              // Make sure the value is never negative.
              $value = max($value, 0);
              $item->setBoost($item->getBoost() * (double) $value * (double) $settings['boost_factor']);
        }

I'm not sure how to test that this is working exactly, but it seems like you can treat dates such as "Authored on" as an Integer in the fields options, then use number-based boosting. This code implies it will "normalize" dates, which is close to the goal here, no?

herved’s picture

Status: Active » Needs review

Pushed a branch with a different take on this: a query-time date boost for the DB backend, instead of the index-time approach in NumberFieldBoost.

- Additive & independent of the keyword score (recip curve, like Solr's recip()):
score += weight / (|now − date| / half_life + 1).
So it expresses recency relative to now and also works without keywords, neither of which the index-time version can do.
- Peaks at now, decays symmetrically (ABS), so future-dated content (e.g.: events) is handled.
- New search_api_db processor date_field_boost (only applies to DB indexes via supportsIndex) that passes its config via query option; the backend applies it in setQuerySort() (only when ordering by relevance) and marks the result uncacheable.
- Two knobs: weight (peak points, in relevance-score units) + half_life_days.
- Kernel tests included.

Note on tuning: because the boost is additive, the right weight is relative to your keyword-score magnitudes. Depending on your other field boosts you may need to set it quite high to have a visible effect (FWIW I'm using 100).

I left the existing date support in NumberFieldBoost untouched (the two can coexist) but happy to deprecate/remove it if you'd rather not have two date-boosting paths. Feedback welcome on the approach/placement.

Disclaimer: I leaned heavily on Claude Code for the implementation, hopefully I steered it in the right direction.

---

Edit: I went additive because the client on my project wants recent items pushed up quite strongly. But offering a choice between additive and multiplicative would be easy to add and multiplicative might make the better default, since it's self-normalizing and closer to what mature search engines tend to do I believe.
Let me know if that's worth doing.

herved’s picture

StatusFileSize
new20.11 KB

Current MR snapshot

anybody’s picture

Thanks @herved that looks great! I left a comment.

But offering a choice between additive and multiplicative would be easy to add and multiplicative might make the better default, since it's self-normalizing and closer to what mature search engines tend to do I believe.

I like that idea!
But I think also there it needs some explanation or examples also in the UI, otherwise one might misunderstand and misconfigure it.

Interested to hear what the others here think about it.

drunken monkey’s picture

Status: Needs review » Needs work

At a glance this looks great, thanks!
I have two requests, though:

  • Multiplicative boosting does sound more useful than additive, so if you need the latter then maybe making it configurable is the way to go. Then you can also remove the “independently of the keyword score” phrase from the processor description and explain the behavior in a bit more detail in the boost mode field.
  • Instead of being specific to the DB backend, this should use a “feature” to allow other backends to support this functionality as well.

I made part of the changes but ran out of time, so please address the rest.

The processor itself is simple enough that it should be fine to only have tests for the DB implementation of the feature, not the processor functionality directly. It would be good to add the processor to ProcessorIntegrationTest::testProcessorIntegration(), though, so we can make sure that the form works correctly. (Though it would need a slight update to assign a test server to the index and make the test backend plugin support the new feature, at least for this test.)

herved’s picture

Assigned: Unassigned » herved

Thank you both for the feedback and review!
I'll continue on this shortly

herved’s picture

Status: Needs work » Needs review

The additive scoring I committed before had a major flaw: it requires knowing the magnitude of current scores beforehand, so we know what date boosting weight to set (hence why I needed a weight of 100).
I committed a multiplicative approach with a slightly different formula. It works very well for my needs.

---

On the new formula

Each result's score is multiplied by a freshness factor:
score = score × (1 + weight / (|now − date| / half_life + 1))
* At "now" the factor is 1 + weight, so recent items are lifted most.
* It decays back to 1 for old or future dates, so the boost only ever lifts a result and never pushes it below its base score.
* weight controls how strong the lift is. 0 disables boosting.
* half_life controls how fast the lift fades (aka the curve of the slope).

Why multiply instead of add, and how this relates to Solr

search_api_solr is effectively additive: it adds a freshness number onto the keyword score. That works for Solr because BM25 keeps keyword scores in a small, predictable range, so a fixed added number is meaningful.

The DB backend's scores are not BM25. They are a simpler TF sum, multiplied by 1000, whose size changes with field count, boosts, and query length. So adding a fixed number is hard to tune.

Multiplying sidesteps that: a 2x bonus is 2x whether the score is 5 or 5000, with no need to know the score magnitude. So this reaches the same goal as Solr, where recent content ranks higher while still being scaled to relevance, but uses multiplication because it is self-normalizing for a backend whose scores are not bounded like BM25.
---

I hope I addressed all remaining points, so this is ready for another round of review,
Thanks

herved’s picture

StatusFileSize
new27.25 KB

Updated MR snapshot

herved’s picture

Status: Needs review » Needs work

The rounding - ROUND() - introduced in the last commit is causing too many ties and causing imprecise ordering.

herved’s picture

Assigned: herved » Unassigned
Status: Needs work » Needs review
StatusFileSize
new27.07 KB

Updated MR snapshot

herved’s picture

Both open threads in the MR are addressed.
I rebased the branch and fixed a few smaller things while re-reading it.

Sorry for the delay, I did not get a notification for your last review.
Ready for another round, thanks!

---

PS: The failing tests look unrelated to this issue, they come from a twig 3.28 deprecation on core's TwigSandboxPolicy::checkSecurity(), which core 11.4.5 already fixed.
Side note: .ignored-deprecations.txt shadows core's .deprecation-ignore.txt (only one ignoreFile is supported) and duplicates most of it, so appending the project-specific patterns to core's list instead would need less manual re-syncing.