Closed (fixed)
Project:
Crawler Rate Limit
Version:
3.0.0-beta1
Component:
Documentation
Priority:
Normal
Category:
Support request
Assigned:
Unassigned
Reporter:
Created:
25 Apr 2023 at 22:05 UTC
Updated:
6 May 2024 at 11:49 UTC
Jump to comment: Most recent
Comments
Comment #2
vaish commentedYou can go to Drupal's Status Report page (/admin/reports/status) and find all the details there. If something is not right you will also see a warning or an error.
Module doesn't use any database tables and doesn't post log messages. That's intentional.
The only place module stores any information is the backend you specified in the module settings. In your case that's Redis. All keys created by this module are prefixed with
crawler_rate_limit:.Comment #3
vaish commentedDid you check module's README.md: https://git.drupalcode.org/project/crawler_rate_limit/-/tree/3.x?ref_typ...
Status Report page is mentioned there. See step 3 under "Install Crawler Rate Limit module".
If you have any specific feedback on how to improve README file, please let me know.
Comment #4
bobburns commentedYes I checked and saw it enabled, but I didn't see any way to know it was working like flood control shows in a table. Thanks for the quick response
Comment #5
vaish commentedI see. You want a proof that module does what it's supposed to do. Here is a simple way to verify that:
In settings.php configure very low number of requests per interval for "bot_traffic". For example 4 requests within 30 seconds. Then use curl on command line to issue requests to your website. First four requests should return response code 200 while the fifth and subsequent ones should return 429. After 30 seconds, counter resets and you will again be able to get 200 but only for 4 requests.
Here is a curl issuing HTTP requests as a bot:
curl --head -A "bot" "https://yourdomain.com/"You can do similar test in your browser but for this you need to use the settings under "regular_traffic". Then just keep refreshing any page on your website and keep track of the responses you get.
On production you can grep your web server's access.log for entries with response code 429. Those will be the ones that were blocked by crawler rate limit.
Note that crawler rate limit, being a Drupal module, processes only requests that are handled by Drupal. It doesn't do anything to requests that are served directly by the web server, like for example requests for images in the public files folder or requests for CSS and JS assets.
Comment #6
bobburns commentedOK good to know - I will try that
Looks like I will need to re-install tarpit and configure the path ban
Comment #9
generalredneckVaish,
I hope you don't mind, but I took the liberty to add to the readme the way I tested this functionality. It uses curl like you suggested but in a loop, allowing a quick one liner. It also uses cache busting as not everyone is going to test this against raw installations, and it's good to see that it works the first time without realizing that something like varnish is eating the request.
Comment #11
vaish commentedThanks @generalredneck. It's great to have this added to the README.