Problem/Motivation
Two preg_replace() calls in AliasCleaner::cleanString() work on bytes (no u modifier) on a UTF-8 string. In byte mode, what \s matches depends on the current locale. Drupal core calls setlocale(LC_ALL, 'C.UTF-8', 'C') in DrupalKernel, and on macOS the C library treats the byte 0xA0 as whitespace in that locale. 0xA0 is the second byte of à (C3 A0), so \s cuts that character in two:
- Ignore words ("Strings to Remove"). When the list contains a word like
àorlà, splitting it leaves a lone0xC3byte in the pattern.mb_eregi_replace()emitsWarning: mb_eregi_replace(): Pattern is not valid under UTF-8 encodingon every alias generation and returnsFALSE, so no ignore word is removed at all, ASCII ones included. With the "Create a new alias. Leave the existing alias functioning." update action, saving existing content then silently creates new canonical aliases. - Whitespace replaced by the separator. With transliteration turned off, every
àin the text is cut as well:Voilà le résumébecomes the aliasvoil?-le-résumé.
On Linux, both glibc and musl (Alpine) do not treat 0xA0 as whitespace in C.UTF-8, so the bug does not show up there, which explains why it went unnoticed. Matching bytes on a UTF-8 string is still wrong on every platform.
Minimal demonstration in plain PHP (macOS):
var_dump(preg_match('/\s/', "\xA0")); // int(0)
setlocale(LC_ALL, 'C.UTF-8');
var_dump(preg_match('/\s/', "\xA0")); // int(1) on macOS, int(0) with glibc or muslReproduced on macOS 15.7, PHP 8.3 and 8.4, Drupal 11.4, Pathauto 8.x-1.15 (same code on 8.x-1.x).
Steps to reproduce
On macOS, with Pathauto installed:
- In Configuration > Search and metadata > URL aliases > Settings, turn off "Transliterate prior to creating alias" and set "Strings to Remove" to
à, de, une. - Run
drush php:eval 'echo \Drupal::service("pathauto.alias_cleaner")->cleanString("Une série de fiches à lire");'
Expected:série-fiches-lire. Actual: themb_eregi_replace()warning, andune-série-de-fiches-?-lire. - Run
drush php:eval 'echo \Drupal::service("pathauto.alias_cleaner")->cleanString("Voilà le résumé");'
Expected:voilà-le-résumé. Actual:voil?-le-résumé.
Proposed resolution
Add the u modifier to these regexes. In UTF-8 mode, \s only matches whole whitespace characters and never a byte inside a multibyte character, whatever the locale.
- The two patterns that split the ignore words list, plus the
preg_replace()fallback used when mbstring is missing, since the pattern contains UTF-8 words. - The whitespace replacement. With
u,preg_replace()returnsNULLon a string that is not valid UTF-8, so that case falls back to the current byte-mode call and its result stays the same.
Note: ignore words are removed after transliteration, so with transliteration on, an accented ignore word still never matches (the text already reads a). That ordering is handled in #3311669: Punctuation processed before replacing strings, not here.
Remaining tasks
- Review the merge request. Two of its kernel tests only fail where the locale treats
0xA0as whitespace, so they are skipped on Drupal CI (Linux). The third one (invalid UTF-8 input) runs everywhere.
User interface changes
None.
API changes
None.
Data model changes
None. On affected sites, newly generated aliases will have the ignore words removed and keep accented characters whole, as their configuration intended.
AI-Generated: Yes (Claude Code was used to help draft this issue summary and to write the fix and its test cases. I reviewed them, and on macOS each new test was confirmed to fail without its part of the fix and to pass with it.)
| Comment | File | Size | Author |
|---|---|---|---|
| #2 | pathauto_ignore_words_utf8_regex.patch | 983 bytes | flocondetoile |
Issue fork pathauto-3628857
Show commands
Start within a Git clone of the project using the version control instructions.
Or, if you do not have SSH keys set up on git.drupalcode.org:
Comments
Comment #2
flocondetoileComment #3
flocondetoileComment #6
mably commentedComment #7
mably commented