Problem/Motivation

Two preg_replace() calls in AliasCleaner::cleanString() work on bytes (no u modifier) on a UTF-8 string. In byte mode, what \s matches depends on the current locale. Drupal core calls setlocale(LC_ALL, 'C.UTF-8', 'C') in DrupalKernel, and on macOS the C library treats the byte 0xA0 as whitespace in that locale. 0xA0 is the second byte of à (C3 A0), so \s cuts that character in two:

  1. Ignore words ("Strings to Remove"). When the list contains a word like à or là, splitting it leaves a lone 0xC3 byte in the pattern. mb_eregi_replace() emits Warning: mb_eregi_replace(): Pattern is not valid under UTF-8 encoding on every alias generation and returns FALSE, so no ignore word is removed at all, ASCII ones included. With the "Create a new alias. Leave the existing alias functioning." update action, saving existing content then silently creates new canonical aliases.
  2. Whitespace replaced by the separator. With transliteration turned off, every à in the text is cut as well: Voilà le résumé becomes the alias voil?-le-résumé.

On Linux, both glibc and musl (Alpine) do not treat 0xA0 as whitespace in C.UTF-8, so the bug does not show up there, which explains why it went unnoticed. Matching bytes on a UTF-8 string is still wrong on every platform.

Minimal demonstration in plain PHP (macOS):

var_dump(preg_match('/\s/', "\xA0"));  // int(0)
setlocale(LC_ALL, 'C.UTF-8');
var_dump(preg_match('/\s/', "\xA0"));  // int(1) on macOS, int(0) with glibc or musl

Reproduced on macOS 15.7, PHP 8.3 and 8.4, Drupal 11.4, Pathauto 8.x-1.15 (same code on 8.x-1.x).

Steps to reproduce

On macOS, with Pathauto installed:

  1. In Configuration > Search and metadata > URL aliases > Settings, turn off "Transliterate prior to creating alias" and set "Strings to Remove" to à, de, une.
  2. Run drush php:eval 'echo \Drupal::service("pathauto.alias_cleaner")->cleanString("Une série de fiches à lire");'
    Expected: série-fiches-lire. Actual: the mb_eregi_replace() warning, and une-série-de-fiches-?-lire.
  3. Run drush php:eval 'echo \Drupal::service("pathauto.alias_cleaner")->cleanString("Voilà le résumé");'
    Expected: voilà-le-résumé. Actual: voil?-le-résumé.

Proposed resolution

Add the u modifier to these regexes. In UTF-8 mode, \s only matches whole whitespace characters and never a byte inside a multibyte character, whatever the locale.

  • The two patterns that split the ignore words list, plus the preg_replace() fallback used when mbstring is missing, since the pattern contains UTF-8 words.
  • The whitespace replacement. With u, preg_replace() returns NULL on a string that is not valid UTF-8, so that case falls back to the current byte-mode call and its result stays the same.

Note: ignore words are removed after transliteration, so with transliteration on, an accented ignore word still never matches (the text already reads a). That ordering is handled in #3311669: Punctuation processed before replacing strings, not here.

Remaining tasks

  • Review the merge request. Two of its kernel tests only fail where the locale treats 0xA0 as whitespace, so they are skipped on Drupal CI (Linux). The third one (invalid UTF-8 input) runs everywhere.

User interface changes

None.

API changes

None.

Data model changes

None. On affected sites, newly generated aliases will have the ignore words removed and keep accented characters whole, as their configuration intended.

AI-Generated: Yes (Claude Code was used to help draft this issue summary and to write the fix and its test cases. I reviewed them, and on macOS each new test was confirmed to fail without its part of the fix and to pass with it.)

Issue fork pathauto-3628857

Command icon Show commands

Start within a Git clone of the project using the version control instructions.

Or, if you do not have SSH keys set up on git.drupalcode.org:

Comments

flocondetoile created an issue. See original summary.

flocondetoile’s picture

StatusFileSize
new983 bytes
flocondetoile’s picture

Issue summary: View changes

mably made their first commit to this issue’s fork.

mably’s picture

Title: Ignore words regex becomes invalid UTF-8 when \s matches the 0xA0 byte (e.g. on macOS), so ignore words are silently not removed » AliasCleaner regexes split UTF-8 characters when the locale treats the 0xA0 byte as whitespace (e.g. on macOS)
Issue summary: View changes
mably’s picture

Status: Active » Needs review