We launched new forums in March 2019—join us there. In a hurry for help with your website? Get Help Now!
    • 28042 ☆ A M B ☆
    • 24,524 Posts
    I’m having an odd problem in Hebrew. All is well with everything utf-8, but certain characters are returning in the search results summary as blocks or question marks. This occurs at the beginning or the end of the summary, although the character at the beginning or the end is not actually the first or last character in the word. It looks as though AjaxSearch is getting confused as to what is whitespace with this character set, and breaks the utf-8 encoding for the affected character.

    I’ll dig into exactly how AjaxSearch generates that summary and report what I figure out. I’m suspecting a subtle PHP bug in the case; there are a number of cases where PHP does not handle utf-8 and other multibyte character sets very well.
      Studying MODX in the desert - http://sottwell.com
      Tips and Tricks from the MODX Forums and Slack Channels - http://modxcookbook.com
      Join the Slack Community - http://modx.org
      • 5811
      • 1,717 Posts
      You are right sottwell.
      I use the html_entity_decode php function to convert html characters into normal characters before text analysing. But this php function doesn’t work with php version lower than 5.0. - see : http://bugs.php.net/bug.php?id=25670

      Hereafter an extract of the function getExtract (line 712 of the classes/search.class.inc.php) which generate summaries:
       if ($this->dbCharset == 'utf8') {
              // convert of all Html entities before extraction
              // require version 5.0 and upper : http://bugs.php.net/bug.php?id=25670
              if ($this->isPhp5) $text = html_entity_decode($text, ENT_QUOTES, 'UTF-8');
              $mbStrpos = 'mb_strpos';
              $mbStrlen = 'mb_strlen';
              $mbStrtolower = 'mb_strtolower';
              $mbSubstr = 'mb_substr';
              $mbStrrpos = 'mb_strrpos';
            }
      
      Some workarounds to overpass the Php 4 limitations seems to be exist. But I hadn’t enought time to test and implement them in the version 1.8 of AS.
        • 13428 ☆ A M B ☆
        • 1,031 Posts
        In one of my installations this bug occours too. The content in this Documents is ’raw’ so there should be no HTML-Entities in $text. Or did you encode them to entities before?
          • 5811
          • 1,717 Posts
          Jako, for the version 1.8.1 I have corrected the following warning: ADDON-12 : A call to the mb_strrpos php function with an empty text

          This warning is due to a call to the mb_strrpos php function with an empty text. This warning occurs with the statement
          $extracts[$i]['left'] = (int) $mbStrrpos($mbSubstr($text,0,$extracts[$i]['left']),' ');
          from line 807 of the classes/search.class.php file.

          It occurs when I tried to find the closest white space on the left of a searchterm found into a text. If the text starts with the searchstring, then the text variable is empty. See http://modxcms.com/forums/index.php/topic,27316.msg174469.html#msg174469

          Not sure but It seems that this correction solve some weirds characters. For instance do a search with "education" on this search page (version 1.8.0) and do the same on this search page (version 1.8.1). In the first result page you have lot of weirds characters not after correction of this warning with the version 1.8.1. And my site runs with Php 4.4.9 under MySql 5.0.51a

          Hereafter is the correction of this issue, but may be now is better for you to wait the version 1.8.1.
                  // find the closest space character on the left & rigth side (west side story !!)
                  for($i=0;$i<$nbExtr;$i++) {
                    // on the left
                    $begin = $mbSubstr($text,0,$extracts[$i]['left']);
                    if ($begin != '') $extracts[$i]['left'] = (int) $mbStrrpos($begin,' ');
                    // on the right
                    $end = $mbSubstr($text,$extracts[$i]['right']+1,$textLength - $extracts[$i]['right']);
                    if ($end != '') $dr = (int) $mbStrpos($end,' ');
                    if (is_int($dr)) $extracts[$i]['right']  += $dr + 1;
                  }
          


            • 13428 ☆ A M B ☆
            • 1,031 Posts
            The Patch seems not to be the solution. Try http://www.partout.info/suche.html and search for ’über’. The first char below the entry ’impressum’ is wrong.

            BTW: The (menu)links on modx.wangba.fr are sometimes wrong. Try to get from http://modx.wangba.fr/index.php?id=51 further by the menu and from http://www.wangba.fr/mod2/index.php?id=141 too.
              • 9869
              • 41 Posts
              sad hm, the problem still exists:
              check this picture: http://img.skitch.com/20081215-xxs6ue6emdk9q4a9mjhp22nnpw.png
              Is there a solution?
                --
                Design Agency - http://fruehjahr.ch
                • 5811
                • 1,717 Posts
                @Flunghund

                What is your Mysql database encoding ? your page encoding ?
                And are you sure that tinyMCE don’t encode accent with html_entities. Check that the configuration of tinyMCE is set to "raw" for html_entity parameter.
                  • 9869
                  • 41 Posts
                  @coroico:
                  DB, Page encoding and settings are utf-8. tinyMCE encodes accents to entities. But if I set it to "raw" its not valid and I have 300 Pages with content alredy filled...!!
                  Any other idea?
                    --
                    Design Agency - http://fruehjahr.ch
                    • 26422
                    • 107 Posts
                    I’m not sure if this is identical but it seems similar. I’m getting weird characters in the search results where regular " (quotes) should be. MY DB and Page are also utf-8.

                    Attached is an example of what is going on.

                    Any ideas?

                    Crawford.
                      Crawford Paul

                      Bridgecourt Web Design - Proudly using ModX CMS
                      Serving Welland, Niagara, Ontario, Canada and the world!
                      http://www.bridgecourt.com
                      • 13094
                      • 58 Posts
                      Hi Coroico,
                      I just (was) updated to 1.8.1 aside the 0.9.6.3 update of modx.
                      Since than the "Umlauts" in the href enclosure weren’t encoded correctly (you know, those stupid UTF-8 replacements like ä instead of ä, for instance). All the rest, like the extract, is displayed correctly.

                      I reinstall the 1.8.1 version from the dispository but with same result.

                      My pages and MySQl are using UTF-8, too.

                      Ooh, and one more question:
                      If no keys were listed for the &whereSearch parameter (the default), why is the introtext part of the result?

                      Kind regards und thank you for your work,
                      iRolf