We launched new forums in March 2019—join us there. In a hurry for help with your website? Get Help Now!
    • 22303 MODX Staff
    • 10,725 Posts
    PMS: The limitation you are speaking of is PHP’s not MODx’s and you can work with multi-byte characters just fine in both in most cases without any problems. MODx does not truncate any fields using string functions, this maybe Ditto or some other third party add-on which we provide as an example of how to script things in MODx. I think you are confusing yourself somewhat.

    The fact still stands, most activities work fine with multi-byte characters and without the mb_string functions; it’s going to be the add-ons you choose for MODx that use various string manipulation functions that are going to be a problem and all that needs be done is make the mb_string functionality optional in the component.
      • 22851
      • 805 Posts
      I apologise, since I should have checked the facts before pointing the finger at php’s inbuilt string functions for the truncation of the description field. embarrassed I wrote that post at 1.44am, which explains why I wasn’t being very coherent! I’ll describe the issue in more detail:

      The truncation actually occurs at the point of entry into the mysql database, since the description column in the site_content table is limited to 255 bytes varchar(255).

      Let me see if I can demonstrate the problem. First, here is a simple html form, with a text input field with maxlength="255" (just like the modx description field). This form posts to itself and displays the results.
      <!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Strict//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-strict.dtd">
      <html xmlns="http://www.w3.org/1999/xhtml">
      <head>
      <title>Maxlength Form Text</title>
      </head>
      <body>
      <form action="<?php echo htmlspecialchars($_SERVER['REQUEST_URI']); ?>" method="post">
      <fieldset>
      <input type="text" maxlength="255" name="text" />
      <input type="submit" />
      </fieldset>
      </form>
      <?php
      if ( isset( $_POST['text'] ) )
      {
          echo '<p>"' . $_POST['text'] . '"</p>';
      }   
      ?>  
      </body>
      </html>
      

      You can type a maximum of 255 characters into this field, whatever the encoding. So, for example, here are 255 single byte characters:
      012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234

      Here are 255 multiple byte characters (these happen to be 3 bytes each in utf-8)...
      012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234

      and here is a mixed bunch of 255 single and multibyte characters.
      かわいいasdasかわい23<>いかわいいかわいいasdasかわいいかわいいかわいいかわいいかssわいいかわいいかわいいかわいいかわいasdasdかわいいかわいいかわいいかasdasdわいいかわいいかわいいかわいいasdasかわい23<>いかわいいかわいい asdasかわいいかわいいかわいいかわいいかssわいいかわいいかわいいかわいいかわいasdasdかわいいかわいいかわいいかasdasdわいいかわいいかわいいkaかわいいasdasかわい23<>いかわいいかわいいasdasかわいいかわいいかわいいかわいいか

      You can submit any of those strings in the html form and you get back what you put it. Now try the same with the MODx description field. Copy and past a string in, save the document, and then check that the string is correctly reproduced from end to end.

      For me at least (modx 0.9.6.2, utf-8 everywhere), the single byte character string works fine.

      The multibyte character is truncated at 255 bytes, which happens to mean that I get back exactly 85 of these 3-byte utf-8 characters:
      0123456789012345678901234567890123456789012345678901234567890123456789012345678901234


      When I try the mixed single byte multibyte string there is a problem. I get 115 characters back, the last of which is an incomplete and invalid utf-8 character:
      かわいいasdasかわい23<>いかわいいかわいいasdasかわいいかわいいかわいいかわいいかssわいいかわいいかわいいかわいいかわいasdasdかわいいかわいいかわいいかasdasdわいいかわいいかわいい�

      This occurs because the string has been truncated at 255 bytes, which, in this case, doesn’t lie on a character boundary.

      It has happened several times now that a description has been written that is ~100 Japanese characters and, after saving, the description has been truncated generating invalid unicode. If, like I do, you serve your pages as application/xhtml+xml to browsers that support it, this error stops all parsing and display of the page by the browser (which is an argument against application/xhtml+xml, I know).

      I know that this is going slightly off the original topic now, but would it be possible that, before modx inserts a string into a database field that has a fixed length in bytes, it could truncate it in a safe way using the mb_substr function, if it is available that is? Doing this on every mysql call could potentially have performance issues, but only doing it when pages are submitted via the manager might be a good compromise.

      Cheers.
        YAMS: Yet Another Multilingual Solution for MODx
        YAMS Forums | Latest: YAMS 1.1.9 | YAMS Documentation
        Please consider donating if you appreciate the time and effort spent developing and supporting YAMS.
        • 25663 MODX Staff
        • 12,272 Posts
        Thank you for your research and providing for great info. Did this get filed into Jira as a bug report? I think it makes sense to do so.
          Ryan Thrash, MODX Co-Founder
          Follow me on Twitter at @rthrash or catch my occasional unofficial thoughts at thrash.me
          • 22303 MODX Staff
          • 10,725 Posts
          Quote from: rthrash at Dec 20, 2008, 09:58 AM

          Thank you for your research and providing for great info. Did this get filed into Jira as a bug report? I think it makes sense to do so.
          Correction, please do not enter this as a Jira bug report, this is not a bug IMHO.  MODx does no substr() on anything, so the problem you describe is a result of the way PHP and MySQL work together.  This is a natural result of the database design and the inability of the HTML form to prevent input beyond the proper number of bytes if you are not serving the data with an encoding that is compatible with your character set data being managed by MySQL.

          Once again, mb_string is not a standard PHP extension that is installed by default and is not available in all environments in which MODx is used, thus it cannot be used reliably for the purpose you describe without making it useless to others that are already using it.

          Regardless, I think what you are experiencing is a result of other problems since no PHP string manipulation is involved at all in your example. You could easily author a plugin to do this for you on save if you need it that bad.
            • 25663 MODX Staff
            • 12,272 Posts
            A plugin would indeed be a great solution. It would basically function like a mini-ManagerManager to watch and count the characters being input into the respective fields. It would then count the number and warn when the number of bytes was exceeded. I’ll see if I can dig up an example that did something similar. OK here’s a starting point if someone wants to play with this idea, which would solve the issue. First a bit of JS to count charset bytes from a very handy site that tackled a similar problem:

              <script type="text/javascript">
                 function checkLength() {
                    var countMe = document.getElementById("someText").value
                    var escapedStr = encodeURI(countMe)
                    if (escapedStr.indexOf("%") != -1) {
                        var count = escapedStr.split("%").length - 1
                        if (count == 0) count++  //perverse case; can't happen with real UTF-8
                        var tmp = escapedStr.length - (count * 3)
                        count = count + tmp
                    } else {
                        count = escapedStr.length
                    }
                    alert(escapedStr + ": size is " + count)
                 }
              </script>


            Next a plugin we used on a project that counted the number of characters entered into the introtext to warn folks when they exceeded the set number of characters:
            <?php
            // Name: IntroText Count
            // Description: Counts the number of characters used in the IntroText field
            // Event(s): OnDocFormPrerender
            
            // Get a reference to the event
            $e = & $modx->Event;
            
            if ($e->name = "OnDocFormPrerender") {
            
            $maxchars = (isset($maxchars) ? $maxchars : ''170'');
            $width = (isset($width) ? $width : ''300px'');
            
            ob_start();
            ?>
                    <script type="text/javascript" src="/assets/js/ezLimitedInput.js"></script>
                    <script type="text/javascript">
            window.addEvent(''domready'', function(){
            $(''mutateContent'').getElement(''textarea[name=introtext]'').addClass(''ezLimitedInput'').setProperty(''title'',''MaxChars:<?php echo $maxchars; ?>|Width:<?php echo $width; ?>'').removeProperty(''style'');
            });
            
                    // <![CDATA[    
                        var oLimitedInput = new ezLimitedInput(
                        {
                            ObjectName: ''oLimitedInput''
                        });
                    // ]]>
                    </script>
            <?php
            $output = ob_get_clean();
            }
            $e->output($output);', 0, '&maxchars=Maximum Characters;text;150 &width=Width;text;300px', 0, ' '),
            ?>


            Here’s an alternate counter thing, twitter style using jQuery: http://tech.karbassi.com/2008/10/27/twitter-style-text-counter-in-jquery/

            And finally, attached find the JS that needs to be stuck in assets/js/ in order to work (after unzipping). I think the jQuery one would be nicer, but the one we built works (and requires Mootools 1.11 or later based on its report).

            This should give you all the building blocks you need to solve the problem without having to do the mbstring dance that’s not universally supported on all installations.
              Ryan Thrash, MODX Co-Founder
              Follow me on Twitter at @rthrash or catch my occasional unofficial thoughts at thrash.me
              • 22851
              • 805 Posts
              Quote from: OpenGeek at Dec 20, 2008, 10:10 AM

              Correction, please do not enter this as a Jira bug report, this is not a bug IMHO. MODx does no substr() on anything, so the problem you describe is a result of the way PHP and MySQL work together. This is a natural result of the database design and the inability of the HTML form to prevent input beyond the proper number of bytes if you are not serving the data with an encoding that is compatible with your character set data being managed by MySQL.
              Quote from: OpenGeek at Dec 20, 2008, 10:10 AM

              Regardless, I think what you are experiencing is a result of other problems since no PHP string manipulation is involved at all in your example. You could easily author a plugin to do this for you on save if you need it that bad.

              The source of the problem is clear to me and I would argue that the problem has got nothing to do with HTML forms. HTML forms have got it right by restricting based on the number of characters. 255 characters is a sensible length for a description field. MODx’s choice of database design allows for 255 characters if you write in English, but as few as 85 if you write in a foreign language and use UTF-8. As I have encountered with Japanese, 85 characters can be too few from time to time. With Japanese, a single Chinese style character can represent several syllables, but I imagine that for other languages, such as Russian perhaps, this 85 character limit could be met quite easily. Ideally, you’d be allowed to have 255 characters independent of the language.

              If MODx wanted to be as foreign language friendly as possible, then it could alter its database design. It could, for example, continue to use varchar, but use fields which hold four times as many bytes as the maximum number of allowed characters. Then 255 characters for the description field could be accepted in all languages.

              Alternatively, MODx could just place a restriction on the total number of bytes. In that case, I would argue that the fact that MODx allows more bytes than that to be entered into the html form and in particular, strings to be truncated in such a way that invalid characters are stored in the database, is a bug which needs fixing. I will report it as a bug on jira tomorrow if it hasn’t already been done.

              rthrash, your javascript solution is excellent and far better than allowing too many characters to be accepted and then truncating the string using php. Please note that the checkLength javascript function appears to be designed only to count the number of bytes for utf-8 encoded strings as far as I can tell. (encodeUri writes the string as utf-8 and encodes the utf-8 byte stream.) Javascript uses UTF-16 internally, where each character is represented by two bytes, with the exception of a certain group which are represented by 4 (2 bytes for the character and 2 bytes for a modifier of that character). Writing a checkLength function for that encoding shouldn’t be too difficult - you could ignore the special group and just limit based on the Javascript length of the string. Writing such a function for other encodings probably would be very difficult however.

              I will definitely try out the plugin and tell you how it goes. Thanks for your help.
                YAMS: Yet Another Multilingual Solution for MODx
                YAMS Forums | Latest: YAMS 1.1.9 | YAMS Documentation
                Please consider donating if you appreciate the time and effort spent developing and supporting YAMS.
                • 25663 MODX Staff
                • 12,272 Posts
                PMS,

                Just for a point of clarification, your concerns about db design are well considered and correct. We actually think the same thing, and have that addressed on the roadmap for a post evo/revo release. For now though we’re trying not to introduce any significant changes as we prepare for a period of transition in MODx in general. For now I think the plugin is the best course, and if you need a larger description field, then create a TV and use ManagerManager to hide the built-in one.
                  Ryan Thrash, MODX Co-Founder
                  Follow me on Twitter at @rthrash or catch my occasional unofficial thoughts at thrash.me
                  • 9869
                  • 41 Posts
                  We are discussing about the PHP5-check in our German MODX-Forum (http://www.modxcms.de/forum/) and found a error about this in the file search.class.inc.php (about line 1064):
                  if ($this->isPhp5) $text = html_entity_decode($text, ENT_QUOTES, 'UTF-8');


                  Hm, and what if its php4? then it should be coded like that:
                  if ($this->isPhp5) $text = html_entity_decode($text, ENT_QUOTES, 'UTF-8');
                  else $text = html_entity_decode($text, ENT_QUOTES);



                    --
                    Design Agency - http://fruehjahr.ch
                    • 5811
                    • 1,717 Posts
                    You forgot to precise that you use html_entity_decode($text, ENT_QUOTES, ’UTF-8’) when you have a database utf8 charset:
                          if (($this->dbCharset == 'utf8') && ($this->cfg['mbstring'])) {
                            // convert of all Html entities before extraction
                            // require version 5.0 and upper : http://bugs.php.net/bug.php?id=25670
                            if ($this->isPhp5) $text = html_entity_decode($text, ENT_QUOTES, 'UTF-8');
                            ...
                          }
                          else {
                            // convert of all Html entities before extraction
                            // require PHP 4.3
                            $text = html_entity_decode($text, ENT_QUOTES);
                            ...
                          }


                    If you add $text = html_entity_decode($text, ENT_QUOTES); this means that implicitly you use ISO-8859-1.
                    Which probably works with German, French and European Western languages that could be coded on one byte length (ISO-8859-1) but may be not with all the languages (ie: japanese, chinese).
                    Do you solve issues with this fix for the German ?
                      • 9869
                      • 41 Posts
                      I have not forgot, I think PHP4 cant understand the additionally statement "utf-8" like in html_entity_decode($text, ENT_QUOTES, ’UTF-8’). But I’m not shure...
                        --
                        Design Agency - http://fruehjahr.ch