I have set the encoding for my MODX MySQL databases to utf8mb4. Since the MODX character set is utf8, there is a mismatch.
I'm not experienced enough in database encoding to know if this is an issue or not. Should I be concerned, or is it fine to leave this as mismatch? Better still, maybe utf8mb4 could be added to the 'modx_charset' drop-down list?
-
☆ A M B ☆
- 24,524 Posts
The MODX charset is used for setting the HTTP headers. I've never heard of utf8mb4 applied to anything but MySQL.
Note that MySQL does not speak the same language as everyone else. When MySQL says "utf8" it really means "some weirdly retarded variant of UTF-8 that is limited to three bytes for god knows what ridiculous reason". If you really want UTF-8 you should tell MySQL that you want this weird thing MySQL likes to call utf8mb4. Don't bother saving on the "WTF!"s. – R. Martinho Fernandes Apr 9 '13 at 9:21
So the issue is that MODX needs to recognize utf8mb4 as valid utf8.
OK, so, this 'mismatch' isn't really, it's a non-issue, and I can 'let sleeping dogs lie'?
-
☆ A M B ☆
- 24,524 Posts
A bit of further research and experimentation turned up a much more serious potential problem.
In MySQL, utf8 is 3 bytes. It's supposed to be 4 bytes, so in MySQL its utf8 cannot support the full character set. Their compromise is utf8mb4, which does use 4 bytes. The problem is in the byte limit of the table data types and the indexes. You can't have as many utf8mb4 characters (at 4 bytes each) as you can of utf8 characters (at 3 bytes each). So what is a perfectly good index in utf8 will be too long in utf8mb4. Some other data types, especially small ones like 'tinytext', are also likely to be affected.
This can cause the MODX setup to fail to create some tables, as now the index may be too long. The setup does not report this error, and appears to complete an installation that actually was not successfully completed - essential tables weren't created.
If there actually is a successful installation, then since the utf8mb4 character set of MySQL is a technical match to proper utf8 (everybody is using 4 bytes per character), technically there is no problem in the mismatch, it's just an issue of the comparison code not taking this into account and returning that literally "utf8" is not a match to "utf8mb4".
I worked on a client's site where almost all the MySQL server variables were set to utf8mb4 and MODX was set to utf8. Everything seemed to work fine, but maybe there are some issues lurking.
Interesting, something new everyday in these forums!
https://dev.mysql.com/doc/refman/5.5/en/charset-unicode-utf8mb4.html
The only cases I have encountered (so far) where utf8mb4 was 'required' is Chinese and Emoticons. There are obscure alphabets that need it. – Rick James May 6 at 20:33
http://stackoverflow.com/questions/30074492/what-is-the-difference-between-utf8mb4-and-utf8-charsets-in-mysql
So, as I understand it, unless you plan to store more than 65,536 codepoints of unicode AND using char instead of varchar there shouldn't be an issue.
-
☆ A M B ☆
- 24,524 Posts
MODX won't install correctly on my localhost. Several tables can't be created because the index string is greater than 1000 bytes. While utf8 is probably sufficient for the vast majority of installations, it really shouldn't break like this if there's a 4-byte character set being used.
@Sottwell Was all you did creating a table to use utfmb8?
Do your server have the InnoDB engine?
The 1000 byte key length restriction is in MyISAM, NOT InnoDB...
http://dev.mysql.com/doc/refman/5.1/en/myisam-storage-engine.html
-
☆ A M B ☆
- 24,524 Posts
I'm just doing a simple standard installation, specifying utf8mb4 for the charset and collation for letting the setup create the database. There's no problem at all using utf8, but using utf8mb4 causes the index errors on several tables, and they are not created.
I am aware that MyISAM uses B-tree indexes, with a limit of 1000 bytes. InnoDB uses something else, with a much smaller limit.