We launched new forums in March 2019—join us there. In a hurry for help with your website? Get Help Now!
    • 22797
    • 134 Posts
    MODx has been a good tool for me to work with over the last couple of years. It’s extensible, powerful, and easy to use for the most part. Among its disadvantages, though, is that MODx is designed around data input more than around data retrieval.

    It’s great that I can keep adding on template variables onto templates, essentially turning templates into pseudo database tables. But that’s also part of the problem. They aren’t real database tables. The data are split up into many tables, and when retrieving documents with many template variables, this involves a lot of queries which slow down data retrieval considerably, e.g. with a lot of LEFT JOIN circumstances--one for each TV (at least I think that’s how it’s set up).

    For sites with complex data relationships, something has to be done. As it is, MODx doesn’t cut it.

    Data retrieval would be a lot faster if each template were represented by a single table of all the TVs for that template only, and if each TV had a specific data type (integer, text, char, etc.), rather than throw all of the data into a single table (modx_site_tmplvar_contentvalues) of the same data type (text).

    This sort of data structure would also allow for much, much more efficient joins of data across templates. As it stands right now, joins that would be quite simple between two tables in a "normal" data structure involve many more queries in MODx, and the result is unacceptably slow. In fact, I’ve resorted to using a homemade plugin that creates a new, separate table for each template for this very purpose. The plugin recognizes when a document is saved or deleted, and sends changes to these other tables. This means that there are always two copies of the data: one in the native MODx structure, and another in the tables created and maintained by my plugin.

    My plugin isn’t ready for the public, in part due to important omissions in the MODx event triggers (onBeforeMoveDoc and onMoveDoc don’t exist, for example; neither do onBeforeDocFormUndelete or onDocFormUndelete). With these important omissions, it’s a bit tricky to ensure the derivative tables are always up to date with the MODx tables.

    I think we can have a structure that is good for both data input and data retrieval. There is more than one way to accomplish this:

    1. Maintain two parallel versions of the data, as I’m doing with my plugin. One version (MODx) is for data input, the other (plugin derivative data) is for data retrieval. This creates a "shadow" or "mirror" of the MODx data. This method has the advantage of not breaking the way MODx currently works. It is backward compatible.

    2. To avoid data duplication, change the way MODx stores template variable data. Put as much data as possible in the same table that is *specific to that template*. Each template would have its own table, with its own structure, with defined data types, lengths, etc. If there are 24 TVs associated with a template, there would be 24 fields in the corresponding TV table (or more, if we include meta data such as id, date modified, etc.)

    3. The same as #2 above, but with a twist: the TV tables could even store a copy of much or all of the information that is currently in the modx_site_content table. I don’t think I’d want to get rid of the modx_site_content table, because having all the basic page info in the same place is good for fast queries, but storing a copy of that information in the TV tables would optimize other kinds of queries. Yes, it would duplicate information, and maybe that’s a good enough reason not to do it, but it makes joins across tables (templates) much easier and faster.

    I would probably advocate for solution #1 as a starting point, mainly for backward compatibility. I could refine my own code and hand it over if it would be useful. But MODx would need to be altered to ensure that there aren’t back door ways of bypassing the plugin as there are currently.

    ----

    Looking into the future:

    I know that MODx 0.97 represents a fundamental shift in the way MODx works, and I welcome all of the changes that I’ve read about, though I haven’t worked much with the code itself. From what I understand, we’ll be able to associate TVs with things other than templates, which is great, even if it does modify my proposal above.

    The basic idea of my proposal is that we need a way to optimize data retrieval by reducing the number of tables and joins for simple page views. This will also reduce the complexity of already-complex data relationships, making queries easier to construct and faster to run.

    If there is another way to do this other than what I’ve proposed, great! I’d love to hear about it.
      • 22303 MODX Staff
      • 10,725 Posts
      Quote from: paulb at Aug 18, 2008, 05:19 PM

      MODx has been a good tool for me to work with over the last couple of years. It’s extensible, powerful, and easy to use for the most part. Among its disadvantages, though, is that MODx is designed around data input more than around data retrieval.

      It’s great that I can keep adding on template variables onto templates, essentially turning templates into pseudo database tables. But that’s also part of the problem. They aren’t real database tables. The data are split up into many tables, and when retrieving documents with many template variables, this involves a lot of queries which slow down data retrieval considerably, e.g. with a lot of LEFT JOIN circumstances--one for each TV (at least I think that’s how it’s set up).

      For sites with complex data relationships, something has to be done. As it is, MODx doesn’t cut it.
      ...

      The basic idea of my proposal is that we need a way to optimize data retrieval by reducing the number of tables and joins for simple page views. This will also reduce the complexity of already-complex data relationships, making queries easier to construct and faster to run.

      If there is another way to do this other than what I’ve proposed, great! I’d love to hear about it.
      I think you are confused about this; there is no proof whatsoever that data retrieval in MODx, even with a complex list of template variables is slow.  Can you provide examples of this?  Saying that MODx doesn’t cut it is simply not the way to approach such suggestions.  First, present some evidence that describes queries that are taking too long.  I have not experienced this unacceptable performance you describe in any complex sites I’ve created with MODx and without further evidence to the contrary, I do not see the purpose in your proposal.

      Let’s establish what you conceive is the real problem before we try to address it, as I do not agree with your evaluation of the data structures.  The queries being executed on MODx requests is fairly streamlined and efficient, and the JOINs do not add any unnecessary overhead.  With the built-in caching mechanism, there are not any queries executed anyway.  Perhaps you are trying to do too much with Template Variables when custom data structures for your site’s needs are really needed.

      As another note, the future MODx data structures are being evolved already and it might be best to consider what you are talking about in conjunction with these changes. These changes will be necessary to support future internationalization and versioning features that have already been designed.
        • 22797
        • 134 Posts
        Let me back up and say that I’ve chosen MODx out of many possible contenders, and have been using it mostly happily for quite some time. It’s a fabulous product, and I even though I am offering some suggestions, I don’t do this with any ill-feeling. Sometimes typed messages appear to convey more negativity than intended, so I apologize if it gave that impression.

        I’ve put quite a lot of data into MODx that is quite tightly integrated. I’ve used template variables for this purpose because that’s the easiest way to make the data editable in a multi-user environment where most users don’t know HTML. You could easily argue that I should use a custom data format, and you’d be right. But I’m using template variables for their ease of maintenance by non-technical users.

        I’ll give a brief overview of some of my data structure. It’s an academic web site:

        • People (name, phone, address, email, bio, photo, etc.)
        • Courses (title, description, course number, etc.)
        • Course instances (sections of courses, with time, room, days, instructor, etc.)
        • Buildings (address, campus,etc.)
        • Rooms (building, room number, etc.)

        There’s a lot more, but this will illustrate the point. On semester schedule pages, there is a list of all the course instances for a given semester. To display this information, I need to pull data from all of the above tables (plus others). The data are linked by the MODx ids of the various documents that contain the relevant information. The course instance has the id of the instructor(s), the id of the course, the id of the building and room.

        There isn’t an easy way to pull all of this information together with publicly available MODx snippets. Ditto can easily get all of the course instances, but then I would still need to get all of the related information, which (if I use only Ditto) would mean running Ditto over and over inside itself essentially, which is very inefficient.

        If each of the items (people, courses, course instances, buildings, rooms) had a separate table, I could write a single -- though somewhat complex -- query to get all of the information at once. I could conceivably do the same thing by querying the modx_site_tmplvar_contentvalues table in a similar way.

        What I ended up doing was creating a single derivative table that contained all of the relevant information, so that I could query it all at once, without having to do any joins. The table contains the peoples’ names, course info, room info, etc. The "hard work" is done by the plugin when I save the document, rather than when the document is viewed on the web. It does slow down the process of saving the documents a little bit, but I’d rather have the slow-down happen there than on the public web site.

        My solution may not be the best one, but it does make queries easier and faster. I know for a fact that the schedule pages are faster now (querying one table) than they were before (querying the TV values multiple times), though I don’t have a side by side comparison because the old version no longer exists.

          • 22303 MODX Staff
          • 10,725 Posts
          It has always been my strong belief that complex data structures like this should be managed in custom data tables. MODx is designed to present views of content formatted for web presentation. IMHO, this extends to Template Variables as well; they are there to extend the presentation layer and allow flexibility in building dynamic sites and avoiding duplication of content and/or scripted logic.

          For projects like this I typically design a set of a tables and a model for those tables with xPDO; this gives me instant access to easily work with the data structure I’ve designed, and provides a foundation for implementing custom data caching strategies in the data model layer. I can then integrate that into MODx in creative ways; including creating modules or front-end admin pages for users to edit that data through custom forms. This both helps reduce the overhead on the main MODx controller that is responsible for dispatching requests, and reduces the quantity of presentation data that is being managed throughout the core, deferring that overhead to custom snippets used to present snapshots of your complex data model in views that are optimized for the web. It also reinforces the separation of business logic and presentation.

          Using the MODx content database as a generic storage solution for very specific, complex data other than that data meant for the presentation layer is not the proper approach IMHO and will result in performance problems. It can certainly work for very simple structures, but looking at your description, I have to say a custom data model is really the best approach for these requirements, especially if you are striving for optimized performance.
            • 18367
            • 834 Posts
            OG,

            just going slightly off topic here;

            If I understand you correctly, what you’re saying is that if I have a whole load of pre-existing db’s containing specific information like customer demographics, etc then I’m better off leaving them in those databases and accessing them as I need them, rather than importing them all into the general modx db?

            I have more than a dozen, they’re not that large, but I was going to import them all into modx.
              Content Creator and Copywriter
              • 22797
              • 134 Posts
              Thank you for your response. It helps me understand how you envision MODx being used.

              One of the reasons that I used TVs this time around was because it allows users to edit the data right in the context of the page. For most things, that makes a lot of sense. They log into MODx, go to the "folder" with the data, open the "page" and start editing. They understand how and why it works, because it’s all right there. They don’t have to go to a separate interface, or module, or intranet, or anything else. From the perspective of non-technical users especially, there is little to learn this way.

              From my perspective, using TVs is handy because I don’t have to create a separate admin interface for everything. I don’t have to create the module or intranet. For the simple content, all I have to do is throw in a few references to the TVs (e.g. [*course_title*], [*credit_hours*] and so on). It’s a rapid application development environment that makes data creation relatively fast and easy.

              Of course, there is a downside in performance on the retrieval end as a result, as the data get more complicated and interconnected.

              Maybe the ideal for me would be to be able to store the data in an external data structure but still be able to edit it from the pages themselves, not just modules. I’d like to be able to go to the web page for a certain course (or person, or building, or room, etc.)  and edit all the related info. Maybe this could be done with an iframe that would show the other interface, or a snippet-like script could be run within the page’s edit interface. Are either of those possible? Maybe with 0.9.7?

              To extend this further, it would be nice to have as much control over the edit interface as possible and suppress certain fields, for example the pagetitle field if I want the pagetitle to be automatically generated from the "course_title" template variable.  Maybe those sorts of things are possible in the next version of MODx.

              Technical question:

              Are all of the TVs automatically retrieved for a document when the document is retrieved, even if the TVs are not displaying on a page? For example, let’s say I have 12 TVs for a given template, all of which have data, but I don’t call that data anywhere in the page. Are all 12 of those TVs queried and retrieved before the page is rendered, even though none of them will be used? (Or, more likely, some will be used for the page, and others are used for other purposes). If it is all being retrieved no matter what, one implication is that the more TVs associated with a given template, the more of an impact this will have on performance. In my experience, this seems to be the case, though there may be other factors to take into account. Another implication is that the data may be retrieved twice if I have another way of retrieving it other than the regular MODx document object model.
                • 22303 MODX Staff
                • 10,725 Posts
                Quote from: markg at Aug 18, 2008, 11:12 PM

                If I understand you correctly, what you’re saying is that if I have a whole load of pre-existing db’s containing specific information like customer demographics, etc then I’m better off leaving them in those databases and accessing them as I need them, rather than importing them all into the general modx db?

                I have more than a dozen, they’re not that large, but I was going to import them all into modx.
                A dozen is not that many. What I’m getting at is the balance between managing web page views, which is what MODx and it’s data structures are designed for, and managing potentially hundreds or thousands of pieces of various categories of data that has meaningful business relationships involved. If it was just a a few simple pieces of data, that would be one thing, but designing web page documents representing every entity in a multi-object model and having the MODx engine responsible for all of the transactional data and business logic involved will exponentially degrade your site’s performance as the quantity of data and/or the rate of changes to the data increase.

                I view Documents in MODx sites (referred to as Resources in Revolution) as ways to simplify these situations in terms of what is to be presented to the user. The business logic and data model exists in and of itself, and in most cases, the overhead of managing this transactional data should not be imposed on your web server or the PHP controller that is responsible for organizing and presenting your web pages. It should be reserved for where it is needed: on the specific web views defined in your site that provide a window into simple individual pieces or complex aggregations of that data, and to forms designed for managing that data and any permissions related to it’s administration.

                If it’s only a dozen or a few dozen simple pieces of data, managed by a single person, that’s one thing, but will this number increase or is that all the data you’ll ever be storing and presenting as individual pages in your web site? And is there only one view of this data you will need, or might you want other views into this same data in the future (other than a Ditto-style summary of each Document)?
                  • 18367
                  • 834 Posts
                  OG,

                  it’s business data/information, that can and most likely will be used for other additional purposes including other applications.

                  And yes they will all grow.

                  So, even if they are few and small now, it seems that it would be prudent to have them in separate dbs.

                  PS Thanks for the detailed explanations.

                    Content Creator and Copywriter
                    • 29774
                    • 386 Posts
                    @paulb

                    In MODx Revolution you are able to manage custom tables within the manager, making use of custom Manager functionality provided by the new ext.js powered user interface, using the new, elegant OO structure:

                    http://svn.modxcms.com/docs/display/~splittingred/2008/06/
                    http://svn.modxcms.com/docs/display/~splittingred/2008/07

                    If you are used to other php frameworks then this will be like a breathe of fresh air. It’s excellent to have such a framework, and there are all sorts of components that will now be much easier to create (ecommerce, forum, newsletters, etc).

                    In the 0.96 release modules are the way to manage custom data structures, but I too have found that the integration with Manager, specifically the site tree, is a problem. I seem to end up using modx documents in the way you have, to model different types of objects, and then create relationships between them. Most of the time this doesn’t seem to impact performance, and I normally write custom snippets with optimised sql to provide views on the data, rather than rely on Ditto, eg:
                    http://www.gmrlaw.com/news/

                    But then my sites for the most part are <2000 pages. Most likely they wouldn’t scale that well.

                    One solution I’ve also considered is to create a dedicated cross reference table for modx documents to speed up joins between related documents. This could be maintained within the relevant Manager pages using an empty TV and the events OnBeforeDocFormRender (to populate the TV) and OnDocFormSave (to update the xref table). I find that it’s the relationships between documents that are the key to performance, rather than the TV data specific to a document, since in a typical master listing view you normally only present a title and summary, and it is the detail view where you need to grab the TV data.

                    Sorry this has turned into a bit of a ramble - this is an interesting subject, I believe it goes to the heart of why modx is attractive for rapid application development.

                    Mark





                      Snippets: GoogleMap | FileDetails | Related Plugin: SSL
                      • 34017
                      • 898 Posts
                      I think the plugin way sounds very optimal for this use compared to a seperate module. I can’t wait to see it.

                      Why I think it’s a good current work-around
                      - simplicity for user. Content changed inside current documents.
                      - data is optimized in your table for the queries on the front-end

                      The only negative I see is duplicated data. But that seems small.
                        Chuck the Trukk
                        ProWebscape.com :: Nashville-WebDesign.com
                        - - - - - - - -
                        What are TV's? Here's some info below.
                        http://modxcms.com/forums/index.php/topic,21081.msg159009.html#msg1590091
                        http://modxcms.com/forums/index.php/topic,14957.msg97008.html#msg97008