We launched new forums in March 2019—join us there. In a hurry for help with your website? Get Help Now!
    • 13808
    • 61 Posts
    I have a site running on Site5 in VPS1 (memory limit 768mb RAM). It's running 2.2.6-pl and is has just under 2900 resources. The site has been running smoothly for months and growing steadily in content to the current point.

    Recently - starting about a month ago after 6-7 rock solid months, we are experiencing some pretty weird behaviour which I'm having issues getting to the bottom of. The site is taking down the VPS it's in frequently (sometimes makes it a few days - has gone down 5 times in an hour) by filling the memory which causes caching which causes the server load to go through the roof and the machine to become unresponsive to the point where it needs to be forcefully rebooted.

    I'd fully understand this if the site was busy, but the most people I've ever seen the site serve at one time is 8 and those are 8 casual browsers - not 8 simultaneous hits on the site. There's almost no-one on the site. And serving those 8 people and me working in the manager, everything is fine.

    Then there's been times where I've watched 'top' on the machine and real time stats in Google Analytics and the server has kicked up numerous PHP processes to the point of killing the server all while Google reported zero people on the site. To this end, as best I can tell it doesn't appear to be traffic load related.

    I'm running a number of extras on the site, but nothing that accounts for multiple processes that seem to start from no traffic (cron etc).

    Site5's advice is to upgrade to raise the memory ceiling (which I'm open to), but I don't feel great about just upgrading when the site is buckling under no traffic and seemingly at random. In my mind there's no guarantee the upgrade will fix the problem when I don't know why it's causing it.

    - I've asked to see the contents of /var/log/messages but because it's a managed service they're not open to sharing the contents.
    - I've used Executor to optimise the page query times right back and this has led to no reduction in server spikes.
    - I've checked the MODX error log repeatedly - all it reports is being unable to cache static resources.
    - I've setup a cron (via CPanel) which logs the load once per minute and the results of watching it for a week shows it going down at all different times of day so it's unlikely to be a daily maintenance process or anything like that.

    Does anyone have any leads to help with further diagnosis? Is there something else valuable I could be logging somewhere or watching for?

    The alternative is to upgrade or move but as outlined above, I'm not convinced about that as a solution when I can't follow why a load of 0 visitors can crash the VPS while a (still low) load of 8 visitors and me in the manager won't.

    Any suggestions very welcome!
      • 22303 MODX Staff
      • 10,725 Posts
      It's possible that robots are spidering your site and hitting a Resource that is broken, or executing a really bad, long-running query. I recently saw a site this was happening too because their site map script was timing out. Just some thoughts, but my gut says you have some bad code on there rearing it's ugly head.
        • 22840
        • 1,572 Posts
        Had exacly the same issue for months with site 5 on a VPS8 and they just kept on blaming sites.

        Heres the long and short of the tickets / argument

        Hello Paul,

        Thanks for your patience. I've gone through the server logs and found following:

        It looks like your scripts isn't closing sessions as it should. That's why a lot of sessions being created for your script and that's causing loads on the server.

        I've found that an account in question is "inglewoo". That account is listed in top processes list on your box

        My Reply

        This is the same answer as we always receive, xxx website is causing the
        issue, so we move the website to another server and the issue persists but
        another website gets blamed !!

        I know there is a issue with the server not closing things as this is a
        known issue to site 5 and they have implemented a "Plaster" fix by creating
        a cron job that runs every six hours to restart the server ( see ticket
        XKZG-00596 ), this "FIX" obviously isn't working as intended so I think
        it's now time for site 5 to actually fix the issues on the server and stop
        blaming websites as we can't just continue to move them to another server.

        Thanks

        Them

        Hello Paul,

        We have checked your VPS and at this time the apache is set to be restarted at each 6 hours. Would you like to try to be restarted at each 3 hours? Otherwise we will have to send this ticket to our Level 3 Support Department so the issue can be checked by a Senior System Administrator.

        Please check and advise.

        Thanks,

        My Reply

        Hi Catalin,

        Restarting the server every 3 hours is just another plaster on the fix,
        this issue was discovered in October and has been ongoing ever since so
        Ideally I would like it escalating to level 3 support and hopefully fixing
        once and for all as it's taking up far too much of my time and client sites
        are constantly un available.

        Thanks

        Them

        Hi Paul,

        Our apologies for this on going issue you have been having. I've been looking over the history of everything, and it seems this may actually be due to a completely different bug that is with Apache it self - and unfortunately while we are not able to reproduce the issue on demand, we do have an internal patch that can be applied to the server to prevent it from happening all together.

        This has been applied to 3 other customer VM's and has worked each time. The issue at hand, assuming this is the same issue, is that when Apache perform a 'self cleanup' on it's children - killing off the older children making room for new ones, it gets caught into an internal loop that sky rockets the load in a matter of seconds - which is why this is hard to catch.

        What I would suggest to do is to remove the current cron that is restarting Apache every 6 hours, upgrade the server to using Apache 2.2.23 along with PHP 5.2/5.3 and 5.4 - and then also run a script that will watch the server, if the load goes above 10.0, it will log as much information about Apache it can along with the incoming connections and restart the service as well after logging is completed as to avoid the server crashing. With this extra information we can determine if it's this particular bug or not, and then patch Apache from there - from the available logs on the server at this time, it does seem as though this is a likely possibility.

        As mentioned, we do not know what triggers this issue with in Apache, and does seem to be very rare - I've only seem this happen on 4 systems fleet wide, with almost 3000 servers.

        As a note with upgrading PHP, along with Apache - php5.3 will become the default, as 5.2 is the default now. Some websites do not like php5.3 and will not function with it. To resolve this issue is simple and quick, although there is a possibility that it would cause some down time which is why I mention it. Note that we could avoid upgrading Apache all together as well, and just utilize the script to catch what is going on and potentially apply the patch. Could you confirm how you would like to proceed with this, and if so confirm your security questions for us as well?

        What is your mother's maiden name?
        What city were you born in?

        Note, that in order to perform the patch it self - we would need to mark your VPS with the no-updates option. Although, this only applies to internal updates - not updates from our upstream provider CentOS for the operating systems packages it self.

        Thanks,

        We then went with the upgrade and have had no further issues on that VPS so we deployed it across all 7 we have with them.

        I personally find that you have to shout and scream at Site5 to get anything resolved ;o)
          • 13808
          • 61 Posts
          opengeek: I had wondered as much about the dodgy code front and of course, spiders wouldn't show up in the Analytics reports. Webmaster Tools is showing some 404 errors (which are handled with a 404 page) but nothing more sinister than that. I've just enabled some more raw logging in CPanel - can you think of a better way to check what might be going on other than cross checking the page loads reported in the access logs around the time of the crashes?

          paulp: that's uncannily similar to our situation and could explain an awful lot. Obviously if this was happening, increasing the server capacity would only get us so far. Thank you so much for weighing in!

          Could you please advise:

          • do you happen to know if the number of PHP processes rocketed during this load spike?
          • are you managing your own VPS or are they managing it for you?
          • roughly how recently did this all go down?
          • are you happy to give me something I can quote (ticket number or something) via PM that I could use to ask specifically if this seem like a similar situation?

          Thanks again for speaking up!
            • 22840
            • 1,572 Posts
            Hi Josh,

            Yes php processes rocketed during the load spikes and they manage the VPS's for us, it was very weird with regards to how often it happened, for example it wouldn't go down for a couple of days but then it would be anything up to 10 a hour, then good for a few hours and then back down again several times

            I'll PM you a couple of ticket numbers now.

            Cheers
              • 13808
              • 61 Posts
              That's exactly how we've found it. Exactly. I'm so hopeful that this will help sort some things for us - easily the strongest lead yet! Thank you! [ed. note: jcurtis last edited this post 13 years, 5 months ago.]
                • 13808
                • 61 Posts
                After a little bit of toing and froing with Site5, they've updated Apache for us. The load spike and crash has happened twice since then so it doesn't seem to have fixed anything this time around.

                I've dialled up all the levels of logging so that I can see what's going on in terms of access around the time of a crash. Hoping a pattern will form that points me to some buggy code I can fix.