I have a site running on Site5 in VPS1 (memory limit 768mb RAM). It's running 2.2.6-pl and is has just under 2900 resources. The site has been running smoothly for months and growing steadily in content to the current point.
Recently - starting about a month ago after 6-7 rock solid months, we are experiencing some pretty weird behaviour which I'm having issues getting to the bottom of. The site is taking down the VPS it's in frequently (sometimes makes it a few days - has gone down 5 times in an hour) by filling the memory which causes caching which causes the server load to go through the roof and the machine to become unresponsive to the point where it needs to be forcefully rebooted.
I'd fully understand this if the site was busy, but the most people I've ever seen the site serve at one time is 8 and those are 8 casual browsers - not 8 simultaneous hits on the site. There's almost no-one on the site. And serving those 8 people and me working in the manager, everything is fine.
Then there's been times where I've watched 'top' on the machine and real time stats in Google Analytics and the server has kicked up numerous PHP processes to the point of killing the server all while Google reported zero people on the site. To this end, as best I can tell it doesn't appear to be traffic load related.
I'm running a number of extras on the site, but nothing that accounts for multiple processes that seem to start from no traffic (cron etc).
Site5's advice is to upgrade to raise the memory ceiling (which I'm open to), but I don't feel great about just upgrading when the site is buckling under no traffic and seemingly at random. In my mind there's no guarantee the upgrade will fix the problem when I don't know why it's causing it.
- I've asked to see the contents of /var/log/messages but because it's a managed service they're not open to sharing the contents.
- I've used Executor to optimise the page query times right back and this has led to no reduction in server spikes.
- I've checked the MODX error log repeatedly - all it reports is being unable to cache static resources.
- I've setup a cron (via CPanel) which logs the load once per minute and the results of watching it for a week shows it going down at all different times of day so it's unlikely to be a daily maintenance process or anything like that.
Does anyone have any leads to help with further diagnosis? Is there something else valuable I could be logging somewhere or watching for?
The alternative is to upgrade or move but as outlined above, I'm not convinced about that as a solution when I can't follow why a load of 0 visitors can crash the VPS while a (still low) load of 8 visitors and me in the manager won't.
Any suggestions very welcome!
-
MODX Staff
- 10,725 Posts
It's possible that robots are spidering your site and hitting a Resource that is broken, or executing a really bad, long-running query. I recently saw a site this was happening too because their site map script was timing out. Just some thoughts, but my gut says you have some bad code on there rearing it's ugly head.
Hi Josh,
Yes php processes rocketed during the load spikes and they manage the VPS's for us, it was very weird with regards to how often it happened, for example it wouldn't go down for a couple of days but then it would be anything up to 10 a hour, then good for a few hours and then back down again several times
I'll PM you a couple of ticket numbers now.
Cheers
That's exactly how we've found it. Exactly. I'm so hopeful that this will help sort some things for us - easily the strongest lead yet! Thank you!
[ed. note: jcurtis last edited this post 13 years, 5 months ago.]
After a little bit of toing and froing with Site5, they've updated Apache for us. The load spike and crash has happened twice since then so it doesn't seem to have fixed anything this time around.
I've dialled up all the levels of logging so that I can see what's going on in terms of access around the time of a crash. Hoping a pattern will form that points me to some buggy code I can fix.